Experiment
Quantisation Trade Offs in On Device Speech Recognition
Where does word error rate break under quantisation, and is the degradation uniform across clinical vocabulary?
- Kind
- Experiment
- Year
- 2025
- Status
- In progress
01
Abstract
A controlled comparison of quantisation strategies for speech recognition models running on consumer hardware, measuring word error rate against latency and memory. The finding that shaped the product is that degradation is not uniform: general conversational accuracy holds up considerably better than drug names, dosages, and abbreviations, which is exactly the vocabulary a clinical system cannot afford to get wrong.
02
Methods
- Word error rate analysis
- Post training quantisation
- ONNX Runtime
- Latency profiling