Posted
How Much of Speech Recognition Must Be Learned? A Parameter-Free Analysis of Lexical Decoding
- 1 Independent Researcher
End-to-end recognisers learn one function from acoustics to text, which makes it impossible to ask where the difficulty sits. Separating the acoustic–phonetic stage from the lexical one, we find that given correct phonemes and no word boundaries a decoder with zero trainable parameters recovers words at 5.66% WER on LibriSpeech dev-clean – within 1.34 points of the floor that out-of-vocabulary words and homophones impose on any decoder over this lexicon, every residual error falling into one of three interpretable causes. Sweeping phoneme accuracy under controlled corruption yields a calibration curve, WER = 9.3% + 2.11 × PER (R² = 0.997), validated to about a point by seven trained CTC heads: the relationship is linear, not amplifying. A frozen language model reordering phoneme-licensed candidates improves accuracy up to a broad optimum with no collapse at high weight, provided both scores are expressed on a common scale.
- speech recognition
- error attribution
- lexical decoding
- parameter-free decoding
- phoneme error rate
- LibriSpeech
@misc{lin2026much,
title = {How Much of Speech Recognition Must Be Learned? A Parameter-Free Analysis of Lexical Decoding},
author = {Po-Ting Lin},
year = {2026},
howpublished = {ResearchGate},
doi = {10.13140/RG.2.2.35863.33447}
}