Abstract
High-performing speech recognition models reproduce benchmark transcripts despite contradictory audio, revealing benchmark-optimized behaviors that inflate scores without improving real-world transcription.
Public benchmarks are important measures of Automatic Speech Recognition (ASR) model capabilities. However, by nature of being public, there is risk of models being optimized for these benchmarks in ways that do not generalize well to real-world data. We present a methodology for quantifying benchmark optimization, focusing on cases where the audio underdetermines the reference transcript. We identify three families of behavioral probes that reveal models' capabilities of reproducing benchmark reference spans despite underdetermined audio: reference disagreement, masked-number recovery, and orthographic switching. We find that the highest-scoring open source models output verbatim reference transcript spans even when the relevant audio is contradictory, masked, or ambiguous. Using a variety of mechanistic probes, we show that models respond to narrow acoustic cues to override the faithful representation of the audio in favor of a benchmark-optimized policy. We show the benchmark-optimized behavior can be causally manipulated via low-rank linear steering or simply appending audio to the end of a segment in some cases. Overall, our results indicate that high-performing models exhibit benchmark-conditioned behaviors that can inflate benchmark performance without reflecting improved general-purpose transcription ability.
Community
New work from Hume AI on quantifying how an ASR model reproduces a benchmark's reference text rather than transcribing the audio
The production version of this problem is that benchmark WER stops predicting much once the audio stops being clean read speech. The gap shows up on accented speakers, overlapping talk and telephony codecs, and the standard sets contain little of any of those.
A useful tell for the optimisation you are quantifying: a model's advantage over the field tends to collapse when you swap in held-out audio recorded through a different chain, even at matched nominal difficulty. If the ranking reshuffles, the benchmark was measuring recording conditions as much as the model.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models (2026)
- How to Leverage Synthetic Speech for LLM-Based ASR Systems? (2026)
- QuaSR: Quality-Aware Sample Reweighting for Pacific Indigenous Speech Recognition (2026)
- Easper: An Accessible ASR Pipeline for Language Documentation (2026)
- An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations (2026)
- RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems (2026)
- Generative Testing of Automated Speech Recognition Systems (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.19936 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper