arxiv:2605.20712
๐ In a Training Loop
Kavya Manohar
AI & ML interests
Speech Recognition, Low Resource Languages, Malayalam
Recent Activity
new activity 18 days ago
adalat-ai/fleurs-ro:v2.0: add Telugu config (466 rows, self-hosted Gemma 4 31B curation) + docs updated a dataset 18 days ago
adalat-ai/fleurs-ro posted an update about 2 months ago
Some bugs teach you more than they cost you. It's a small bug with a big lesson I keep running into: speech tools are built and benchmarked on English, so their limits are quietly calibrated for English. You only find the edges when you work in the languages they weren't tested on.
Learnt things the hard way, wrote it down so you don't have to.
https://huggingface.co/blog/adalat-ai/whisper-token-limit