I don't think the challenge with speech to text was size of the binary. In my experience the challenge is understanding my 84 year old Croatian father with a sagging mouth after a stroke, when he's trying to write his autobiography.
I just setup Windows speech to text for him last week and it's great to see how he can write an entire page in 10 minutes, it would take him days using the keyboard.
But every single sound he makes with his mouth ends up on the page too.
> I don't think the challenge with speech to text was size of the binary
The usecase for small models like this is making on-device STT/TTS more accessable. This is important if your usecase is sensitive to either privacy or latency, but this comes at the cost of quality.
My experience has been that these small TTS models are unexpectedly good if your audio is in distribution (western accents, higher quality audio, common vocabulary), but pretty quickly degrade as you move outside of that. They often dont support more complex features such as diarization, multilingual, or realtime streaming either.
There are many different challenges, each requiring their own solution. I, for one, really miss the old Google Assistant on my Android phone. It would very reliably play most songs that I wanted to hear on Spotify. Gemini fails at this almost every time, and is significantly slower. It's actually a difficult problem, as the songs people want to hear are regularly being released, are often associated with uncommon names, or have words in unusual orders, so normal LLM style tools just don't cut it.
I hope you will find a solid solution for your father.
I was researching STT for people with speech disorders two years ago and essentially everything was boiling down to three problems at the end of the day - data scarcity, irregularity of way of speaking and thus constant ambiguity in translation, and individual differences in speech patterns among patients.
Have you tried playing music from his childhood for 20 minutes a day before his writing sessions?
In some cases, this may improve function for a few hours. Best regards =3
i mean for something this small, it can be fit into a l3 cache on a cpu and be essentially always on various purposes
Unrelated: I love your username.
Sorry about your father. He needs a dictation model, not a general purpose speech-to-text model. They ignore umms and ahhs, change things like “an elephant, no a monkey, went up the tree” to “a monkey went up the tree,” support saying punctuation aloud sometimes, etc.
Gemini team just released Gemini 3.5 Transcribe that’s supposed to be good at this; it’s available via api: https://blog.google/innovation-and-ai/models-and-research/ge...