What the demo does not do is show streaming output of transcribed text as we are speaking and recording (before we hit stop). That is an essential feature IMO for most general purpose live STT apps.
My mind is boggled by how many implementations miss this.
Handy has Nemotron Streaming and it works fabulously, FWIW. I’ve vibed a kind-of-working Deepgram API server into it but haven’t gotten around to finishing it. It’s something that should exist IMO!
This is the main reason I lean on Deepgram over local services.
What do streaming implementations do when a bigger context reveals a different interpretation/parse? When I use Whisper in the terminal, I can see it going back and correcting itself. Are corrections off the table for a true streaming transcription?