You know English doesn't work like that. The word you're saying only becomes clear with the surrounding context. Eg, 'there' vs 'their'.
But this is still possible to do if you track the whole run of text. You could replace all of it each time so it LOOKS like it’s streaming but earlier words also change. I’m hoping the streaming models do this eventually.
I believe the built-in iOS dictation already does this.
Handy already supports streaming transcription models, and you can see the words in the small Handy pop-up while you are talking.
So in general this definitely works. Handy is just missing the feature to insert these streamed words into the app where the cursor is.
Model should be able to understand where logical sentence ends, to stop buffering, and optionally rewrite some of the test that has already been output.
It may be interesting to have it immediately insert the words, even if they are wrong, and when a sentence is finished, replace what has been written with the final corrected sentence.