Speech-to-text models predict the next token of text from the preceding tokens of text and the current tokens of speech.
Thanks, I did some learning and it fell more into place.
Thanks, I did some learning and it fell more into place.