Users often stack a transcription model on top to get the voice prompt, then decode to actions. Think of Alexa and Siri.
Thank you.
Thank you.