I find their (your?) example of using it as a separate API for expressions a bit odd. If you're already outputting voice surely other metadata could be outputted at the same time?
In text mode I'd use structured output for this, or even get away with instructing the model to [emote] or even just map emojis to emotes.
Great demo though, I'm going to try out the image stuff now. Super rad that it's multimodal decision making; "does this pipe need to be inspected?"->"yes","I want a closer look","no" etc!
Edit: impressive stuff! I gave it a bunch of examples:
- things that humans should/should not eat and it got this correct (including rejecting rat poison)
- probability staff member should be called to a train platform (people standing vs. child looking over the edge)
I find their (your?) example of using it as a separate API for expressions a bit odd. If you're already outputting voice surely other metadata could be outputted at the same time?
In text mode I'd use structured output for this, or even get away with instructing the model to [emote] or even just map emojis to emotes.
Great demo though, I'm going to try out the image stuff now. Super rad that it's multimodal decision making; "does this pipe need to be inspected?"->"yes","I want a closer look","no" etc!
Edit: impressive stuff! I gave it a bunch of examples: - things that humans should/should not eat and it got this correct (including rejecting rat poison) - probability staff member should be called to a train platform (people standing vs. child looking over the edge)