logoalt Hacker News

brianyu8 • yesterday at 5:06 AM • 1 reply • view on HN

Yes! Check out the 20 second mark of this video for a demo of how fast image tasks can be: https://x.com/OpenAIDevs/status/2107573382229188645


Replies

fennecfoxy • yesterday at 10:35 AM

I find their (your?) example of using it as a separate API for expressions a bit odd. If you're already outputting voice surely other metadata could be outputted at the same time?

In text mode I'd use structured output for this, or even get away with instructing the model to [emote] or even just map emojis to emotes.

Great demo though, I'm going to try out the image stuff now. Super rad that it's multimodal decision making; "does this pipe need to be inspected?"->"yes","I want a closer look","no" etc!

Edit: impressive stuff! I gave it a bunch of examples: - things that humans should/should not eat and it got this correct (including rejecting rat poison) - probability staff member should be called to a train platform (people standing vs. child looking over the edge)

➕ show 1 reply