logoalt Hacker News

zmmmmmtoday at 7:05 AM2 repliesview on HN

> It jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation.

I'm confused, videos contain images and audio ...?


Replies

ibottytoday at 7:19 AM

That's most likely a disagreement on terms. In the media world, video is only the moving images, not audio. This is separate from images, that are meant to be still images.

PxldLtdtoday at 7:22 AM

It's more a comment about the feature detection I think; all image, video and audio input contribute to the same weights/activations that can produce image, video and audio output.