Parameter count split is interesting:
740M total (270M text, 170M vision, 300M audio)
Makes sense...
Vision is just processing a still image (why it's by far the smallest). Text requires dealing with the entropy of human language. Audio is meaningless without time.
Makes sense...
Vision is just processing a still image (why it's by far the smallest). Text requires dealing with the entropy of human language. Audio is meaningless without time.