Interesting. Wasn't Deepseek's founder saying that they had explicitly decided not to focus on multimodal models at all and were going text-only because they believed it was enough to achieve AGI?
I think you're thinking of Dario saying this about image generation.
I think you're thinking of Dario saying this about image generation.