logoalt Hacker News

didibuslast Monday at 2:37 AM1 replyview on HN

Most of the models are multi-modal and trained on images no? That's what they claim at least.


Replies

skygazerlast Monday at 5:42 AM

You’re right. Modern frontier models are now multimodal. I used often as weak a hedge, because I know at least his gpt3.5 turbo and llama3.1 generated pelicans were from text only models without image training. The chinese models are interesting, because before their vision models existed they may have been distilling text only models from text output of American vision models, so they could have benefited from the teacher model’s vision capability without being vision models themselves.