logoalt Hacker News

ameliustoday at 12:35 PM2 repliesview on HN

What I want is a model that is trained with data that is openly available, where the data is curated by academia. I don't want corporate crap in my AI (unless it has been filtered properly).


Replies

thevintertoday at 12:37 PM

I'm ready to stand corrected, but I'm pretty positive that such a process would require 1) an insane amount of work and 2) wouldn't produce anything close to SOTA results because of the lack of training data.

It is my understanding that - sadly - the insane amount of copyrighted works and corporate crap is a prerequisite for having a corpus that is big enough

show 1 reply
dorkypunktoday at 1:26 PM

There are models that do that, for example the Olmo family of models, although they have Gemma 3 performance levels for that matter.