“Information wants to be free“.
It’s not “theft of labor”; the work was already done. If anything it is theft of “intellectual property” (aka “copyright infringement”), if you believe that is a thing, but not of the “labor” that went into it.
My personal take: anyone producing content, everyone’s creativity, is fed by something that others did before. We’re all standing on the shoulders of giants composed of previous generations and their “content’s” distribution and dissemination. I have an immense gratitude for all the labor before me that I was and am allowed to partake; without that, I would be nothing. Sharing information is an act of love; gatekeeping it is short-sighted greed. New technologies have always “killed” previous “labor”, out of which new opportunity grows. I just wished the collected data was public. I hope we all get a mega-leak at some point.
Just because “information wants to be free” is a thing people say doesn’t mean it’s true.
Well the future we seem to be getting is “information wants to be free for the first ten thousand tokens, then $1 per million tokens after”.
> Sharing information is an act of love
Most AI companies are not sharing it, though. They appropriated it and resell it.
There’s a difference between an individual creating something and the industrialization of creation. You can’t scale the creation of a single person 1000000x by the snap of a finger but you can with machines. This has severe implications.
The key problem is that IP is either proprietary to the creator or it is a commons type of situation.
Even if you agree with the former exploiting the commons for personal profit is... not good.
One could make the argument that if these LLMs were all open weight it would be okay, but to keep the result of the training private and proprietary is not fair.
I sort of agree, and i think strengtening IP Law is probably not great. But I do think it's very fucked that building generative ai is only possible by taking the works of countless artists and craftspeople and then the model produced from that data immediately gets deployed to destroy the careers of the people whose, work was vital to it being created, without compensation for them, while making a few evil nerds richer than god. I think if you work at one of these labs you owe an enormous debt to society and your earnings should be redistributed among it.
>"Information wants to be free".
I agree wholeheartedly and in keeping with that, I call upon frontier AI labs to release both their weights and training sets.
[flagged]
"I just wished the collected data was public. "
That's the entire contention here. It's a double standard. Companies will sue the living hell out of anyone taking their IP, whether it's code or art, yet they have no qualms taking all the data they need from anyone and everyone. It was already a problem before, i.e. artists getting paid very little for work that companies profit a lot from like musicians or digital artists, but now with AI it's on steroids.