I just don't understand people saying "but a human learning from a book isn't illegal".
How do people not understand that some laws only make sense at a certain scale? One human learning from resources and being added to the labour pool is not the same as an infinitely copyable entity doing the same thing. One has negligible impact on the demand for the original, and the other replaces 99% of the demand."
And creating a rule that says you cannot train on any material unless the rights holder authorises it via license is not complicated. That will creat a amrketplace where creators can decide the price for their content. It's just inconvenient.
“Information wants to be free“.
It’s not “theft of labor”; the work was already done. If anything it is theft of “intellectual property” (aka “copyright infringement”), if you believe that is a thing, but not of the “labor” that went into it.
My personal take: anyone producing content, everyone’s creativity, is fed by something that others did before. We’re all standing on the shoulders of giants composed of previous generations and their “content’s” distribution and dissemination. I have an immense gratitude for all the labor before me that I was and am allowed to partake; without that, I would be nothing. Sharing information is an act of love; gatekeeping it is short-sighted greed. New technologies have always “killed” previous “labor”, out of which new opportunity grows. I just wished the collected data was public. I hope we all get a mega-leak at some point.
I wouldn't have a problem with working off the fruits of other people's labor because most of us are essentially doing that everyday anyway, the issue is that big tech companies (want to) reap all the benefit and create profit from something that should be accessible to everyone. Everything is getting privatized -- housing, water, electricity, and now, thinking and knowledge. We are heading towards a world where you have to pay even more excessive fees just for existing and for completing any basic task.
Google has been scraping everything from us since day one. Meta, Microsoft, Github, Slack, Reddit, StackOverflow, big and small, every single app that interacts with people uses our own data to make money and create walled gardens. I haven't seen a single one opening their silos to the world. That's our data, we produced it, you captured it and now you think it's yours
So no, your cries for regulating others because you are losing the race won't work this time.
If someone asked what is 'the largest theft of labor in human history' I would have thought slavery.
The most shocking point is that they have a Microsoft exec who knows what he's talking about.
In case of programming.
How much do the current LLMs invent solutions for user tasks, how much they just copy and adopt existing open-source solutions from from Github and other code repositories?
This not a problem for open-source code under permissive software license, but works derived from open-source code with copyleft software license should be also under copyleft license.
Could the biggest commercial benefit of LLMs be just working around limitations of copyleft licenses?
What is the monetary value of human work put into copyleft software and later used to train LLMs? It's hard to estimate, but the study "Estimating the Total Development Cost of a Linux Distribution", estimated that it would cost $1.4 billion to develop the Linux kernel alone.
https://consortiuminfo.org/metalibrary/estimating-the-total-...
I have this weird vision of an alternate reality where governments (say, National Archives) are the ones creating the models as a public service and then the rest of the industry is just commoditized pricing of hosting them, competing with value add bits. And we’re on here reading articles about how the latest release of the EU model does a better job generating maps now and the new Canadian model seems to apologize less and whatnot.
Yet we still tell students to buy textbooks. The individual must always pay. The corporation can do whatever the hell it wants.
The hypocrisy of this new world is already catching up to us.
The problem isn’t just “stealing the fruits of human labor”, it’s also driving down the value of human skills and even taking away human jobs.
It's humanities collective knowledge and work. That's why nobody should ever buy the narrative of distillation being a crime or theft. It should be a human right to distill these models. Distillation should be being provided as a service.
I remember when open source software, and Linux in particular, was the threat to the world according to Microsoft execs.
Are we considering what is the shelf life of information?
If you build a building, the expense on materials determines longevity. If you build a city. The robustness of government and the economy in it determines the property taxes and value of property over time.
If you make or cook food. The majority of the nutritional value of it goes to the initial consumption. Once the food has stayed out without refrigeration it is taken over by bacteria and fungi. Refrigeration seems to be paywalls. Once the information is out it accumulates at exponential rates - the amount of text on the internet does not diminish but increases. Some people may “prune” old content away, but that is rare. Human attention is somewhat a fixed number. Thus text left out is not consumed, but sits idle and decays in accuracy and value over time. The fresh content of valuable should be in a fridge. If not valuable it is released - thus scavengers and those hungry and motivated to dig can consume it. If spammy and sales-y / propaganda-y which a lot of content farms are doing, the goal is for it to be consumed by the masses and push the zeitgeist to buy its premise. That’s Sugar or addictive shelf-stable junk foods. AI model companies are the bacteria / fungus/cockroaches/rats of the information dumpster. They sneak out any remaining energy from content that would otherwise be buried by other content and try to give it a second shelf life - one reachable and accessible and consumable by humans. They make alcohol. Alcohol is addictive. Ir may mess with your brain - it may make you lazy. It will sneak in bad decisions because it lowers your judgement. It is repurposed food, not the one you are used to injesting. It may even have its own agenda - depending on how the information is reprocessed. And it also has a shelf life since humanity continues to have new insights and people keep getting new alcohol brands to try.
The largest theft of labor in human history … and it’s to do away with the laborers by making a device that produces labor substitute, with full awareness that the substitute produced is not fit for the purpose of making more such devices.
It’s like burning all the crops for heat, which you use to boil the oceans for salt, which you use to salt the earth so no more crops can grow.
If AI wants to destroy humanity it better get its boots on, or else AI companies might get there first.
I think so too. The only way to redeem this theft would be to force all AI companies to open source their models if they cannot prove that copyrighted material was not used to train them.
I think it’s more like ‘The absolute maximum possible degree of theft’ there can’t be larger, it’s everything current and past.
I understand the sentiment and partly agree. But also, the original has not gone anywhere. You're free to accumulate knowledge in the old way just as before. So maybe it's not theft of knowledge that we should be angry about, it's something else harder to define.
I remember techchrunch.com making the argument that IP Infringment != Theft in the music piracy era.. how quickly the tide turns :)
Microsoft is one to talk.... Remember when MS trained copilot on all your github code?
Are those factory workers we saw photos of now, wearing cameras to capture the movement of their hands stitching getting compensated for a generations worth of wages? Do they even have any choice but to give away the copy-right to their labor?
They are not selling the information. They are selling a service for easy access to that information.
These are two different things.
Note: am not an AI fanatic.
I don’t mind these companies scraping my content.
But for love of god, my blog changes at most every couple months. You don’t need to scrape it every few minutes.
Correct solution here is to make sure royalties are embedded in the AI responses (and work output). These should be appropriately priced and go back to the owners of the IP. If the IP is no longer owned then it can be free use.
There should be a carveout for non-profit or government AI.
If corporations weren't already owning the consumer, with AI it does this by many orders of magnitude. If something isn't done to prevent AI from being used to farm the masses for data, we will be living in a sci-fi dystopia without a doubt.
I think it's okay to advance humanity, but they can GTFO when they then try to ban distilling and open models.
Yet he also helps destroy all those jobs. The thief is calling "Catch the thief!".
He does not see the moral dilemma here?
As someone who thinks that genAI is harmful, I deeply resent that any of my work has been used to help train it. I will never forgive these companies for forcing me to contribute.
There will be a point where companies will not need to scrape any content. Agents will create endless streams of probes, and they will end up solving all kinds of knowledge problems.
Copyright infringement, if this even were that, is not theft. Chattel slavery is the largest theft of labor in human history.
How is this different from Microsoft scraping to build Bing?
Honest question. There is a line in the sand somewhere apparently.
You have to admit there is now some lovely schadenfreude to be had from the whole ‘Chinese free LLM companies be stealing our theft! Stop them!’ whining.
AI overall is the ultimate piracy crime.
I wonder what a token cost would be if AI companies were to pay royalties to every author who made their business even possible.
I do not understand what "theft" they are talking about. Those AI bots were scraping publicly accessible internet.
Publicly. Accessible.
Of course there are some parts of the publicly accessible internet which host content that may be considered illegal or has been obtained illegally. If those AI bots used such content as well, it is fair to call it out as wrong, in my opinion. But that is a separate topic.
Blindly calling scraping of publicly accessible internet a "theft" is, in my opinion, disingenuous. Especially when coming from a company operating a web search engine. Which itself has its own bots scraping the same parts of the internet 24/7.
"The question of whether AI firms can legally use copyrighted material to train AI has no clear answer, but judges have been largely favorable to AI companies’ arguments that training constitutes “fair use.” This legal rule lets people use copyrighted work without permission in certain cases, like parody, news reporting, or criticism. Earlier this month, the Trump administration contributed a brief in defense of OpenAI’s unlicensed use of copyrighted material to train its LLMs. "
So training can make it legal as well. Interesting...
And here we are, just watching and doing nothing..
Throughout history we’ve been able to retell stories, to copy content, to create shallow clones or synthesis
It is only now in human history that we are able to create nearly perfect copies, and we’ve been taxed incredibly for this with overpriced everything.
Copying isn't stealing you babies
Spiderman pointing
Microsoft executives levelling "tone-deaf" up in realtime.
Imagine the parthenon marbles. When they were looted it was even a celebrated act, but they are still stolen in the british museum centuries later.
Information wants to be free and all that but there's a sense in which AI really is real intellectual property theft in an ethical sense compared to others and of _course_ it was Facebook who steals from everyone where Zuckerberg personally approved it
Their bots are also apparently the worst. Google does not put huge strain on your public-facing website (I think). Facebook does, they're incredibly malicious about it
Because it was, it completely defaced all copyright and similar laws, like there is ZERO ground to stand against China now regarding theft... it's so weird how this is being allowed.
But what about M$ owning Github and doing the same with its content? Github even did not deny scanning private repositories. (Gitlab denied the same when asked). So...
"The Net interprets censorship as damage and routes around it."
[flagged]
It's the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity. So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.
It is said that at the heart of every great fortune there is a great crime, so it should be no surprise that the most valuable companies on the planet will most likely result from this crime. And given that justice can be bought by those with the most money you can forget about anything coming of this.