logoalt Hacker News

PunchyHamster • today at 12:15 AM • 2 replies • view on HN

They can't be ethically sourced and good at the same time.

The current models intelligence depends on massive training dataset of essentially stolen data


Replies

mehrzad • today at 12:45 AM

While that is true, theoretically a regulation could be enacted that output tokens must focus on STEM research and other practical tasks and the LLM must refuse tasks outside of those areas, just as Claude disallowed cybersecurity tasks. Obviously this would never happen, but the theft of the training data wouldn’t matter as much if the usecases were less sinister.

zzzeek • today at 12:41 AM

openai and anthropic trained on actually stolen data since it was pirated datasets.

google OTOH already had a lot of this dataset in their possession (e.g. Google Books etc), still questionably licensed for how they used it, but not quite as bad. They did apparently break through NYT paywalls and stuff like that though, still theft.