If the output of this is even somewhat coherent, it would disprove the argument that mass amounts of...

InvisibleUp • yesterday at 5:49 PM • 1 reply • view on HN

If the output of this is even somewhat coherent, it would disprove the argument that mass amounts of copyrighted works are required to train an LLM. Unfortunately that does not appear to be the case here.

Replies

HighFreqAsuka • yesterday at 6:18 PM

Take a look at The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text (https://arxiv.org/pdf/2506.05209). They build a reasonable 7B parameter model using only open-licensed data.

➕ show 1 reply

alt Hacker News

Replies