logoalt Hacker News

zzzeek • yesterday at 11:57 PM • 9 replies • view on HN

I follow anti-LLM discourse quite a lot, and across the main bulletpoints: energy/carbon emissions, content worker harm, job displacement, deskilling, mental health effects, and copyright/plagiarism, the plagiarism one seems to have the most attention, and it's also the most solvable, if there were only more serious effort on ethically sourced models that can actually do the real science / math / code work that is what LLMs are best at. The whole world of LLMs to create videos/books/art/literature is where most of the offense is (the video/imagery side of it is where most of the energy/carbon emissions problems are too. and content worker harm).

I really wish there'd be a split among these disciplines (science/math/code vs. videos/art/literature) - one is vastly more problematic than the other.


Replies

UqWBcuFx6NV4r • today at 12:04 AM

Yep. I’d probably be a lot less chastised in some circles for using Claude Code at work if it wasn’t misconstrued as being in support of, I don’t know, encroaching on the hypothetical commissions of a chronically online instagram furry artist or something.

It is very tiring to say “I don’t necessarily disagree with you about AI ‘art’, but in my field—which you do not understand, and in which the underlying build process is often not the creative output—AI presents very real productivity gains” for the umpteenth time.

I am skeptical of there being sufficient data to build “ethical” training datasets, and I’m confident that much of the same contingent will (somewhat rightfully) argue that ‘second-generation’ copyrighted AI material has already irreversibly made its way into every modern dataset.

➕ show 7 replies
altermetax • today at 12:12 AM

I don't really see the difference, code is protected by copyright (or copyleft) as much as art is, and yet the LLM scrapers use it without scruples. Same goes for math and science publications.

➕ show 1 reply
CapsAdmin • today at 12:33 AM

In my experience, the science/math/code crowd don't care about copyright/plagiarism as much as the video/art/literature crowd, so the first crowd turns a blind eye to most of the latter crowd talks about.

Code can be art, and copyright/plagiarism is real. It sort of boils down to how much it bothers us.

harimau777 • today at 1:04 AM

I don't think it is likely that they could get enough data without stealing. It would be incredibly costly to have to pay artists to church out art just to train an AI.

hardbass • today at 12:59 AM

I hope thats possible but I am not sure if proper knowledge of science can be had without also learning other literature (and vice versa).

asa123 • today at 12:02 AM

i think math is close to art (just to be a contrarian, but kind of really)

its somewhat funny that math people are in a conundrum as to support or not support but this might partially be because some wish to believe that math itself is and can be useful and therefore accelerating is good

but the art people have no such delusions so they’re just strictly against

imo proof writing is more akin to art than coding/tech but…

vouaobrasil • today at 12:11 AM

> I really wish there'd be a split among these disciplines (science/math/code vs. videos/art/literature) - one is vastly more problematic than the other.

I disagree that they can be separated. Practically, I think they can't. Because the mere invention of new tools inspires even more AI advancement and that in turn will cause the other side (artistic side) to degenerate even more.

I'm anti-LLM all the way, 100%, no exceptions. Zero tolerance.

➕ show 2 replies
PunchyHamster • today at 12:15 AM

They can't be ethically sourced and good at the same time.

The current models intelligence depends on massive training dataset of essentially stolen data

➕ show 2 replies
singpolyma3 • today at 12:10 AM

Models aren't in school so plagiarism doesn't really apply

➕ show 1 reply