logoalt Hacker News

lazideyesterday at 6:09 PM2 repliesview on HN

I don’t think they will improve, there is too much incentive to poison the datasets going forward.

A lot of the models up to this point have been benefitted - like Google did - from essentially ‘pre SEO’ internet.

Now the same tools are being used to generate nigh infinite good sounding bullshit, which poisons the dataset in all sorts of hard to detect ways.

To add insult to injury, the human experts are also not as. Naive, and have many incentives to poison their own input in subtle ways too.


Replies

brokencodeyesterday at 6:34 PM

I seriously doubt that data set poisoning will be a real limiter in model performance.

For one, if your website/book is poisoned, who is going to trust it for anything at all, much less for training models?

For two, all the major AI labs hire or contract for subject matter experts to create curated data sets, evaluate model performance, etc.

Unless they hire malicious experts, this will provide a growing, high quality data set that should drown out any poisoned pretraining data.

show 3 replies
rvnxyesterday at 6:16 PM

Human doctors use LLMs to diagnose too

OpenEvidence claims

    "More than 40% of U.S. physicians use it daily, and it handled around 20 million clinical consultations per month. Over 100 million Americans were treated by a doctor using it in 2025."
https://www.cnbc.com/2026/01/21/openevidence-chatgpt-for-doc...
show 3 replies