logoalt Hacker News

pbuiyesterday at 6:14 PM5 repliesview on HN

I'm not sure what to think... I uploaded a PyHPC workshop paper I wrote in 2011 and it said 27% machine.

I also uploaded my PhD dissertation from 2012 and got back 40% machine, which is just barely below the 42% threshold.

I don't publish anymore... but does this mean I wrote like a LLM or did LLMs learn from me? :p

Update: I also uploaded a IEEE CLUSTERS paper I wrote in 2015 and it came back 74% machine written :|


Replies

NitpickLawyeryesterday at 6:41 PM

When "detectors" first started popping up all over the place, all of them rated the declaration of independence as 100% AI written, so... Yeah, these things just don't work. And what's even more dangerous is that people that don't understand how any of it works use these tools, and accuse people of using AI, sometimes with grave consequences. Students have been through this, at all levels of education.

show 3 replies
sean_pedersenyesterday at 11:59 PM

"Any attempt to build AI generated content (deep fake) detection systems is flawed, since the outputs of such a system may be used to train an even better fake data generator. This leads to an equilibrium state of digital uncertainty: nothing in the digital realm can be deemed as real anymore - only as digital. I do not care if a digital artifact is human or AI made - I only care if it is useful to me. Useful content is on point, factual and at best surprising (teaches something new)." - https://seanpedersen.github.io/posts/digital-uncertainty/

miohtamatoday at 5:33 AM

It means machine text and human text cannot be distinguished from each other.

It's just text.

dgellowyesterday at 7:54 PM

Could it be that your papers are literally in the training set?

cansofgreaseyesterday at 6:26 PM

It's been trained on a work and then distilled, the false positive noise has to be absurd.