One larger problem here is the value of a research paper is rarely the specific knowledge it adds but in the process of researching that adds to the collective knowledge+experience of those involved, especially training graduate students. AI papers shortcut this entirely. Academia has a lot to answer for this too by making papers the currency of success. AI generated papers are almost shortcut learning at a full system level.
> Separately, our group has been exploring approaches along these lines to make such evaluations more scalable
Actually, that sounds like an interesting idea for peer review in general, to include an interview between referees and authors. If it saves one round of rebuttals/reactions, it needn't even consume a lot more of everyone's time if you're doing those things properly. What it would undermine would be blindness, but something's gotta give, and it was already on its way out.
This is the journal's policy on LLM use by authors [1]:
> LLMs may be used as general-purpose assistive tools. Whichever tools are used, authors are fully responsible for content on which they are listed as (co-) authors. This includes, but is not limited to, content generated by LLMs that could be construed as plagiarism or scientific misconduct (e.g., fabrication of facts). Low-quality contributions (be they submissions or reviews) that appear to be largely LLM-generated will be closely examined for evidence of the issues mentioned previously, such as scientific misconduct. LLMs are not eligible for authorship. We will periodically revise this policy as new information about the use of LLMs in the scientific process becomes available.
While it doesn't outright encourage using LLMs, it's right at the door, and IMO a policy this weak is actively contributing to the problem the article's author is complaining about. In my opinion any policy weaker than "using LLMs to generate any part of your submission is not allowed and considered a serious breach of ethics" is insane. People like to say that such policies are unenforceable, but that's really not the point (at first), since there are other things like (somewhat ironically) p-hacking that are pretty hard to detect but still widely recognized as unethical. We haven't exactly solved p-hacking either, but at least most of us can agree that p-hacking should be eliminated.
It's hard for me not to read between the lines here. Maybe it's the tinfoil talking, but it being a machine-learning journal, it probably embodies a generally pro-AI philosophy, and thus may not want to discourage too much of it...
It may also be worth noting that this journal apparently uses AI itself on the reviewing side [2]. I'm not claiming this is super unethical or anything as long as the main review is human (although I have concerns), it probably should be part of the conversation.
[1]: https://jmlr.org/tmlr/editorial-policies.html
[2]: https://medium.com/@TmlrOrg/ai-reviews-at-tmlr-for-assessing...
IMHO (not a paper writer, but read a lot during my grad school years), the Genie is out of the bottle. The only way forward, as I see it, is using LLMs for reviews also. Basically, filter all submitted papers with an LLM and ask it to summarize it, find the biggest weaknesses and main strong points, etc. that a human can then use to review the paper. Basically, LLM-as-a-reviewer .
Personally, I would love to see a conference where people are explicitly encouraged to use LLMs for doing the work and writing the papers, and LLMs are used to review them too.
Peer review has its historical issues, but the landscape of science and science-publishing has changed. New problems of authorship and authorial-understanding are now challenged by LLMs writing (at least) good sounding papers - some of which might be of acceptable quality in subject (I am not against AI in the sciences; some of the math work has been great). On the other hand: I am against authors not understanding their own work. High repute journals may need to add "oral exams" to the paper acceptance process...
I'm not familiar with the world of academic publishing, so I want to ask: how is the industry making sure that submissions aren't at least partially AI-generated?
Is it standard practice for authors to have to defend their submissions via interview like this? If not, why not?
Does the vetting process vary with the quality of the publisher?
As an outsider, it's extremely worrying that anyone would even attempt to submit an AI-generated paper for publication in an academic journal. At that level I would have assumed literally everybody should know better than to even try.
Is it really “research” - as in expanding human knowledge - if nobody understands it? The point is deepening human understanding, not producing research papers
> All reviewers, Action Editors and Editors-in-Chief for TMLR are unpaid volunteers.
He is not an unpaid volunteer.
He's an associate professor at the prestigious Carnegie Mellon University. He is not paid by the journal, but he is paid a salary by the university, and the university expects that a small part of his academic work is to serve as an editor in academic journals.
FYI in case the author is reading, https://www.cs.cmu.edu/~nihars/preprints/greCAPTCHA.pdf is a dead link.
EDIT: I found a live link on arxiv https://arxiv.org/html/2609.20481v1
Sounds like I shouldn't be so hard on AI when it hallucinates things based on this data? :)
I wonder how the ratios would change for papers at different parts of the review process. For what fraction of published papers are the authors unable to answer basic questions about them?
I think authenticity and trust will command a (larger) premium in this new age of slop.
The article highlights how only one out of ten paper’s authors were able to answer questions thoroughly and at a high level. This indicates an overwhelming percentage of authors are slopping up their work with AI and submitting it without even reading it.
The Medium comments on this post are also on point. Running the same experiment with accepted papers is a good control. Running a similar experiment with reviewers would be interesting, but more obnoxious because they are not being paid.
I would keep a private blacklist (shadow ban) the authors who wasted several hours of a reviewer's time to prove they were not legitimate. The existence of such a list would be problematic, though.
Could the same system we use here be applied? Accepted authors could "vouch" for "dead" papers in case they were "auto-killed"?
This system is broken and providing more evidence that it is broken isn't much of a step towards fixing it.