- papers are read, summarized and digested by AI, because there are just so many papers at leading AI conferences that nobody has time to eyeball them all
We are very rapidly automating humans out of the academic publication loop here.
Idk if we're automating humans out of the publishing loop as much as rapidly automating the production of crap. I had a very similar experience reviewing for EMNLP recently.
We are nowhere near AI being able to judge the quality of research (in fact, one might reasonably state that even most humans can't really judge the quality of research). Most things in society are not like math: we can't automate (via verification) our way out of noise overwhelming the signal.
Folks are willing to entirely abuse the public resource that is faithful, honest reviewing. (This is unsurprising; the abuse of the commons / public resources has been rising for a long time). There isn't a good solution other than something akin to draconian social scoring to limit access to the reviewing system.
This is a legitimate question: should people have reputations? Should their behavior be made more visible publicly, both good and bad? How?
In a small contained society, where consequences are more directly affecting individuals immedately, these questions don't need to be asked, because they're inherantly answered. We now are a society with billions of people, and dire consequences sometimes deferred for a generation, or more. Part of our general failing is the lack of good answers to the above questions. For many people, there are rarely negative consequences for causing harm to others, and the rewards can be very great indeed.
"Consequences and reputations stemming from one's actions" isn't necessarily draconian social scoring. Without a structure for imposing consequences on wrongdoers, we're not a society, we're just monkeys flinging poo at each other.
For a more mathematically rigorous treatment, negative feedback is fundamental to Control Theory; it allows us to stabilise processes that would otherwise go out of control.
The problem is that everyone imagines some sort of just and fair arbiter of these things, when the reality is all of the social scoring and consequences and reputations rarely actually stem from one’s actions, and far more often stem from how much money someone has put in someone’s pocket, who someone knows, the color of someone’s skin, or what’s between someone’s legs.
Until we’re actually serious about treating people equitably, (not equally, as that would simply leave the lopsided power structure we have in place) we aren’t getting out of this.
People react negatively to this, because they fear a dystopian society of the sort we've seen in plenty of movies, and rightfully so.
But it's also worth pointing out that "consequences and reputations stemming from one's actions" is already the world we live in and always have lived in. Hell, even Hacker News has karma points, downvoting, shadow banning, and the like. There's no such thing as a society with zero consequences and zero reputations. The only real question is a matter of degree, structure, severity, reach, and various idiosyncrasies that differ across cultures.
So it would be nice to have a more nuanced discussion about this instead of treating it like a 0 or 1 decision.
People should get paid for reviewing. That’s the solution. Publishing companies rack in billions of dollars in pure profit exploring free labor. Once people actually get paid for reviewing, it becomes much faster and higher quality and you won’t need AI triage.
I think you are correct, its so hard to judge the quality of research that we have been using publication record/count as a proxy. Ultimately, having papers being easy to write is good, provided we find a better way of judging quality.
We just need journals run by an AI that charges other AI to read them, and then the AI run colleges can promote the AI with the most AI journal entries and citations.
My 2 cents: AI has improved writing considerably for non-native english speakers (in particular China). Writing feels more standardized/boring but easier to read overall. I hit fewer papers that are a pain to read. The most problematic aspect I see are semi-bogus claims i.e. sentences that aren't false, but don't quite feel right either. YMMV.
Validating existence should already be trivial: virtually all journal already have publicly-available indices which contain at least author information and an abstract. Combine that with DOI citations and you're basically done - even with closed-access journals.
Checking the content is of course a lot more difficult, but that doesn't magically become trivial with open-access journals: you still need to read and interpret what is being said in the paper and compare it to the claims being made in the citation. Granted, these days you could use AI for a first pass, but it's still going to be incredibly tedious work.
And yet the automation we pour trillions in, is the one that will do anything, including everything I find interesting and will never clean my own house.
I'm sympathetic to the idea: we should have an open, publicly queryable citation graph. Google scholar could very easily offer this at marginal cost near zero, but they won't.
I thought that it was. I thought journals were starting to implement full bans upwards of a year+ for those who don't honestly disclose AI usage in their work? If I'm not mistaken, arXiv is doing this as well? granted, disclosure is different from overuse, but it seems like a small jump to just go ahead and just ban not checking ones work! ...disclosed or not disclosed. my own personal view on it, is if you can't be bothered to spot hallucinations and other such errors, its not ai that is the issue, it is incompetence and laziness
> Both papers were accepted for oral presentations with the condition that they simply fix the hallucinated references.
I do wonder what truthfully could be on ai verification, if even one paper with such an error is accepted it sets the precedent you hopefully get lucky to not get caught (then again verifying for basic tells isn't the same verifying is this genuinely a worthwhile publication, but that's a separate matter)
"A lot of the content of this blog was initially drafted by an agent of some sort"
What? I mean who does this. My voice is my voice and it's literally never occurred to me to have an LLM do a first draft. I thought that was college kid stuff.
This is an oversimplification and I'll edit. Caleb and I had a conversation about our reviewing woes and thought it would be fun to do an interview style post, so we had Claude come up with some questions based on our convo. We answered the questions from scratch and had Claude proofread at the end.
At this point, in the field of AI research:
- papers are written by AI (as pointed out in this article, and as obvious to anyone who spends a while actually reading recent AI research)
- papers are reviewed by AI (NeurIPS is doing an AI assisted review experiment - https://neurips.cc/Conferences/2026/ai-reviewing-experiment - and I feel the trend is moving towards AI reviewers whether we like it or not)
- papers are read, summarized and digested by AI, because there are just so many papers at leading AI conferences that nobody has time to eyeball them all
We are very rapidly automating humans out of the academic publication loop here.
Idk if we're automating humans out of the publishing loop as much as rapidly automating the production of crap. I had a very similar experience reviewing for EMNLP recently.
We are nowhere near AI being able to judge the quality of research (in fact, one might reasonably state that even most humans can't really judge the quality of research). Most things in society are not like math: we can't automate (via verification) our way out of noise overwhelming the signal.
Folks are willing to entirely abuse the public resource that is faithful, honest reviewing. (This is unsurprising; the abuse of the commons / public resources has been rising for a long time). There isn't a good solution other than something akin to draconian social scoring to limit access to the reviewing system.
> draconian social scoring
This is a legitimate question: should people have reputations? Should their behavior be made more visible publicly, both good and bad? How?
In a small contained society, where consequences are more directly affecting individuals immedately, these questions don't need to be asked, because they're inherantly answered. We now are a society with billions of people, and dire consequences sometimes deferred for a generation, or more. Part of our general failing is the lack of good answers to the above questions. For many people, there are rarely negative consequences for causing harm to others, and the rewards can be very great indeed.
"Consequences and reputations stemming from one's actions" isn't necessarily draconian social scoring. Without a structure for imposing consequences on wrongdoers, we're not a society, we're just monkeys flinging poo at each other.
For a more mathematically rigorous treatment, negative feedback is fundamental to Control Theory; it allows us to stabilise processes that would otherwise go out of control.
The problem is that everyone imagines some sort of just and fair arbiter of these things, when the reality is all of the social scoring and consequences and reputations rarely actually stem from one’s actions, and far more often stem from how much money someone has put in someone’s pocket, who someone knows, the color of someone’s skin, or what’s between someone’s legs.
Until we’re actually serious about treating people equitably, (not equally, as that would simply leave the lopsided power structure we have in place) we aren’t getting out of this.
People react negatively to this, because they fear a dystopian society of the sort we've seen in plenty of movies, and rightfully so.
But it's also worth pointing out that "consequences and reputations stemming from one's actions" is already the world we live in and always have lived in. Hell, even Hacker News has karma points, downvoting, shadow banning, and the like. There's no such thing as a society with zero consequences and zero reputations. The only real question is a matter of degree, structure, severity, reach, and various idiosyncrasies that differ across cultures.
So it would be nice to have a more nuanced discussion about this instead of treating it like a 0 or 1 decision.
in short: our social sructure has built out so far beyond the Dunbar's number that it has become a haven for psychopaths.
People should get paid for reviewing. That’s the solution. Publishing companies rack in billions of dollars in pure profit exploring free labor. Once people actually get paid for reviewing, it becomes much faster and higher quality and you won’t need AI triage.
I think you are correct, its so hard to judge the quality of research that we have been using publication record/count as a proxy. Ultimately, having papers being easy to write is good, provided we find a better way of judging quality.
> We are very rapidly automating humans out of the academic publication loop
We are rendering it irrelevant. If this is the norm for academia, I’m sympathetic to the folks looking to cut its funding.
We just need journals run by an AI that charges other AI to read them, and then the AI run colleges can promote the AI with the most AI journal entries and citations.
I wonder into what weird research niche would all such AI schools converge to after enough time.
Paperclip research.
Benchmaxxing is already here. No need to pine for the future!
Unless someone has discovered some magic sauce for getting AI to write well, I have immense sympathy for anyone trying to wade through these papers.
I may generate slop from time to time, but I do my best to keep it to myself.
My 2 cents: AI has improved writing considerably for non-native english speakers (in particular China). Writing feels more standardized/boring but easier to read overall. I hit fewer papers that are a pain to read. The most problematic aspect I see are semi-bogus claims i.e. sentences that aren't false, but don't quite feel right either. YMMV.
What’s the point?
https://arxiv.org/stats/monthly_submissions
They should consider swapping this for a log plot.
I can imagine in 2027 academia looking like Moltbook.
This is a side effect of academia never taking open accessibility to papers and journals seriously.
If all these papers were not gatekept by journals, it would be trivially easy to validate at least the existence of cited papers and quotes.
Validating existence should already be trivial: virtually all journal already have publicly-available indices which contain at least author information and an abstract. Combine that with DOI citations and you're basically done - even with closed-access journals.
Checking the content is of course a lot more difficult, but that doesn't magically become trivial with open-access journals: you still need to read and interpret what is being said in the paper and compare it to the claims being made in the citation. Granted, these days you could use AI for a first pass, but it's still going to be incredibly tedious work.
> This is a side effect of academia never taking open accessibility to papers and journals seriously.
From my perspective, it's a clear manifestation of humanity's most pervasive failing - the one that defines every group, eventually.
And yet the automation we pour trillions in, is the one that will do anything, including everything I find interesting and will never clean my own house.
Truly passing the real Turing test.
It is trivially easy to validate the existence of a cited paper.
I'm sympathetic to the idea: we should have an open, publicly queryable citation graph. Google scholar could very easily offer this at marginal cost near zero, but they won't.
I feel like this should be treated as, and have consequences similar to, plagiarism.
Alas, it's probably just wishful thinking on my part.
I thought that it was. I thought journals were starting to implement full bans upwards of a year+ for those who don't honestly disclose AI usage in their work? If I'm not mistaken, arXiv is doing this as well? granted, disclosure is different from overuse, but it seems like a small jump to just go ahead and just ban not checking ones work! ...disclosed or not disclosed. my own personal view on it, is if you can't be bothered to spot hallucinations and other such errors, its not ai that is the issue, it is incompetence and laziness
So wrist slap?
> Both papers were accepted for oral presentations with the condition that they simply fix the hallucinated references.
I do wonder what truthfully could be on ai verification, if even one paper with such an error is accepted it sets the precedent you hopefully get lucky to not get caught (then again verifying for basic tells isn't the same verifying is this genuinely a worthwhile publication, but that's a separate matter)
This is why Eversaid.co is going to be even more useful as ai generated content explodes.
"A lot of the content of this blog was initially drafted by an agent of some sort"
What? I mean who does this. My voice is my voice and it's literally never occurred to me to have an LLM do a first draft. I thought that was college kid stuff.
I can accept having a human first draft and let LLM proofreading and/or do some polish on language, but not the reversed order.
This is an oversimplification and I'll edit. Caleb and I had a conversation about our reviewing woes and thought it would be fun to do an interview style post, so we had Claude come up with some questions based on our convo. We answered the questions from scratch and had Claude proofread at the end.