Without a thriving mathematical community to point out these things it would have stayed broken. With automated math that community as tao pointed out is at risk.
The bigger question is why there was internal pressure to rush such a historic launch without having someone in the company, anyone, check the proofs first.
This concerns the Hodge conjecture (millennium prize related) paper. Seems to me like PhD
nerds weren't confident bosses pushed ahead anyway.
What makes you think that nobody checked the proofs first? It's not like someone checking it once without spotting any mistakes means that nobody else will find any mistakes either.
Tweeter checked it with Astra. It seems like OAI could have pointed their own instance at it before launch. Because the source of the tip is likely someone at OAI, my guess is that they actually did check. But after the launch.
Because people are finding errors using other LLMs. This implies that if they spent a miniscule fraction of the enormous pile of money they spend making this pile of slop they'd find the errors. They didn't want to find errors. They want to build hype for an IPO.
My assumption is that they checked the proofs vigurously, but now a way broader community is taking a look with professionals from the relevant subfields, and different agent setups / models.
Proof of what? There is a sign error in one of the proofs, OpenAI acknowledged it and withdrew three papers (two relied on the result).
I agree LLM review is also fallible (as is human review) but the interesting part to me is that finding this sign error before publication should have been table stakes for OpenAI, it’s their own model that found the sign error.
I’m curious what was in the original prompt and what was in the prompt that led to finding the sign error, I think it matters a lot for understanding the dynamics here
> With automated math that community as tao pointed out is at risk.
If they can be automated, they are not necessary. If they are necessary, they won't be fully automated. It's a pretty simple experiment to run, the math "community" should bear with us. Darwin would be proud.
I find the paper about beating O(n log n) for integer multiplication also quite fishy, not sure but it seems like too good to be true, I feel like there must be a subtle flaw in that. Maybe that's just me hating these small numbers in the paper, but it seems wrong, unnatural even! I would be similarly skeptical about a physics paper that claims to be able to exceed the speed of light by a tiny fraction. There's no reason n log n is the natural limit here but I see a few good intuitions so having something else that can't be represented in an elegant form seem very "unmathematical" to me.
Surely it's a galactic algorithm that you can't physically run? You wouldn't get a constant as small as 2^{-182} without some other numbers elsewhere being incredibly large.
From the "Introduction" section of that paper: "The constants and thresholds in the
construction are extremely large".
If you do a dump like this all of it should be formalized, there's simply too much material to review by hand and additionally it is AI-written which makes it hard to read compared to human work.
Not a mathematician but surely if a problem I was working on had an AI also working on it, I would want to know as early as possible - even with flaws or gaps. What advantage is it to me to be less informed?
I can prompt ChatGPT right now and ask for mountains of more "mathematical work"; thousands and thousands of pages of nonsense for you to review. So you can "be informed".
Apparently the write-ups are garbage (as in very hard to read). I feel like they could've had AI fix that up at least somewhat. Maybe they'll reinvest more in writing ability now.
I think the obscurity is a feature, not a bug. They don't want a headline where 250 of these results are invalidated overnight. They want rejections to trickle out and be buried.
flooding the system with parts that may be incorrect hurts the whole process and will get people to tune out (like politics). Can't see the International Mathematical Union making a similar mistake because it would do reputational harm. But the models don't care about their reputation. Bad for the layman - like me - to know what to make of all this.
I’m conflicted. I guess we’ll see what the final total is once an enormous level of unpaid human effort is expended verifying the AI outputs. A little sad if that’s the future of math.
It kind of reminds me of when tech giants open source a project as a means of putting a positive spin on abandonware. “Here’s the source! Any problems are yours to fix now. You’re welcome”
Where are you getting unpaid from? Almost all people who are qualified to analyze the results are paid researchers. And if it is unpaid, then it sounds like they're looking it over for their own reasons, and that's fine?
I also fail to see the issue you have with releasing abandoned source. In what world is that bad? That obviously is a gift and should be encouraged. e.g. id software's history of doing that has meant their work stays alive forever.
Researchers normally don't pay each other to read each other's work. It's a symbiotic relationship. If they think the AI results are nonsense, they could ignore it like any other crank. If they think it seems plausible and it's relevant to them, they can try to understand it. Seems fine?
yes, but in this case, openai does a lot of PR how their models solve important math problems. if it goes unchallenged, parts of society would think that is true. what would happen if we scale it and 10k companies dump 10k papers every month claiming solved math problems. how is this scalable?
We need the companies to humanly review their papers. in the same way as at other companies we use humans to review the papers.
Well, per another comment in the thread, some 20% of their solutions come with formalization, so there's a very high chance (probably higher than typical asks of research mathematics) that they did solve the problem. And that also presents a pretty easy solution to the scaling issue: demand formal proofs.
(If you're going to object that it's difficult to validate the statement of the problem, please first state your level of experience doing so. It's getting tiring seeing people raise this objection and claim that a statement is just as hard as a proof over and over who don't seem to actually know any math and have never tried to write anything in Lean)
Why's that? Do you think that institutions and funding agencies won't cover researchers' use of advanced models, or what? The group I worked in in undergrad had millions of dollars of equipment for doing experiments. I'd have to imagine they could get funding for a few thousand dollars in tokens for the theoreticians to have AI assistance.
> Almost all people who are qualified to analyze the results are paid researchers.
This _might_ have been true somewhat in the past (although it wasn't), but it's completely false today. Anyone with access to a sufficiently advanced model has the capabilities of analyzing these papers/proofs. It's no different than reading a codebase you might not be fully familiar with, and checking it for correctness (give an engineering analogy).
This hardcore gatekeeping of math (and by extension STEM) fields MUST stop.
I mean, I have a decent math background, but I would struggle greatly to attempt to even tell you what most (or any) of the conjectures in e.g. number theory are about at even the highest level. I can't imagine a layman would have any hope.
Like I was reading some about adele rings last night, which is already going to be quite a concept for a layman to be able to even slightly describe. Then you can layer on that apparently they're locally compact, so we can talk about harmonic analysis on the additive group. Like, come on now, 99.99% of people have no hope of ever following along.
I’m unsure. The whole academia is built on unpaid human efforts. Journal writers are unpaid, and institutions paid for their papers to be published by for-profit publishers. Journal reviewers are paid the bare minimum, certainly unproportional to their efforts and expertise.
The product at the end is important, but so is the process. Few of the things that would happen along the way are happening here, so it's harder to justify the value of the deliverable when there is a failure.
well, it might be that these proofs are correct or it might be that people aren't bothering to spend a lot of time checking whether they are correct. OpenAI already has a pretty bad reputation in the mathematics community for how they are approaching this process, they seem to be more interested in creating a story for their IPO than advancing math.
But can the others even be "disproven", given that they apparently are so messy and awful that no humans can follow them? Shouldn't the onus instead be on OpenAI to prove that they're right, instead of hundreds of mathematicians wading through slop?
Exactly, I was glad to see these withdrawals, its a natural part of a healthy ecosystem of scientific review, hypothesis, claim, test, refute, extend, withdraw, its the heart of science.
IMO If you take out all the stupid human aspects mostly related to fear, egos, etc, we should brace the imperfect and helpful tools, whatever they are, improve them so they are as easy as possible to review, and keep that core scientific discovery loop going
Retraction and withdrawal are different. Retraction is when you publish something, it passes peer review, is published, and some time later its publication is undone, often by an editor or some other person because some fraud was uncovered.
Withdrawal is akin to submitting a paper to peer review and then when you’ve noticed mistakes, you decide to take the paper back and correct it.
Reject is when someone else notices the mistakes and tells you to take it back and correct it.
Withdrawal and reject happen all the time in a scientist’s career. They don’t necessarily mean the scientist is doing bad research, just the research was not ready. Retract usually means something more.
By dumping the papers, OpenAI skipped the typical peer review process, so peer review should be understood as what’s going on now as mathematicians look over the papers and find flaws.
I guess I feel two ways about this. It does could be a "working with the garage open" sort of thing where these are known to be unverified and they're just letting people see the sausage being made.
But that also means marketing and PR should shut the fuck up until things are verified.
Picture a remake of Good Will Hunting, where Will is an AI and, instead of getting the mathematical formulas correct, he just mass dumps a bunch of nonsense and the professors have to go through it all, pointing out where it is wrong. The professors know that AI Will isn’t as smart as people say, but their funding depends on it, and if they prove it, which they can do easily, it may just tank the whole economy, causing a depression, all because they have history’s dumbest President in office.
Those results were not dumped to advance math. Those results were dumped to generate positive press for OpenAI.
Now it's on actual mathematicians to figure out whether the proofs are bogus or not. But the mistakes will never reach the same level of public attention as the original positive press, so for OpenAI this is good anyway, consequences be damned.
In a sense this is a microcosm of AI usage in the wild, ignorance is laundered through LLMs, and it is left for those that still have knowledge to figure out what makes sense.
OpenAI is a horribly negligent company. The threat it poses to humanity is not that their models will be a superintelligent singularity that will take over the world, is that their negligence and greed has real world consequences that they really don't give a fuck about. Like right now, their shitty math bruteforce is just keeping actual specialists occupied trying to figure out what is bullshit from what is not. For free, mind you.
Tao and most anti/critical ai math folks remind me a bit of the brhamins in the hindu caste hirarchy system... whilst others fought (kshatriyas), farmed/traded (vaishyas), built things, cleaned roads etc (shudras), the brahmins were the high priest doing science, stronomy, religion ...
AI is this strange modernist machinary that kind of threatens that brhaminic role... its almost like the vatican vs post industrialization world .. where they still have to keep making the case for why religion/priesthood/god is important... even as the tech/science world starts operating on totally different terms...
This metaphor might be applicable if the AI slop machine was in fact producing novel output. It seems to be getting invalidated as people dig through the wall of meaningless text surrounding the actual results.
thats kinda the point tho... im not making a claim about whether ai produces novel output, im talking about the social role of the people arguing about what counts as knowledge ,, novelty in the first place... Tao et all ,,
This is pretty standard. Important to note that these "errors" (* not really errors) themselves were caught by an LLM, further proving their usefulness.
* The reason you shouldn't consider the withdrawals to be caused by errors is because this is pretty standard in math and development. "Errors" like this are a core aspect of science and it happens _all the time_. And from my research LLMs have a far lower error rate than even the best human scientists.
It's so funny to see people on Hacker News talk about academic communities that they clearly are not part of. To be clear to anyone reading, no, retracting papers is _not_ standard in math. In math, people submit papers after they have checked with colleagues, they give seminars, they submit it for peer review, etc. It is not shotgunning papers and retracting incorrect ones on the regular. Retracting papers is rare and embarrassing.
I wonder why the heck were they provided / uploaded then. Perhaps just a fast and loose play-out on their part. What I don't understand is how come engineers / scientists working on these are okay with this kind of attitude.
This company is just irresponsible. We (at least, Americans that vote and can therefore decide indirectly what’s legal) should not allow them to continue.
Irresponsible is an adjective that means lacking a proper sense of responsibility, or acting without thinking about or caring about the potential consequences of one's actions.
They do not care if they waste everyone’s time or flood the common with slop. They did not spend their own time to verify that the Lean proofs correspond to the natural language proofs. They did not even spend their own time to verify that all of these proofs are well written.
They instead are mining the unrenewable resource of open math problems.
However, seeing it another way is easy if you are financially motivated by their upcoming IPO (see how easy it is to invent motivations for comments?).
it's like if you vibe coded something and the onus is now on the reviewer and the reviewer tells you your work contains bugs and is messy - that is not acceptable from the reviewers pov - why should the reviewer spend all his human effort, a scarce resource, reviewing your code while you've spent barely a fraction of his effort generating this. OpenAI is a trillion dollar company, surely they can verify stuff before pushing it out? The problem is not that they are solving the problems, they don't care at all about the actual process of doing mathematics. Using compute to mine problems and throwing results in github and letting human reviewers spend effort to correct these is not going to win them any favor. If OpenAI really cared about math, they would have someone on their side who actually understood the results they produced and was able to verify their correctness and educate others.
Withdrawing it isn't irresponsible. I don't even think releasing it is, because that is often the best way to find problems with it; expose it to the wider research community. Doing PR victory laps based on unverified, unfinished work, on the other hand, is. It is well known, even scientifically demonstrated, that retractions or refutations get much less attention than the initial exuberant PR announcement. And thus this practice contributes to misinformation.
And just to preempt the kneejerk whataboutism: Yes, all of academia does this to varying extents, and the mainstream media are also complicit. And that is also irresponsible. And no, that is not an excuse for OpenAI. Especially when you consider that OpenAI actively portray themselves as some kind of moral arbiter on AI and "doing good for humanity". They should be held to the extraordinarily high moral standards they purport to hold themselves to.
The greatest proof ever written by man, Andrew Wiles' FLT, was published with a flaw before being retracted and reworked. It's _completely_ normal in mathematics to find flaws in the argument. That's what peer review in mathematics is _for_.
Besides which it's not clear if these papers are considered published or preprint since they appear in no journal, so it's not really a retraction.
It's also very very normal to post preprints on Arxiv before peer review, so it's not the case that mathematics is kept private during review.
Without a thriving mathematical community to point out these things it would have stayed broken. With automated math that community as tao pointed out is at risk.
Have you got a link to someone pointing it out? It looks like they're still going through formal proofs so likely found the problems that way.
Yes: https://x.com/ElliotGlazer/status/2108026240582246600
So someone ran a different LLM to find an issue they'd find anyway during formalisation? That's not the same as relying on thriving community.
The bigger question is why there was internal pressure to rush such a historic launch without having someone in the company, anyone, check the proofs first.
This concerns the Hodge conjecture (millennium prize related) paper. Seems to me like PhD nerds weren't confident bosses pushed ahead anyway.
What makes you think that nobody checked the proofs first? It's not like someone checking it once without spotting any mistakes means that nobody else will find any mistakes either.
Tweeter checked it with Astra. It seems like OAI could have pointed their own instance at it before launch. Because the source of the tip is likely someone at OAI, my guess is that they actually did check. But after the launch.
Because people are finding errors using other LLMs. This implies that if they spent a miniscule fraction of the enormous pile of money they spend making this pile of slop they'd find the errors. They didn't want to find errors. They want to build hype for an IPO.
My assumption is that they checked the proofs vigurously, but now a way broader community is taking a look with professionals from the relevant subfields, and different agent setups / models.
No. Someone received a tip, presumably from inside OAI. He then used OAI Astra to check..
That’s not proof though is it? If the original LLM output is fallible surely the LLM review of that output is also very much fallible?
Proof of what? There is a sign error in one of the proofs, OpenAI acknowledged it and withdrew three papers (two relied on the result).
I agree LLM review is also fallible (as is human review) but the interesting part to me is that finding this sign error before publication should have been table stakes for OpenAI, it’s their own model that found the sign error.
I’m curious what was in the original prompt and what was in the prompt that led to finding the sign error, I think it matters a lot for understanding the dynamics here
Output of LLM can be infallible* even if LLMs themselves make mistakes. Ditto with humans.
As much as anything can be infallible.
> With automated math that community as tao pointed out is at risk.
If they can be automated, they are not necessary. If they are necessary, they won't be fully automated. It's a pretty simple experiment to run, the math "community" should bear with us. Darwin would be proud.
This statement assumes local maxima don’t exist and greedy short term optimization always leads to long term benefit.
“If they can be automated, they are not necessary.” That’s a pretty interesting take as eventually everything could be automated.
You've abandoned every part of yourself to the siren's song of efficiency and automation, huh?
To witness an arson and rejoice reveals an ugly kind of sadism.
Let the evolution run its course. Let the stream find its way. Let the forces restore the equilibrium.
I find the paper about beating O(n log n) for integer multiplication also quite fishy, not sure but it seems like too good to be true, I feel like there must be a subtle flaw in that. Maybe that's just me hating these small numbers in the paper, but it seems wrong, unnatural even! I would be similarly skeptical about a physics paper that claims to be able to exceed the speed of light by a tiny fraction. There's no reason n log n is the natural limit here but I see a few good intuitions so having something else that can't be represented in an elegant form seem very "unmathematical" to me.
I didn’t read the paper but can’t this just be additionally with doing the actual muls? Or was it a nonconstructive proof?
Surely it's a galactic algorithm that you can't physically run? You wouldn't get a constant as small as 2^{-182} without some other numbers elsewhere being incredibly large.
From the "Introduction" section of that paper: "The constants and thresholds in the construction are extremely large".
If you do a dump like this all of it should be formalized, there's simply too much material to review by hand and additionally it is AI-written which makes it hard to read compared to human work.
Not a mathematician but surely if a problem I was working on had an AI also working on it, I would want to know as early as possible - even with flaws or gaps. What advantage is it to me to be less informed?
I can prompt ChatGPT right now and ask for mountains of more "mathematical work"; thousands and thousands of pages of nonsense for you to review. So you can "be informed".
Apparently the write-ups are garbage (as in very hard to read). I feel like they could've had AI fix that up at least somewhat. Maybe they'll reinvest more in writing ability now.
I think the obscurity is a feature, not a bug. They don't want a headline where 250 of these results are invalidated overnight. They want rejections to trickle out and be buried.
Not sure why you're being downvoted, because this is exactly what I think is happening.
Almost like they care less about the maths and more about a deadline for marketing purposes.
flooding the system with parts that may be incorrect hurts the whole process and will get people to tune out (like politics). Can't see the International Mathematical Union making a similar mistake because it would do reputational harm. But the models don't care about their reputation. Bad for the layman - like me - to know what to make of all this.
And added 6 new ones. Might want to add that to the headline.
Looks like those are new formalizations of existing proofs, not new ones
3 mistakes (so far) out of ~400 is still a pretty good hit rate
I’m conflicted. I guess we’ll see what the final total is once an enormous level of unpaid human effort is expended verifying the AI outputs. A little sad if that’s the future of math.
It kind of reminds me of when tech giants open source a project as a means of putting a positive spin on abandonware. “Here’s the source! Any problems are yours to fix now. You’re welcome”
Where are you getting unpaid from? Almost all people who are qualified to analyze the results are paid researchers. And if it is unpaid, then it sounds like they're looking it over for their own reasons, and that's fine?
I also fail to see the issue you have with releasing abandoned source. In what world is that bad? That obviously is a gift and should be encouraged. e.g. id software's history of doing that has meant their work stays alive forever.
Unpaid by OpenAI. In service of their PR goals they’re creating a bunch of work for others and not contributing to compensating that work.
OpenAI isn’t paying mathematicians to verify all these papers. It’s dumping the papers trying to nerd snipe them into checking it for free.
Researchers normally don't pay each other to read each other's work. It's a symbiotic relationship. If they think the AI results are nonsense, they could ignore it like any other crank. If they think it seems plausible and it's relevant to them, they can try to understand it. Seems fine?
yes, but in this case, openai does a lot of PR how their models solve important math problems. if it goes unchallenged, parts of society would think that is true. what would happen if we scale it and 10k companies dump 10k papers every month claiming solved math problems. how is this scalable?
We need the companies to humanly review their papers. in the same way as at other companies we use humans to review the papers.
Well, per another comment in the thread, some 20% of their solutions come with formalization, so there's a very high chance (probably higher than typical asks of research mathematics) that they did solve the problem. And that also presents a pretty easy solution to the scaling issue: demand formal proofs.
(If you're going to object that it's difficult to validate the statement of the problem, please first state your level of experience doing so. It's getting tiring seeing people raise this objection and claim that a statement is just as hard as a proof over and over who don't seem to actually know any math and have never tried to write anything in Lean)
Formalization only helps you if you prove that the formalization is correctly implemented, that the model didn’t subtly mess up or cheat.
Someone still has to read the formalization.
> It's a symbiotic relationship
That's right, and the difference is that this one is parasitic.
Why's that? Do you think that institutions and funding agencies won't cover researchers' use of advanced models, or what? The group I worked in in undergrad had millions of dollars of equipment for doing experiments. I'd have to imagine they could get funding for a few thousand dollars in tokens for the theoreticians to have AI assistance.
> Almost all people who are qualified to analyze the results are paid researchers.
This _might_ have been true somewhat in the past (although it wasn't), but it's completely false today. Anyone with access to a sufficiently advanced model has the capabilities of analyzing these papers/proofs. It's no different than reading a codebase you might not be fully familiar with, and checking it for correctness (give an engineering analogy).
This hardcore gatekeeping of math (and by extension STEM) fields MUST stop.
I mean, I have a decent math background, but I would struggle greatly to attempt to even tell you what most (or any) of the conjectures in e.g. number theory are about at even the highest level. I can't imagine a layman would have any hope.
Like I was reading some about adele rings last night, which is already going to be quite a concept for a layman to be able to even slightly describe. Then you can layer on that apparently they're locally compact, so we can talk about harmonic analysis on the additive group. Like, come on now, 99.99% of people have no hope of ever following along.
I’m unsure. The whole academia is built on unpaid human efforts. Journal writers are unpaid, and institutions paid for their papers to be published by for-profit publishers. Journal reviewers are paid the bare minimum, certainly unproportional to their efforts and expertise.
We are all reverse centaurs now.
I would say not. For a mathematicians, having to retract more than 2 papers in a lifetime is already a big issue in their career.
But have you formalized every single one of your results? And if you haven't what odds would you put on one of them not working out, if formalized?
But is it equivalent to withdrawing post-publication or is it more akin to not passing peer review with major revisions requested?
Most mathematicians don't produce 400 papers in a 48 hour window either so I'm not sure comparisons are helpful
The product at the end is important, but so is the process. Few of the things that would happen along the way are happening here, so it's harder to justify the value of the deliverable when there is a failure.
I bet it's a lower rate of errors than the typical human math paper of 2025
well, it might be that these proofs are correct or it might be that people aren't bothering to spend a lot of time checking whether they are correct. OpenAI already has a pretty bad reputation in the mathematics community for how they are approaching this process, they seem to be more interested in creating a story for their IPO than advancing math.
But can the others even be "disproven", given that they apparently are so messy and awful that no humans can follow them? Shouldn't the onus instead be on OpenAI to prove that they're right, instead of hundreds of mathematicians wading through slop?
The onus is on formal verification when it comes to computer generated results. So far only 22% of the papers have it.
Exactly, I was glad to see these withdrawals, its a natural part of a healthy ecosystem of scientific review, hypothesis, claim, test, refute, extend, withdraw, its the heart of science.
IMO If you take out all the stupid human aspects mostly related to fear, egos, etc, we should brace the imperfect and helpful tools, whatever they are, improve them so they are as easy as possible to review, and keep that core scientific discovery loop going
How many mathematicians need to retract ~1% of their papers?
Retraction and withdrawal are different. Retraction is when you publish something, it passes peer review, is published, and some time later its publication is undone, often by an editor or some other person because some fraud was uncovered.
Withdrawal is akin to submitting a paper to peer review and then when you’ve noticed mistakes, you decide to take the paper back and correct it.
Reject is when someone else notices the mistakes and tells you to take it back and correct it.
Withdrawal and reject happen all the time in a scientist’s career. They don’t necessarily mean the scientist is doing bad research, just the research was not ready. Retract usually means something more.
By dumping the papers, OpenAI skipped the typical peer review process, so peer review should be understood as what’s going on now as mathematicians look over the papers and find flaws.
It would turn out to be a pure energy and time wasting exercise … the math community would want to keep away from it.
Proof by authority works until human mathematicians actually run the code. Back to prompt engineering.
Whose names are on these papers?
They try to be clever and put company name as the author.
Not a great sign when nobody wants to attach their own name to words thrown into the world.
I guess I feel two ways about this. It does could be a "working with the garage open" sort of thing where these are known to be unverified and they're just letting people see the sausage being made.
But that also means marketing and PR should shut the fuck up until things are verified.
Picture a remake of Good Will Hunting, where Will is an AI and, instead of getting the mathematical formulas correct, he just mass dumps a bunch of nonsense and the professors have to go through it all, pointing out where it is wrong. The professors know that AI Will isn’t as smart as people say, but their funding depends on it, and if they prove it, which they can do easily, it may just tank the whole economy, causing a depression, all because they have history’s dumbest President in office.
SMBC had a funny take on the Good Will Hunting AI joke, recently.
https://www.smbc-comics.com/comic/equations
All that AI compute and resources and they couldn't hold onto releasing the papers until a legitimate peer review was performed?
Just had to get that PR stunt out to bump their valuation.
Vibe coding math research is just next-level AI slop.
Mind you, this is a competent AI company that's making these mistakes.
I can only imagine what non-technical people are putting out in production via vibe-coded AI slop.
I think it's equal parts PR stunt and negging.
People won't want to seriously peer review an AI study unless it's already out there potentially spreading misinformation
These are not mistakes.
Those results were not dumped to advance math. Those results were dumped to generate positive press for OpenAI.
Now it's on actual mathematicians to figure out whether the proofs are bogus or not. But the mistakes will never reach the same level of public attention as the original positive press, so for OpenAI this is good anyway, consequences be damned.
In a sense this is a microcosm of AI usage in the wild, ignorance is laundered through LLMs, and it is left for those that still have knowledge to figure out what makes sense.
OpenAI is a horribly negligent company. The threat it poses to humanity is not that their models will be a superintelligent singularity that will take over the world, is that their negligence and greed has real world consequences that they really don't give a fuck about. Like right now, their shitty math bruteforce is just keeping actual specialists occupied trying to figure out what is bullshit from what is not. For free, mind you.
Imagine how many wasted hours in peer review these AI math papers will cause for the 1% chance of reaching a transformative idea.
Another discussion: https://news.ycombinator.com/item?id=50002650
Tao and most anti/critical ai math folks remind me a bit of the brhamins in the hindu caste hirarchy system... whilst others fought (kshatriyas), farmed/traded (vaishyas), built things, cleaned roads etc (shudras), the brahmins were the high priest doing science, stronomy, religion ...
AI is this strange modernist machinary that kind of threatens that brhaminic role... its almost like the vatican vs post industrialization world .. where they still have to keep making the case for why religion/priesthood/god is important... even as the tech/science world starts operating on totally different terms...
This is weirdo pseudo-neo-religious slop.
This metaphor might be applicable if the AI slop machine was in fact producing novel output. It seems to be getting invalidated as people dig through the wall of meaningless text surrounding the actual results.
thats kinda the point tho... im not making a claim about whether ai produces novel output, im talking about the social role of the people arguing about what counts as knowledge ,, novelty in the first place... Tao et all ,,
"a sign error" LOL
This is pretty standard. Important to note that these "errors" (* not really errors) themselves were caught by an LLM, further proving their usefulness.
* The reason you shouldn't consider the withdrawals to be caused by errors is because this is pretty standard in math and development. "Errors" like this are a core aspect of science and it happens _all the time_. And from my research LLMs have a far lower error rate than even the best human scientists.
It's so funny to see people on Hacker News talk about academic communities that they clearly are not part of. To be clear to anyone reading, no, retracting papers is _not_ standard in math. In math, people submit papers after they have checked with colleagues, they give seminars, they submit it for peer review, etc. It is not shotgunning papers and retracting incorrect ones on the regular. Retracting papers is rare and embarrassing.
Would you care to share your research? What was your methodology? Did you write it up? Is it published?
Or is it the kind of research you'd prefer to keep shrouded in mystery?
So much for the "it's lean verified" defense.
The withdrawn papers were not lean verified nor claimed to be.
I wonder why the heck were they provided / uploaded then. Perhaps just a fast and loose play-out on their part. What I don't understand is how come engineers / scientists working on these are okay with this kind of attitude.
I suspect they are not okay with this way of working.
I’m sure the outsized pay packages help quite a bit
They’re okay enough to do the work and collect a paycheck and stock options. I don’t think their arms are being twisted that hard.
How come mathematicians are ok with human mathematicians are ok with that kind of sketchy submissions? That's pretty bad too.
... they're not? Why would you think mathematicians are happy about sketchy or bad submissions.
Lies travel around the world before the truth has time to put it’s sneaker’s on…
The headlines keep the hype train arunnin
But that's the argument that was used when people here were skeptical about the results.
not these results
Can you point to where that defense has been made?
This company is just irresponsible. We (at least, Americans that vote and can therefore decide indirectly what’s legal) should not allow them to continue.
This type of sentiment is almost always motivated by fear over the potential negative personal economic impact from AI (e.g. losing employment).
Publishing a math paper and then unpublishing it is not "irresponsible". It's just a math paper.
Irresponsible is an adjective that means lacking a proper sense of responsibility, or acting without thinking about or caring about the potential consequences of one's actions.
They do not care if they waste everyone’s time or flood the common with slop. They did not spend their own time to verify that the Lean proofs correspond to the natural language proofs. They did not even spend their own time to verify that all of these proofs are well written.
They instead are mining the unrenewable resource of open math problems.
However, seeing it another way is easy if you are financially motivated by their upcoming IPO (see how easy it is to invent motivations for comments?).
It is irresponsible when nobody checked it for correctness first. They just copy pasted what the slop machine spit out.
How is withdrawing a paper irresponsible?
it's like if you vibe coded something and the onus is now on the reviewer and the reviewer tells you your work contains bugs and is messy - that is not acceptable from the reviewers pov - why should the reviewer spend all his human effort, a scarce resource, reviewing your code while you've spent barely a fraction of his effort generating this. OpenAI is a trillion dollar company, surely they can verify stuff before pushing it out? The problem is not that they are solving the problems, they don't care at all about the actual process of doing mathematics. Using compute to mine problems and throwing results in github and letting human reviewers spend effort to correct these is not going to win them any favor. If OpenAI really cared about math, they would have someone on their side who actually understood the results they produced and was able to verify their correctness and educate others.
Withdrawing it isn't irresponsible. I don't even think releasing it is, because that is often the best way to find problems with it; expose it to the wider research community. Doing PR victory laps based on unverified, unfinished work, on the other hand, is. It is well known, even scientifically demonstrated, that retractions or refutations get much less attention than the initial exuberant PR announcement. And thus this practice contributes to misinformation.
And just to preempt the kneejerk whataboutism: Yes, all of academia does this to varying extents, and the mainstream media are also complicit. And that is also irresponsible. And no, that is not an excuse for OpenAI. Especially when you consider that OpenAI actively portray themselves as some kind of moral arbiter on AI and "doing good for humanity". They should be held to the extraordinarily high moral standards they purport to hold themselves to.
They only withdrew it after someone pointed out that it was a flawed paper.
In one case by asking Astra to review it.
Not exactly encouraging that they did their homework before publishing results.
This is how science is meant to work.
[delayed]
Most good scientists work hard to prove results to themselves before going public.
Peer review is then done _in private_ before publication as a check on quality and significance.
Retracting a paper is pretty embarrassing. And not considered science as usual.
The greatest proof ever written by man, Andrew Wiles' FLT, was published with a flaw before being retracted and reworked. It's _completely_ normal in mathematics to find flaws in the argument. That's what peer review in mathematics is _for_.
Besides which it's not clear if these papers are considered published or preprint since they appear in no journal, so it's not really a retraction.
It's also very very normal to post preprints on Arxiv before peer review, so it's not the case that mathematics is kept private during review.
Retraction is after publication and peer review, no?
Yes. A researcher who achieves a reputation of making mistakes carries that reputation.
No, scientists check their work before releasing it. Slop is still slop if you wrap it up as an academic result.
Withdrawing it, in itself, might not be, but it kind of skips over the part where they have a paper that warrants it in the first place.