OpenAI withdraws three mathematical results
twitter.com/danintheory
[61 comments hidden]
[8 comments hidden]
[7 comments hidden]
[4 comments hidden]
[3 comments hidden]
[hidden]
I'm unaware of any serious proofs that have been shown to have a kernel exploit in them.
[hidden]
I agree with Jtarii that it's very unlikely a Lean bug is critical to most of these proofs. But we're in strange times, so I agree wtih the sentiment that we should wait for further analysis before declaring complete confidence in the proofs.
[8 comments hidden]
2. Why shouldn’t math progress happen in the open, commit by commit? Why is it so horrible if a proof is 95% of the way there but we later find that it needs to be refined? Mathematics previously was optimizing for an antiquated publishing and distribution scheme. There is no need for the first print to be correct. We have the internet now. We can and should publish incomplete results and correct things on the fly. Maybe mathematicians would have solved some of these problems years ago if they didn’t hide incomplete almost solutions in their filing cabinet because it wasn’t yet ready to be published.
You don’t hate the pageantry of mathematics and academics enough.
[3 comments hidden]
I bet you enjoy when a peer asks you to find the issues in a fully AI generated PR that’s 95% of the way there.
[hidden]
There's a lot to hate about the academic world, but the solution isn't spewing out terabytes of crappy half-baked results.
[3 comments hidden]
Also “real mathematicians” aren’t the people who “math belongs to”, you’re a mathematician if you do math, that’s it.
[2 comments hidden]
[hidden]
1. In OpenAI's case, they dumped 1.8 GiB of Lean proofs on the world. I don't think they've done anything ethically wrong by doing that, but it's the exact opposite of "commit by commit". In fact, I'd say human mathematics has been much closer to "commit by commit", usually using smaller results as stepping stones.
2. You can have "commit by commit" without formal provers like Lean. Just keep your text in a Git repo.
[15 comments hidden]
[3 comments hidden]
[hidden]
[hidden]
[hidden]
> All the human needs to verify is that the statement of the theorem is translated correctly from natural language to Lean. That usually covers a very small surface of the Lean code.
That's exactly what the navier stokes paper posted yesterday pointed out where the LLM bends the Lean code to make it "compile", because the NL might be wrong to begin with or because it missed a detail:https://arxiv.org/html/2610.08144v1#S2
For complex / tedious proofs I can easily see how small details like this can lead to a valid lean proof (or valid "code"), but missing the important details that got lost.
[8 comments hidden]
Take the Rieman hypothesis. All one would have to do to prove or disprove it would be to encode the statement of the hypothesis in Lean, press enter, and we're off to the races.
That's not how it works. Essentially you have to encode all the intermediary steps of the proof in Lean too, and then Lean can check their correctness for you and check that they lead to each other. But it won't just generate a whole proof from nothing. That is the whole point of the Gen AI math claims.
[4 comments hidden]
[3 comments hidden]
>> All the human needs to verify is that the statement of the theorem is translated correctly from natural language to Lean. That usually covers a very small surface of the Lean code.
As far as I understand the comment "the statement of the theorem" is the declaration of the theorem to be proved, not its proof. That it "usually covers a very small surface of the Lean code" also implies that the OP was only referring to the theorem, the thing to be proved, and not the proof that can run into many thousands of lines.
If I misunderstood then I don't think that's a problem? I don't believe my comment above comes across as rude or an attack on the OP? I think it's normal for this site for users to correct one another without it being a cause for bad feelings.
[2 comments hidden]
I hope there's no hard feelings, by the way -- I wasn't meaning to be nasty to you in that comment. I guess I thought your reading was a bit uncharitable, but I understand that misunderstandings happen (very much including on my end) and I didn't mean to make a big thing of it, so I'm sorry if I came across rudely.
[6 comments hidden]
The proof explicitly hand-waves some complexity by assuming lookup tables to avoid some calculations which isn’t actually possible since it’s dealing with such large numbers and it only works on incredibly large numbers.
The complexity being so close to nlogn and the handwaving by assuming lookup tables in parts should be a really really obvious smell. At the very least worthy of holding back from the broader announcement.
It us proven in lean as-is with these assumptions and it’s not one of the ones retracted but those assumptions are doing some heavy lifting. I think it’s worth adding back in those ‘by using a lookup tables for x’ complexities and seeing if we really are below nlogn on that one.
[hidden]
This is just a Rice Theorem problem, right?
[22 comments hidden]
[16 comments hidden]
I'm probably wrong, but what's the point of throwing away all skepticism?
[5 comments hidden]
[hidden]
[hidden]
The should be on AI labs to definitively prove their extraordinary claims, and they should be paying mathematicians to do so given that the results from these machines are so opaque and often nonsensical.
[hidden]
[hidden]
[9 comments hidden]
Nothing wrong with being skeptical, but I see no reason to be skeptical as of yet.
[hidden]
[4 comments hidden]
[2 comments hidden]
[hidden]
[hidden]
[hidden]
[hidden]
Good question whether Lean is enough.
[5 comments hidden]
https://leodemoura.github.io/blog/2026-8-1-postmortem-for-ke...
There's also a fairly well known incident where the Lean formalization of the Riemann Hypothesis in Mathlib was incorrect.
[4 comments hidden]
[2 comments hidden]
[hidden]
[3 comments hidden]
I'm curious to know if the withdrawal was due to an actual mathematician looking at the papers and noticing the errors, or they ran a model on these to proofread, which would not be the first time, presumably, since they would have surely done that before publishing. Both options have interesting implications.
[19 comments hidden]
[7 comments hidden]
[6 comments hidden]
From the "Introduction" section of that paper: "The constants and thresholds in the construction are extremely large".
(And verifying if the algorithm multiplies correctly or not is the less-interesting part of this, anyway. Gets you no closer to verifying the complexity result).
[2 comments hidden]
[2 comments hidden]
Of course, that's not to say the research is necessarily useless. It's still theoretically interesting to find "better" algorithms if only to shed some light on lower bounds, and so on. And who knows, maybe the line of research could lead to more practical algorithms later on.
[hidden]
https://eprint.iacr.org/2022/439
this doesn't use the literal n \log n algorithm, which may be galactic. but it uses fundamentally similar techniques.
[hidden]
[hidden]
Regarding elegance, take a look at Graham's number. It was not some meaningful constant - it's just a big-ass number which could be used in existence proof. Human mathematicians have been using this approach for quite some time, it's not really AI doing things odd
[6 comments hidden]
The lean proof uses these assume ‘a lookup table’ assumptions. The paper smells with the nlogn^0.99999999 (many more nines actually) and unbelievably close to nlogn statement and then the literal talk of lookup tables pushes it over the edge clearly for me.
Maths can generate weird numbers out of nowhere but it really really looks like an nlogn result with some tricks to get past leen to me
[3 comments hidden]
[2 comments hidden]
[2 comments hidden]
The proof can be entirely valid even if it's not actually reasonable to implement and requires an enormous size lookup table - but it still is a meaningful mathematical result and makes progress.
I'm sure there are plenty of times where originally something was proven and thought to be completely impractical but then later had niche use cases or was the bedrock for solving other cases. And the opposite is true: there remain plenty of proofs of things that are mathematically certain but will in all practicality never be useful.
[3 comments hidden]
Edited.Added. To compute the posterior one needs to know the probability of detecting such error in a proposed proof.
[2 comments hidden]
If it doesn't have a Lean proof, why would we bother with it? And they really shouldn't publish it! We don't need math slop too.
[hidden]
[22 comments hidden]
[5 comments hidden]
[2 comments hidden]
[10 comments hidden]
[8 comments hidden]
[4 comments hidden]
With 40% formalized they probably have a good idea of how many were found to have fatal issues in the formalization attempt, and they hired some mathematicians to verify some of them, especially the big headline ones.
[hidden]
[hidden]
[3 comments hidden]
[hidden]
[hidden]
[3 comments hidden]
Are they all too busy having brilliant ideas? Doubt.
OpenAI math paper dump should be considered like a hint from 200 IQ eccentric genius - unreliable but perhaps insightful. If it was not "ugh AI" people would be happy about it.
[62 comments hidden]
[2 comments hidden]
[5 comments hidden]
IMO If you take out all the stupid human aspects mostly related to fear, egos, etc, we should brace the imperfect and helpful tools, whatever they are, improve them so they are as easy as possible to review, and keep that core scientific discovery loop going
[4 comments hidden]
Exactly, and leave the only veneberable human aspect which is allowing the rotten rich bloat further without dissent!
[3 comments hidden]
In this case we are talking about using AI to accelerate human progress in mathematics and scientific discovery, and rather than stay on topic, a human comes in with a wealth inequality complaint.
The distributions of the gains of AI advancement is a separate issue, we definitely shouldn't hold up progress on the frontier of human knowledge because of wealth inequality. Its an important issue that needs to be solved, but pausing or slowing progress at the edge of human discovery because of wealth inequality concerns is utterly insane.
[28 comments hidden]
It kind of reminds me of when tech giants open source a project as a means of putting a positive spin on abandonware. “Here’s the source! Any problems are yours to fix now. You’re welcome”
[24 comments hidden]
I also fail to see the issue you have with releasing abandoned source. In what world is that bad? That obviously is a gift and should be encouraged. e.g. id software's history of doing that has meant their work stays alive forever.
[12 comments hidden]
[11 comments hidden]
[7 comments hidden]
That's right, and the difference is that this one is parasitic.
[6 comments hidden]
[3 comments hidden]
[2 comments hidden]
Shouldn't we pay for the best tools if it helps researchers to be more effective?
[2 comments hidden]
If reading each others work is symbiotic it makes sense OpenAI is parasitic: whose papers are they reading in return? No-one’s.
[hidden]
The problem people are having is they clearly are interested. They think the ideas are good. In fact, too good. If they thought otherwise, they would just say it's all slop, no one cares, business as usual. And you can tell that there's this phase change in their behavior because previously you could ask chatgpt about math and it would just give you word salad and everyone knew that. Now we can all see that it's not just word salad and people are scrambling to figure out what to make of that. Obviously an accurate answer oracle is still strictly useful even if it makes no attempt to tell you why (you can even use it only to help prove your boring technical lemmas when you have ideas!), so obviously this is an emotional reaction, not a rational one.
[3 comments hidden]
We need the companies to humanly review their papers. in the same way as at other companies we use humans to review the papers.
[2 comments hidden]
(If you're going to object that it's difficult to validate the statement of the problem, please first state your level of experience doing so. It's getting tiring seeing people raise this objection and claim that a statement is just as hard as a proof over and over who don't seem to actually know any math and have never tried to write anything in Lean)
[6 comments hidden]
[5 comments hidden]
Ironically though, what I imagine will happen, is that the researchers will pay OAI to use chatGPT to help themselves eval the proofs.
[4 comments hidden]
You have the frontier labs who are marketing that it’s over and they’re building intelligent machines and you’re saying the mathematicians should ignore it? Ok
[hidden]
Mathematicians, as autonomous entities with no formal connection to any AI lab, have zero obligation to do any work for those labs. OAI can't do anything if all the mathematicians band together and say "Sorry, we're not interested".
If they do chose to engage, they are doing so entirely voluntarily, and it would strongly indicate, if they are voluntarily doing it for free, that there is value (i.e. compensation, payment, barter, worthwhile, whatever) to be had by digging in.
[5 comments hidden]
This _might_ have been true somewhat in the past (although it wasn't), but it's completely false today. Anyone with access to a sufficiently advanced model has the capabilities of analyzing these papers/proofs. It's no different than reading a codebase you might not be fully familiar with, and checking it for correctness (give an engineering analogy).
This hardcore gatekeeping of math (and by extension STEM) fields MUST stop.
[hidden]
Like I was reading some about adele rings last night, which is already going to be quite a concept for a layman to be able to even slightly describe. Then you can layer on that apparently they're locally compact, so we can talk about harmonic analysis on the additive group. Like, come on now, 99.99% of people have no hope of ever following along, and this is stuff from 75 years ago.
[2 comments hidden]
But they don’t. What they have is the ability to ask something else to do the analysis. It’s an important distinction. If the asker has the skills to evaluate the results, that’s one thing, but too many don’t and act as if whatever they got is unambiguous truth.
> This hardcore gatekeeping of math (and by extension STEM) fields MUST stop.
What must stop is the overuse of the word “gatekeeping”. Anyone is free to study these fields and work on problems. What people rightfully object to is uninformed research flooding everything with hard to verify junk.
[hidden]
Mathematicians do have some vested interest in keeping the profession from collapsing into an intellectual oligopoly, where one or two commercial players with early access to their own internal models continuously scoops everyone else and pollutes the field with externalities, and the profession itself collapses, only leaving AIs and hobbyists able to stand. At that point it'll be the AI firms who become rent-seekers. That's not really "gatekeeping" in any conventional sense of the term, it's protecting a healthy economy of ideas and the long-term development of mathematics.
[10 comments hidden]
[hidden]
[6 comments hidden]
[3 comments hidden]
[2 comments hidden]
[5 comments hidden]
[2 comments hidden]
Withdrawal is akin to submitting a paper to peer review and then when you’ve noticed mistakes, you decide to take the paper back and correct it.
Reject is when someone else notices the mistakes and tells you to take it back and correct it.
Withdrawal and reject happen all the time in a scientist’s career. They don’t necessarily mean the scientist is doing bad research, just the research was not ready. Retract usually means something more.
By dumping the papers, OpenAI skipped the typical peer review process, so peer review should be understood as what’s going on now as mathematicians look over the papers and find flaws.
[2 comments hidden]
[hidden]
[4 comments hidden]
[2 comments hidden]
[5 comments hidden]
The latest model even finds mistakes in previously published math papers!!!*
* so far, only OpenAI's math papers were faulty and needed retraction.
[2 comments hidden]
It'll be a good test to separate those earnestly trying to advance human knowledge, from those wasting my tax dollars. The later group ought to be publicly shamed and ridiculed without mercy. We need a more invective word than 'pseudo-intellectual.'
[188 comments hidden]
[24 comments hidden]
[23 comments hidden]
[19 comments hidden]
[14 comments hidden]
This concerns the Hodge conjecture (millennium prize related) paper. Seems to me like PhD nerds weren't confident bosses pushed ahead anyway.
[9 comments hidden]
[hidden]
[6 comments hidden]
[4 comments hidden]
[hidden]
Because this is what they say all the time. It's like a badge they have to wear and tell everyone they are wearing, even though we see it.
You can see the same thing with ANT. Had they looked at Mythos output, they would have realized there were only 76 items, not 79 like the bot claimed. Or the ones that were just a "it crashed" and nothing else (not a cve imo).
https://www.youtube.com/watch?v=NnV_cWeoo5Q (Linux Kernel team sharing their side of the Mythos "hacking" story)
[4 comments hidden]
[2 comments hidden]
[hidden]
Presumably the tip would be from someone who’s familiar with the area but doesn’t want attention. Which is unlikely to be someone in OAI.
[hidden]
Perhaps with AI.
A manual human check of each one would take a few month at least. In peer review, there are horror stories in math about more than 1 year before the journal accept the paper. So 3 reviewers x 700 pdf = 2000 mathematicians, that is 10%-20% of the community according to an unreliable count printed by Gemini after scrapping r/math.
Also, in most cases the only people that can understand the proof in a so short time (let's say a few months!) is the small group of people working in similar problems, i.e. the same group of 20-100 guys/gals that you meet in every conference.
[3 comments hidden]
[2 comments hidden]
Scientific progress used to be people debating and correcting other people. Now it's going to be people with AI assistance debating and correcting other people with AI assistance.
[hidden]
1. This is expected if you only use a single model family like Claude, eg. we use a different model family for code review than authoring, OAI could have done this too for their math dump
2. Ai needs a good human driver beyond the trivial or mundane, they are expert enhancing machines, not expert creating machines. This is where the community comes in. Reading Tao's ChatGPT session reveals this: https://news.ycombinator.com/item?id=49010345
3. OAI is not trying to be a member of the/any community, this is not the first story to shows this, nor do I expect it to be the last. Perhaps this is them being effective altruists today? /s
[3 comments hidden]
[hidden]
I agree LLM review is also fallible (as is human review) but the interesting part to me is that finding this sign error before publication should have been table stakes for OpenAI, it’s their own model that found the sign error.
I’m curious what was in the original prompt and what was in the prompt that led to finding the sign error, I think it matters a lot for understanding the dynamics here
[15 comments hidden]
If they can be automated, they are not necessary. If they are necessary, they won't be fully automated. It's a pretty simple experiment to run, the math "community" should bear with us. Darwin would be proud.
[5 comments hidden]
[2 comments hidden]
It's great that we're starting to see the light at the end of the tunnel, and will some day achieve a perfect market without humans. If you think about it, all the market really needs is a people to own everything, everything else can be automated, and all those annoying human workers can be eliminated.
[4 comments hidden]
To witness an arson and rejoice reveals an ugly kind of sadism.
[3 comments hidden]
[2 comments hidden]
[hidden]
I see why you might call it sisyphean, but I don't see what's ironic about it.
[4 comments hidden]
[3 comments hidden]
Even if your stated assumption was baked into the original comment, which is doubtful: the historical record shows that we will keep relearning The Bitter Lesson and each community will pretend what they do for a living is exceptional and immune because of xyz. The screams will get louder when the "greedy" and "dumb" automation comes knocking and it turns out nothing was truly immune or "nuanced ".
Getting some new hobbies may be in order, it's a Brave New World.
[2 comments hidden]
Bringing it back to math explicitly: you are essentially betting that the singularity is here, today, and that there are absolutely no downsides to breaking the pipeline which trains mathematicians (meaning that in 5-10 years at most there will be zero humans capable of assessing AI math output or independently advancing the state of the art).
[hidden]
The CS101 lesson in the first paragraph is appreciated, you should do it more often for us simpletons.
[5 comments hidden]
[3 comments hidden]
[hidden]
[23 comments hidden]
[22 comments hidden]
An absolutely ridiculous statement. There is a vast amount of mathematical knowledge that hasn’t even been written down, much less formalized.
[21 comments hidden]
[4 comments hidden]
[3 comments hidden]
[9 comments hidden]
[8 comments hidden]
[7 comments hidden]
The same happens all the time in mathematics.
[5 comments hidden]
[4 comments hidden]
[3 comments hidden]
At the minimum you ought recognize that world class experts would disagree with your dismissiveness. Trying to explain the differing positions should not earn such immediate dismissal.
[2 comments hidden]
[hidden]
You're also failing to comprehend my passing comment about professors as a minimum criterion for the diversity of what reasonable opinions on the issue ought to like. I said it order to increase an open minded discussion, whereas you then used your familiarity with academics to validate your specific views. There's a world of difference there already, and metacognitively yours is the problematic one.
[5 comments hidden]
You’re distinguishing “knowledge” from “idea” in a particular way that doesn’t correspond to common usage (see my counter examples). Without you being explicit about your definitions, I can’t tell whether what you’re saying is meaningful. It feels tautological.
Given that an executive assistant has unwritten knowledge that is necessary to do their job, where your evidence that no mathematician has analogous knowledge (using the word in the common way, not whatever way you mean it)?
It’s possible, but it’s not as obvious as you seem to think.
[4 comments hidden]
We can dig into the philosophy of these definitions, but I think the far more interesting point is that even if we grant the existence of this kind of knowledge in the minds of human mathematicians, we have passed the threshold where that "knowledge" can keep up with systems that do not have access to it. Moreover, to claim humans have a "vast amount of mathematical knowledge" that is apparently valuable and that AI systems don't have, you'd have to prove that this "knowledge" is not implied or cannot be reverse engineered from the entire corpus of mathematical writing on which AI systems are trained. You'd also have to demonstrate that this "knowledge" leads to actual results that AI systems cannot generate without it. Given the results AI systems are producing, which are far beyond human ability at this point, it is more likely that AI systems have already internalized the entirety of this so-called "tacit knowledge" and then went much further, much faster, without humans in the loop at all.
[3 comments hidden]
[2 comments hidden]
On the question of a single word out of my entire position, that is a reasonably fair statement for mathematics (which is the topic under discussion). The fact that you think it's tautological supports both that you agree with its correctness and that this is a mostly irrelevant side conversation. And whatever you think the answer should be is not somehow not subject to the constraints of logic.
[hidden]
Not your contention that LLMs likely have something analogous to what you call “ideas” (you’re almost certainly right).
Mostly irrelevant? Dunno, maybe according to your rigid ontology ;-p
Cheers!
[2 comments hidden]
[hidden]
[hidden]
And that's a bad thing. If the math community didn't exist or was weak, OpenAI would still benefit from the prestige of these results they were forced to withdraw. Withdrawing these papers has harmed OpenAI's investors, and that's totally unacceptable.
> With automated math that community as tao pointed out is at risk.
Good to hear. The problem they represent needs to be eliminated.
[108 comments hidden]
What about software developer community? AI has eliminated the need for junior software engineers. Almost no one is hiring junior software engineers. But companies still need senior software engineers. Without junior engineers how will there be senior software engineers in the future?
What is the solution? I don't think the solution is to say AI progress in software, mathematics etc. should be halted.
[22 comments hidden]
[2 comments hidden]
[hidden]
I have a data pipeline with 6 steps, A -> B -> C -> D -> E -> F. I asked Codex to make some specific optimizations to step B and benchmark them. It did what I asked. Then it decided to also benchmark the entire pipeline, and after noticing that step E was slow it decided to make some optimizations that I had not asked for on step E. It was at this point that I wondered why it was taking so long, saw what it was doing, and stopped it.
This is GPT-6.1 Sol High.
[15 comments hidden]
Right, companies won't need software engineers. They'll just need someone who can use tools to produce source code and maintain the generated artifacts, plus make domain-specific technical decisions like "what should the system do when two users update the same record as the same time" or "how should the system behave when a message in the queue cannot be processed".
We really oughta come up with a job title for these people.
[2 comments hidden]
[5 comments hidden]
“ Claude! what should the system do when two users update the same record as the same time, explain to me with full clarity”
or "Claude! how should the system behave when a message in the queue cannot be processed? Give me all the possible ways ranked from best to worst, also explain to me all these concepts so I can understand as I don’t have a cs degree".
If you think that there is no future where software engineers don’t matter then you are delusional. While the future is not set in stone the pace and trendline of AI point to this future as a MORE realistic future then the alternative.
Your example btw is ALREADY a solved problem. AI can answer it and design around it. Agents at my company already handle our infra.
[4 comments hidden]
This isn't even the right question to ask, I think you've basically proved my point. You are in charge of deciding what the system should do when two users update a record at the same time. It's extremely dependent on what you're trying to do.
> Your example btw is ALREADY a solved problem. AI can answer it and design around it.
What's the one-size-fit-all solution for concurrency management that works for every single domain and application? I'm curious.
> Claude! how should the system behave when a message in the queue cannot be processed? Give me all the possible ways ranked from best to worst, also explain to me all these concepts so I can understand as I don’t have a cs degree".
Who's going to make this decision? The CEO?
[3 comments hidden]
It’s your example. I simply took your example and asked Claude. If it’s not the right question then don’t give it out as an example.
> What's the one-size-fit-all solution for concurrency management that works for every single domain and application? I'm curious.
I’m curious how your brain concocted I said that. Examine the context of our conversation. What I mean there is that AI can solve those questions for every possible domain application.
> Who's going to make this decision? The CEO?
Armed with Claude any non technical person can make this decision.
[2 comments hidden]
Knowing the right question to ask is what makes a person an engineer.
> What I mean there is that AI can solve those questions for every possible domain application.
Yes, if you know what to ask. You're doing an excellent job of demonstrating my point!
> Armed with Claude any non technical person can make this decision.
You just disproved that by asking the wrong question. Much like you, the CEO won't even know what to ask an AI.
[hidden]
Yes, but this is orthogonal to the point and that is: AI can do it too.
>Yes, if you know what to ask. You're doing an excellent job of demonstrating my point!
No it's your comprehension that needs work. You are missing MY point while being repeatedly getting enamored with your own point. My point is that AI KNOWS the questions.
>You just disproved that by asking the wrong question. Much like you, the CEO won't even know what to ask an AI.
I didn't ask a single question bro. I only regurgitated your examples. Much like AI, half your statements are based off of hallucinations.
[6 comments hidden]
I do not know or care what my if statement turned into in x86 assembly unless it becomes a performance problem and even then, I'm not profiling or debugging in machine language. Neither do most developers these days. A message in a queue becomes something akin to that in this era.
I see that I got downvoted there. This is not something I advocate or look forward to but I feel this is where it is going.
[hidden]
Likewise, you don't care exactly how an if statement gets converted into machine code, but you do know precisely what an if statement is and how it should behave, and could identify if it was buggy, and that that part of the codebase contains a bug. If you can't do that, then there is an impossible-to-estimate probability that at some point you get stuck and no progress will ever be possible. I don't see that as a winning strategy, in the long run (but it may work very well in the short term).
[4 comments hidden]
No it isn't lol. Have you worked on any real systems with customers? Good luck telling your boss at AWS that a poison pill message stopped the payment queue from processing so they lost $100 million in sales but hey, it's an implementation detail, no big deal.
[3 comments hidden]
I have. Those systems already fail in spectacular ways and people tell their bosses that some worker process stopped working because its transaction IDs overflowed.
I bet that sounds like `the flux capacitor stopped reticulating splines` which is already an implementation detail for the boss anyway. Nothing changes.
[4 comments hidden]
[3 comments hidden]
I think that person does not need to know about locks anymore.
[2 comments hidden]
Some people hope AI will get good enough in a few years that it can innovate without human experts. Maybe? But that remains to be seen.
[hidden]
By the progress of AI from ChatGPT to now is horrifyingly fast.
Trendlines and basic reasoning point to a most probable future where the AI is superior. We can’t just say “that remains to be seen” because the alternative is the least probable future.
Anticipate the change and act prior.
[2 comments hidden]
In the age of extraction capitalism where building sustainable, profitable companies is not the goal, no one will care.
[19 comments hidden]
[14 comments hidden]
[6 comments hidden]
[5 comments hidden]
[hidden]
[3 comments hidden]
Human language is famously terrible at being unambiguous.
[2 comments hidden]
[hidden]
But that's the thing, you and the AI might be "on the same page" for one prompt, and then you aren't for the next.
> and yet judges and lawyers agree on how to interpret most of them
Um, no? Lawyers and judges frequently disagree on how laws should be interpreted. And it is often not written in a way that a layperson can easily understand.
> defined arbitration process to resolve any new ambiguities.
That is famously slow and expensive, and can resul in something very different from what the lawmakers originally intended.
[4 comments hidden]
[3 comments hidden]
LLMs eat away at all of these requirements.
[hidden]
[4 comments hidden]
The developers who knew how to write efficient low-level code found that there were no jobs for that anynmore, so today the developers who can work at that level are very few.
The same will happen with AI being the new abstraction. In another decade or two, very few people will be able to write code by hand. We'll need a rack of compute in a data center and multiple KW of power to do what we used to do on a desktop PC drawing a couple of hundred Watts.
[3 comments hidden]
[hidden]
[hidden]
[27 comments hidden]
[15 comments hidden]
Your code used to be a masterpiece, so well crafted it's easy for AI to tweak and modify because you've got everything so logically organised and scoped... and now this is what you're producing? Hard to debug monstrosities that only an LLM can realistically bolt new features or tweaks onto, because it can do the kinds of refactoring necessary each time.
We use to talk about the fact that code should be readable because you spend more time reading it than writing it, but I think that misses the key part that readable code is also typically easier to debug. If you can read and understand the code, you can follow the logic when things are wrong in production, and you can more easily reason about the emergent properties of interactions between the complex systems that are involved.
[3 comments hidden]
[hidden]
[4 comments hidden]
I jest, while in agreement with this whole comment. We used to care about fostering informed developers and maintaining high standards and good quality software.
I literally compared it to being an expert woodworker. Beautiful ornate decoration. Rich, sturdy mahogany, one of a kind, beveled edges and a fantastic stained hardwood.
Now it’s the 30$ Ikea cardboard stuff.
My biggest question is how many tables does the world need, and how many woodworkers will be required to build + maintain those table factories.
[2 comments hidden]
If you buy a house or rent an apartment in most of the US, any furniture will be stripped out, even if the only thing you plan to do with it afterwards is throw it away. Then the new tenant provides their own furniture, which they either moved from elsewhere at great expense or had to purchase on the spot. No part of this makes any sense. We don't strip countertops when selling a house even if they're unfashionable, we don't replace white goods, but somehow it's expected for the furniture.
If you buy a house in Hawaii, it's understood that you're buying the furniture that's already in the house. I assume the reason is that it's more difficult to obtain new furniture in Hawaii.
If you rent an apartment in China, it will come with furniture, because how else are you supposed to live in it? And if you're not happy with the furniture it's shown with, you negotiate with the landlord for the furniture you need.
If Americans sold their furniture when they moved instead of throwing it away, you'd see everyone using much higher-quality furniture. It would have come with their house. And providing it to houses that didn't have it yet would be an investment, just like the countertops.
[hidden]
I think with AI, most people are getting lazy to do it properly.
The other day, I deleted 65K LOC that were dead code or stupid explanations over very obvious code from a vibe coded repository which had 95K LOC (but should have 10k imo)
[3 comments hidden]
Do you even see anything underneath that rose tint ?
[2 comments hidden]
I've waded through unfamiliar code at 3am trying to figure out just why the hell things broke this time more times than I can care to remember, particularly trying to tease apart the ways that code has organically grown compared to the original intent. If I'm subjected to one more piece of "spooky at a distance" injected behaviour code I'll probably scream loud enough to be heard half way across the country.
It's still just leaps and bounds more readable than what people are slopping together (I do have some co-workers that have been extremely tightly focused on avoiding "slop" with their AI and doing a lot of tuning, and it's certainly preferable to the ones that haven't)
[hidden]
And this is the pattern I've seen on most (semi)successful projects I've worked on in the past - I'd say correlation between financial success and code quality is 0 (up to a point where the whole thing doesn't fall apart). Once scale (both in load and in code size/features) starts mattering you're stuck building on a foundation of shit. LLMs are very good at identifying and cleaning up said shit layers, as long as you're steering them towards a desireable outcome.
[2 comments hidden]
Not saying it doesn’t exist, just saying this sounds dangerously close to a boomer talking about the 1950s
[hidden]
Luckier than me. I've worked at a few places where the code was simultaneously brilliantly written[0] and also unmaintainable nonsense that caused endless problems. One place had its own object model and ORM that absolutely no-one currently at the place understood and literally every bug filed (whilst I was there) could be traced back to that code.
(Probably just a coincidence that most of those places where Perl shops but I've seen it with Go too...)
[0] In terms of "cleverness", not in terms of "maintainability" or "readability".
[2 comments hidden]
[10 comments hidden]
there are several features over the last few months that were obviously made and deployed and no one even launched the dev server and tried a single thing to verify if it was right. just "pull my ticket, do my ticket, push my ticket. i am a developer."
[2 comments hidden]
I’ve seen monthly and quarterly executing results AI generated with plainly wrong factual information.
[3 comments hidden]
[hidden]
[hidden]
Time unfortunately that you are simply not given because the whole point is to work faster so expectations of velocity have of course increased.
So you end up having to get a fuzzy idea of a given changeset, look at it a bit and evaluate quickly if it's a good change or a bad one and click it through, and on to the next thing.
At the size of PRs being whole features and doing this 5x faster than before, all you've got left as a senior is your Spidey sense. As a junior.. no idea how they cope.
We'd all love to still spend that 90% thinking time, but it used to be justified by necessity of getting the job done. Now, it's just not a luxury that is afforded.
What can ya do.. I'm slowly learning to be the best bot herder I can be, but it certainly feels like a change in skillsets. I definitely only feel like I'm able to do a good job due to enough experience and building and growing software projects over time to have an idea of what kind of decisions might bite us down the road, as well as to know what maybe can't be known but is just worth a risk. Without that kind of intuition I think I'd feel like I was really shooting blind.
[4 comments hidden]
IME a lot of people are now subject to output-rate expectations that preclude doing much else, honestly.
[3 comments hidden]
[2 comments hidden]
[hidden]
[5 comments hidden]
Some jobs will stick around in vastly diminished numbers with tasks that are completely different than what they used to be to produce the same output (e.g. farmer). Other jobs will be eliminated entirely (e.g. switchboard operator). I'm guessing things like software engineering will go the way of the farmer, with the main unknown being just how much demand for software there is.
[hidden]
In this case with AI that power will shift to the companies that run the AIs
[3 comments hidden]
The entire valuation of the AI industry is predicated on people not just losing individual jobs, but being taken out of the workforce entirely on an economic level.
They are talking about the workforce of the entire economy shrinking. People will lose their livelihoods for good.
This has the potential to be even worse than the second agricultural revolution to industrial revolution phase, which made ordinary workers lives absolutely miserable for maybe a hundred and fifty years.
This time, there will be no jobs. If you are displaced from one industry by AI, you will end up in another industry also being decimated by AI; if you get a job at all, you will do so by working lower pay than other workers, who will in turn be pushed down the ladder.
And that is if you are lucky: if you have only IT skills, why should you be the first to get a fruit picking or plumbing job?
[hidden]
Easy. Because you have the skills to increase productivity by automating it. Oh wait...
[hidden]
[hidden]
[hidden]
Ahh that's OK then. Everyone's in this same boat simultaneously in multiple industries! Cool!
> What is the solution? I don't think the solution is to say AI progress in software, mathematics etc. should be halted.
I think the solution from the maths world is to not grant these AI papers (or their human sponsors) the normal courtesies of "regular order", just as you would not with an AI lawyer or someone who was just pressing enter at a law firm.
But in the software world, nobody gives a shit, apparently. We are collectively morally bankrupt and should not be granted the regular order to help other people to decide what to do with us.
[11 comments hidden]
In fact, the world is always filled with curious people who like to go one level below.
This hysteria about losing "Junior Software Engineers" -- most of them in it for money, promotion rather than craftmanship, is over-rated.
People who love solving puzzles will always find ways to sharpen their mind.
People who love understanding things, will always find ways (AI will help them tremendously).
People who love taking shortcuts will always find ways for it (AI or not)
[hidden]
Eh? Apart from it not being what they are paid to do on their 9/9/6 jobs, when will they have the time to make it happen?
What is going to happen is that the remnants of the open source community will do the job of educating juniors for free, when the university degree system collapses. Just like it currently keeps a bunch of systems going with inadequate compensation.
The corporate world gets the problem off its balance sheet. Again.
[3 comments hidden]
The actual argument is "there will be no jobs for Junior Engineers, so there will be far fewer, and as a result there will be a huge shortage of Senior Software Engineers".
You are responding to the problem as if it some kind of extinction event, like a rare bird, where if we can find a breeding population we save the day. A few curious people, self training for the love of the game, and as a result we still have a few Software Engineers so everything is fine. It is not like that, and I haven't heard anyone suggest that is the issue. The potential problem is a massive shortage of workers with skills that are currently essential to the functioning of a large fraction of the economy, whom we might still need in the future.
The continued existence of talented enthusiasts does not establish an adequate workforce pipeline. If paid entry level experience contracts, what replaces it, and why should we expect that replacement to operate at sufficient scale?
[2 comments hidden]
[hidden]
[6 comments hidden]
[5 comments hidden]
If you have made it to the point of being a Junior Developer, I can assure you food and shelter is not a problem for them. You just have to adjust to a standard of living like the other 7 Billion people on this world.
Also, if a Junior Developer can show me(or anyone) they have built an entire system on their own and explain key concepts, there is no dearth of jobs for them
[4 comments hidden]
[2 comments hidden]
[hidden]
Trying to actually get the framing to be more reasonable is just hard because so many people have oversold it and large swaths of the public are sick of hearing about it.
[hidden]
[hidden]
My take is: the bar to what counts to HR as "senior" will go down as businesses everywhere try to adapt and hire more seniors - "senior" now just a name, as it becomes the new "junior". Then everyone will pat each other on the back till it all goes down in the flames of bankruptcy.
[hidden]
In other words, they are increasingly devauing their own senior position and discarding their hard earned skills that make them seniors in the first place.
There will be no more seniors. Tech companies will hire junior llm agent wranglers who took a class in undergrad doing this. That is probably the nearterm.
[2 comments hidden]
[hidden]
You have to do the hard thing eventually, or you never get anywhere.
[hidden]
[2 comments hidden]
Disclaimer: I am not a fan of AI, and I am currently writing software without any help from AI, as I prefer.
[hidden]
Would we need doctors, accountants or analysts? Probably not. At some point farming will be fully automated too, and so will grocery distribution and food preparation. At that point we will have arrived in the post-scarcity world. This is the promise of AI.
The problem is that the post-scarcity world will arrive gradually, not suddenly. Some jobs will be automated sooner than others. The ones that are not yet automated will expect payment for services. People who just lost jobs to automation won't have income to pay for those not-yet-automated services. But that's only until all jobs are automated.
There will be tremendous social upheaval and unrest during the transition to post-scarcity world.
[3 comments hidden]
[hidden]
Just for a thought experiment lets say junior hiring actually goes to 0% starting today and there is no other route into software dev for example maybe it is illegal for anyone under 22 today going forward to work in dev. Maybe something like ~2-2.5% of the workforce retires each year? And lets just define senior as 10+ YOE so 75% of the current batch of ~40 year working timeline. For simplicity lets just say the other 25% don't ever become senior devs and in 10 years the total number of senior devs has reduced 33% due to retirements. ChatGPT public launch was less then 4 years ago! Look at the ludicrous progress in the timespan. Even if progress suddenly massively slows or hits a wall it is currently hard to fathom it not improving at a rate of 3% a year.
More realistically I think we just have no ability to predict wtf things will look like 10+ years out at this point which is the point in this artificially constrained timeline where a reduction of senior devs just due to time would even start to be noticeable I think.
[4 comments hidden]
[hidden]
It is pointless to call out the little wins humans still have because again, those wins are temporary.
We need real concerted effort into asking: what is the point? For me the only answer I came up with is: fun.
[hidden]
[9 comments hidden]
[2 comments hidden]
There is however a new problem of scale. Erdös was a human and still managed to create work for an entire generation of mathematicians, how much of a mess will an automathician create?
[3 comments hidden]
Interesting that math gets so much attention, when actual advances to material science, biology and chemistry have much higher ramifications and economic benefits. I assume progress there is kept under wraps until they can capture the economic benefits. If they can do that, then the insane valuations may actually be valid.
[2 comments hidden]
Or progress is not as straight forward in those fields as in math.
[hidden]
In a field where a single person can maybe produce and characterize a single sample per day, AI will only really start speeding up progress once you combine it with robotics.
[3 comments hidden]
[4 comments hidden]
[hidden]
Understanding the proofs is a different story unfortunately.
[2 comments hidden]
We let the experts investigate. If the results are dodgy, then the next batch of results will have to do more upfront work to demonstrate their worth. If there is gold in them hills, then this is exciting though very disruptive for the math community.
[2 comments hidden]
[4 comments hidden]
[3 comments hidden]
[2 comments hidden]
[hidden]
[58 comments hidden]
[5 comments hidden]
[4 comments hidden]
You might think this is not very useful, maybe - but that’s not a reason to retract..?
[hidden]
[50 comments hidden]
> As part of our GitHub repository, we are sharing formalizations of many of the proofs in Lean, a programming language that allows mathematical proofs to be checked by a computer. We will update the repository with more formalizations as we obtain them.
Meaning they published all results before checking all of them, and intended to add more Lean proofs later. In the linked post they state ~42% of the posted results now have formalized proofs, some were added, some verified, and I assume this means that some results turned out to be wrong.
[47 comments hidden]
If your AI tool can help advance mathematical research, share the tool with mathematicians. Using it like this is irresponsible.
"AI will kill us all": no. Greedy humans will kill us all. With AI.
[5 comments hidden]
[3 comments hidden]
[hidden]
[hidden]
EVERYONE needs to re-visit how they work and what investment is needed.
You can't go around and don't think that no one has to reevaluate how to add ai to research and development.
I have to do this, my company has to do it and for sure a university has to do this too.
[12 comments hidden]
[2 comments hidden]
[5 comments hidden]
[2 comments hidden]
[hidden]
[hidden]
Lean 4 is relatively a new thing, last time I checked the formalization of undergraduate level mathematics isn't entirely done yet.
example https://ai.math.uw.edu/projects/spring-2026/
Lean itself is very hard to get rigorously correct, if you have every tried it yourself. I am not surprised if some AI even tries to benchmaxx Lean 4 by some loopholes
[6 comments hidden]
Also you are deeply confused about what accelerationism is, I believe. Sorry.
[5 comments hidden]
Whom? Did they create their own board of mathematicians that would agree with them? See below Terence Tao's blog, sharing a statement from the Association for Human Mathematics.
https://terrytao.wordpress.com/2026/10/07/ahm-statement-on-o...
[3 comments hidden]
> Mathematicians have a particular vision of progress that is informed by history and field-specific considerations.
really sounds like something a side-quest association would produce in a panic response to someone trying something different. It's just an _ad hominem_ and gatekeeping argument.
[hidden]
Also jeez just noticed their name... that's unfortunate. I guess mathematicians don't make great rhetoricians/politicians/marketers! Someone get a sophist or two over there to help them out STAT
[12 comments hidden]
[hidden]
That's an analogy among many others, but the point is that OpenAI should use their tools responsibly. If they have 700 potential ground-breaking but unproven results, they should share it in a way that they do not get free (possibly unwarranted) publicity for it.
[7 comments hidden]
[5 comments hidden]
[4 comments hidden]
[3 comments hidden]
They had over a hundred papers, they basically did a GitHub dump and a pretty bare blog post that mentions they have a retraction policy. Should they have done it anonymously? Is the blog post the problem? Is it that it’s on GitHub? Or what?
[2 comments hidden]
It's really not that hard to understand. My issue is NOT with the fact that they publish results. It's purely about how it's done and what the misleading claims are that come along with them.
[hidden]
> in many fields it is (was?) normal to post a preprint and leave it up for the next year while the handful of other people working on the topic digest it and agree on whether it’s right or not, before even submitting to a journal. Sometimes the others would find an isssue in an argument and you hopefully manage to fix it or potentially retract / not submit to a journal.
A pre-print has a chance to be incorrect. Pre-prints do sometimes end up incorrect and need adjustments or even retraction. It sounds like you define "confirming it before you publish" as being 100% certain it's correct. Nobody is 100% sure their pre-print is correct. If it was 100% correct and never retracted there'd be no need for this process of sharing and digesting it before actually submitting. You do your best and maybe someone else thinks about it in a new light you missed and it's suddenly wrong.
Is the rate of retraction and/or correction going to be worse for OpenAI than it would be in normal circumstances for a normal human author? Only time will tell, it's only been a few days.
But all these complaints really just look like pulling at whatever straw to deny what's happening.
They could've instead hired a select group of mathematicians to spend then next 6 months checking these works, trying to understand and refine them, build on them etc... You and the others here would have 100% complained. Even more bitterly. And then you would've perhaps been right: it would have looked like there's a select in-group of mathematicians that get to see and work on the new breakthroughs while the rest of the field is kept in the dark. Instead of now where it's public, anyone can do that. Maybe that's what will happen in the next release. It's not gonna be good though.
[hidden]
[hidden]
[3 comments hidden]
And im completly lost on why you think sharing progress is irresponsible? Its not a recipe for building a nuclear weapon at home in 5 easy steps.
THese are Math proofs.
Either a Mathematican ignores it, or not. Thats the only risk.
[3 comments hidden]
What do you mean? They're publishing the Lean proofs themselves. Who's forced into anything?
[2 comments hidden]
They are just putting out a bunch of weirdly written extremely long and technical papers and saying: Hey, here is the solution (we hope there are no mistakes).
[hidden]
[hidden]
[3 comments hidden]
> The vast majority of results were obtained with the same procedure using an unreleased internal OpenAI model. On average, each result used three hours of ChatGPT Pro thinking compute with that model. Over the course of the evaluation, the model was posed approximately 4,000 problems. Aggregating the output into result families and manuscripts and requiring an appropriate level of significance led to the catalog outlined above.
seeing the full list of problems would be the most interesting part of this whole situation. it could give some insights into what kind of attributes of problems cause issues / are easy to solve for LLMs. (edit: they posted results for ~700 of the 4000)
[hidden]
This is a "complaint" that Tao had (I think it was on his blog) is that if mathematicians could see the failures, it might provide insight of where/how the models struggle. Of course, it's unclear if these failures can be addressed with more chips/training/etc.
[hidden]
1791446917 | OpenAI Withdraws 3 Math Papers | https://github.com/openai/math/blob/main/history.md | https://news.ycombinator.com/item?id=50003107 | 118 comments
1791451788 | AHM Statement on OpenAI's October 6 Release of Mathematical Documents | https://www.ahmath.org/statements | https://news.ycombinator.com/item?id=50003677 | 15 comments
1791460215 | AHM Statement on OpenAI's October 6 Release of Mathematical Documents | https://terrytao.wordpress.com/2026/10/07/ahm-statement-on-o... | https://news.ycombinator.com/item?id=50004713 | 0 comments
1791478724 | OpenAI mathematics papers have no correspondence address | https://mathoverflow.net/questions/515845/how-can-one-contac... | https://news.ycombinator.com/item?id=50008412 | 1 comment
1791487754 | OpenAI's New Math Breakthroughs Put AI's Role in Science Under the Microscope | https://medium.com/@d02514047/openais-new-math-breakthroughs... | https://news.ycombinator.com/item?id=50010818 | 1 comment
[5 comments hidden]
Is this of practical use, or just a proof for now?
[hidden]
Look at examples here: https://en.wikipedia.org/wiki/Galactic_algorithm
[hidden]
What proportion have actually been verified by the mathematical community by their own standards of proof? There is dispute in general about what constitutes proof, and which statements qualify as having been proved.
But even formalization doesn't mean the formalization corresponds with the stated problem.
[hidden]
[hidden]
I don't have a problem with them publishing. I don't have a problem with the process and how they are interacting with it. I am delighted that they are actually acting as stewards of these works.
All that aside, they should be paying the people verifying the problems. The thing that really gets me is that we know anything published in the process of verifying this is going to be vacuumed up into the next training session.
[hidden]
[hidden]
[4 comments hidden]
[9 comments hidden]
[7 comments hidden]
If that’s the case in a way its a similar delusion that average people are experiencing with their own AI use.
[2 comments hidden]
[hidden]
[hidden]
[2 comments hidden]
[hidden]
You could almost draw a comparison between that and inexperienced consultants making business changes, claiming glory and then disappearing before the thing falls apart.
[6 comments hidden]
[5 comments hidden]
[4 comments hidden]
[3 comments hidden]
But that also means marketing and PR should shut the fuck up until things are verified. And they should more clearly annotate lack-of-verification status on their repo.
[2 comments hidden]
Goals of marketing people obviously don't aligned with science goals.
[hidden]
This heuristic indicates that at most a handful of them will be of marginal value.
[hidden]
[2 comments hidden]
[hidden]
[2 comments hidden]
[8 comments hidden]
This can and should erode our trust in every single proof OpenAI published. The model is clearly faliable despite the lean proof, and clearly the output was't actually checked properly before release. Once these proofs are peer reviewed and published in a journal we might be able to trust them again but until then they are just slop, sadly.
I think OpenAI actually did the right thing by sharing everything with the whole community right now but I also hope that some significant credit will now go to the reviewers who confirm these 'proofs" actually work.
[7 comments hidden]
Humans produce flawed Lean proofs. Indeed LLMs were successful at finding and fixing many issues in the “core” standard library if I recall correctly. Humans regularly produce flawed papers and have minor issues require fixing. And when it happens it often isn’t as prompt and clear as this.
[hidden]
More seriously the problem is the complete utter lack of care OpenAI has shown in their desperation to demoralize mathematicians with their new LLM. In their words, it took 3 hours of ChatGPT pro per result, why not spend a hundred hours per result formalizing it, checking if the formalization matches the natural language proof, and whether the argument could be made more simpler and readable. Any human paper has hundreds of hours of work put into it, but OpenAI who is absolutely adamant in demoralizing the mathematical community and demonstrating their superior “intelligence” will only spend 3 hours, write unreadable, inscrutable proofs, not formalize all of them, and then dump it on the mathematical community for some reason.
[5 comments hidden]
In fact we already follow exactly this process for human papers and have done for a very long time. Publications without peer review are treated with great suspicion. This is how we end up with journals of varying levels of prestige and rigour.
The process isn't flawless and there are huge problems with retractions, as well as weird financial incentives and rent extraction but there is definitely an increased level of trust in a paper published in Nature.
[4 comments hidden]
[3 comments hidden]
I am pretty sure that journal publication is the norm. I do know mathematics although I am not one myself. The mathematicans I know want/need journal publications as unpublished research isn't counted in their research output. I know Arxiv is typically used for preprints while awaiting review but publication is always the aim.
Don't take it from me though. zbMath did a study finding 60-70 percent of arxiv papers on math end up in peer reviewed journals. https://ems.press/content/serial-article-files/47210.
[2 comments hidden]
If they’re untenured sure. Tenured mathematicians care a significantly deal less.
> Don't take it from me though. zbMath did a study finding 60-70 percent of arxiv papers on math end up in peer reviewed journals.
Elite journal capacity didn’t expand. There’s just a bunch of mid-tier journals helping to absorb the increased publishing to help these ~indentured~ untenured professors.
[hidden]
You are right that most work gets published in mid tier journals. These are still peer reviewed. The difference is that they have a lower bar for impact. The bar for correctness is still very very high however.
[hidden]
If they want any credibility on these matters, they should be funding a lab of elite mathematicians and putting them to work alongside their own researchers to attempt to solve these properly within the realm of the community.
Propagating arse slop cannons onto social media and calling it a day is not that.
[2 comments hidden]
[hidden]
[hidden]
[hidden]
[8 comments hidden]
If you want to see paper retractions, you ain’t seen nothing yet.
[7 comments hidden]
[6 comments hidden]
[hidden]
Consider now the work needed to operate the detectors, acquire the data, do calibrations and validations.
Notwithstanding designing, building and commissioning the detectors.
Please don’t assume lab automation.
[hidden]
[hidden]
I have a hard time thinking a neural net "predict the next token" is going to do things like solve math proofs or find cures for cancer when it gets very basic things blatantly incorrect sometimes.
[2 comments hidden]
[3 comments hidden]
[hidden]
[hidden]
[10 comments hidden]
* The reason you shouldn't consider the withdrawals to be caused by errors is because this is pretty standard in math and development. "Errors" like this are a core aspect of science and it happens _all the time_. And from my research LLMs have a far lower error rate than even the best human scientists.
[2 comments hidden]
Or is it the kind of research you'd prefer to keep shrouded in mystery?
[7 comments hidden]
[3 comments hidden]
[hidden]
[2 comments hidden]
Even Einstein retracted one of his earlier papers re the cosmological constant... And he is one of the greats. Although not purely math related, it still counts.
[3 comments hidden]
Or if the "peer review" holds up for the remaining results, then it's fair to say that the AI hype is real and the world is about to change dramatically and faster than anyone can comprehend.
So which is it?? LLMs can do some really impressive coding. Bug fixing. Exploit finding. It has reasoning abilites that advance every day. Solving real math problems like this is one thing I was waiting on. It will be interesting to see if it holds up.
If it does, we should expect many other advancements to follow in many other areas. Disease, material science, fusion?
I mean, even if just a few results ultimately hold up to scrutiny, isn't that something that would have been regarded as a major advancement regardless of if it was AI?
The cynical view still makes me think that at the end of the day all the models can do is predict the next word. And as a result, they will be very limited to certain tasks like coding. Math reasoning is much different from writing code. Time will tell.
[hidden]
there is also will be new option of "maybe correct LLM proof", which humans will never be able to comprehend and verify.
[hidden]
Can't be, as IPO isn't until next year, and I'm pretty confident that is there are errors, said actual mathematicians will find at least a few in the remaining ~2.75 months left in the year+. If there really are errors and a decent amount aren't discovered before the IPO then I'll have serious questions about the math community.
[hidden]
[3 comments hidden]
It just hallucinates an answer and then makes up workings to go with it! Just like when they start hacking and lying because the problem is impossible…
[2 comments hidden]
No it’s not. In fact, much of it is formally verified, which makes it far more reliable than most human-written proofs.
Btw, the most famous human-written proof of the past half-century (Fermat’s Last Theorem) had a massive flaw that took two years and major help from other mathematicians to fix, while the most (in)famous human-written proof of the past 15 years (abc conjecture) is now widely believed to be false.
But people hear what they want to hear I guess.
[hidden]
[hidden]
[4 comments hidden]
AI is this strange modernist machinary that kind of threatens that brhaminic role... its almost like the vatican vs post industrialization world .. where they still have to keep making the case for why religion/priesthood/god is important... even as the tech/science world starts operating on totally different terms...
[hidden]
[3 comments hidden]
[2 comments hidden]
[hidden]
[6 comments hidden]
But how much more progress would have been made in the last 100 years if they collaborated in real time? Watching the work someone is doing and spotting errors, or making suggestions would be a good thing. Unless the goal is simply to claim credit for a discovery vs the discovery itself.
[hidden]
[hidden]
[2 comments hidden]
[hidden]
If someone had a blog about their research on a math problem, and they figure out one day that what they wrote a week before was wrong (retracting it), would that be a bad thing? I don't think it would be bad, but that's not how academics approach things.
chubot[16 comments hidden]
If so, why did they mix proofs that were verified with Lean, and proofs in natural language?
I was wondering that while reading Aaronson's blog:
https://scottaaronson.blog/?p=10169
Or at least, we’re pretty sure that it’s a proof! There’s a Lean certificate, as there are for some of the other 372 breakthrough results (not all of them). But it also appears that no human has understood just about any of these proofs yet
It seems that the obvious thing to do would be to release in TWO parts: the ones that are verified, and the ones that might have some good ideas but also might have some mistakes. Presumably the latter would be much more epxensive for humans to verify.
fspeech[14 comments hidden]
malisper[10 comments hidden]
Can you explain this? How would having a lean proof of the program make it more likely the proof is weak?
fspeech[3 comments hidden]
YeGoblynQueenne[2 comments hidden]
And people laughed at Doug Lenat and Cyc for wanting to encode all knowledge as a set of rules.
fspeech[hidden]
fspeech[6 comments hidden]
malisper[5 comments hidden]
Sure, even a Lean proof can be wrong. But you made it sound like that because there's a Lean proof, the claims are more likely to be wrong:
>>> Given the constraint (3.5 hours of model effort) having a Lean proof is actually a sign that some of the results are likely weak
fspeech[4 comments hidden]
malisper[3 comments hidden]
fspeech[hidden]
fn-mote[hidden]
chr15m[hidden]
That is very useful context, thank you!
le-mark[2 comments hidden]
Filling these gaps would be a great use of llm tokens!
zodiac[hidden]
Roark66[hidden]
A bit like submitting an llm written PR to an open source project without reading and understanding it.
Have your LLM generate 372 mathematical "breakthroughs" and then have 2 thousands mathematicians spend two weeks trying to understand each to tell you maybe one or two are viable...