When mathematicians say a problem is “open,” they don’t mean unfinished — they mean decades of the smartest humans alive haven’t been able to crack it. OpenAI just published results claiming its models made meaningful progress on ten of those problems, spanning geometry, cryptography, and complexity theory. That’s not a benchmark score or a chatbot improvement. That’s AI doing math that humans genuinely couldn’t do. If even half of what they’re claiming holds up to peer scrutiny, this is one of the more significant moments in the short history of AI-assisted science.
Why Mathematical Research Is Such a Hard Test for AI
There’s a reason AI labs love showing off math scores. Math is unambiguous. Either a proof is valid or it isn’t. You can’t hallucinate your way to a correct theorem the way you might generate a plausible-sounding but wrong historical fact. This makes mathematics one of the few domains where AI capability is genuinely measurable — no human rater subjectivity required.
For years, the story on AI and math went like this: models could solve olympiad-style competition problems, but real research mathematics — the kind where the goal isn’t known, the path isn’t clear, and the tools haven’t been invented yet — was firmly out of reach. OpenAI’s new announcement challenges that story directly.
The company hasn’t just claimed benchmark improvements. It’s pointing to specific, named open problems in specific subfields, with results it says are new contributions to human mathematical knowledge. That’s a very different kind of claim, and it deserves careful attention — plus a healthy dose of skepticism until the math community has had time to dig in.
What the Ten Advances Actually Cover
OpenAI hasn’t released a single model or product here. This is research output — results generated with AI assistance across several branches of mathematics and theoretical computer science. The problems span three broad areas:
- Geometry: Progress on structural problems involving shapes, spaces, and their properties — the kind of work that feeds into fields from physics to materials science.
- Cryptography: Advances touching on the mathematical foundations that secure digital communications. This one has obvious real-world stakes and will likely attract the most scrutiny.
- Complexity theory: Questions about what computers can and cannot compute efficiently — the deep theoretical layer beneath everything from algorithm design to AI itself.
What’s notable is the breadth. These aren’t ten variations on the same theme. They’re problems from different communities, with different toolkits, and different standards of proof. For an AI system to contribute meaningfully across all three suggests something more general is happening than narrow domain memorization.
How OpenAI’s Models Approached These Problems
The approach, based on what OpenAI has shared, involves using reasoning-focused models — likely in the o-series family — to explore proof strategies, generate candidate solutions, and verify logical steps. This isn’t the model spitting out an answer in one shot. It’s closer to an extended research process where the AI iterates, checks its own work, and pursues multiple paths before converging on something defensible.
This matters because it’s structurally different from how these models perform on standard math benchmarks. On those tests, speed and accuracy on known problem types is what counts. Here, the task is open-ended: find something new. That’s a harder cognitive challenge, and it’s the one that actually matters for science.
The Cryptography Results Are the Most Consequential
I want to flag the cryptography piece specifically, because it sits at an unusual intersection of pure mathematics and immediate practical relevance. Cryptographic systems — including the ones protecting your bank account, your messages, and arguably national infrastructure — rest on mathematical problems believed to be hard. Progress on those underlying problems, even theoretical progress, is something security researchers take seriously.
OpenAI hasn’t claimed it broke any existing cryptographic scheme. But advances in the mathematical foundations of cryptography can shift how researchers think about long-term security assumptions. Expect the cryptography research community to scrutinize these results closely over the coming months.
What This Actually Means — And What It Doesn’t
Here’s the thing: OpenAI announcing mathematical breakthroughs is very different from those breakthroughs being validated. The formal mathematics community moves slowly by design. Peer review for a significant result can take months or years, and even then, subtle errors in proofs have survived review before being caught later. The history of mathematics has no shortage of “solved” problems that turned out not to be.
That said, OpenAI has presumably had internal mathematicians verify these results before going public. They’d be taking an enormous credibility hit if the proofs collapsed under scrutiny. I wouldn’t be surprised if several of these hold up — and if they do, the implications run deep.
For the AI Research Community
If AI models can genuinely contribute to open problems in mathematics, the question of what constitutes “human-level” intelligence gets a lot more complicated. Mathematics has long been treated as a proxy for deep reasoning — not just pattern matching, but actual structured thought. Results here would push back hard against the critics who argue current models are sophisticated autocomplete.
It also reframes how labs like Google DeepMind — which has its own serious mathematics research program through AlphaProof and related work — are competing with OpenAI. DeepMind has been arguably more transparent about its mathematical AI research; OpenAI countering with ten open-problem advances is a significant move in that ongoing rivalry.
For Mathematicians and Scientists
This is where things get interesting and a little uncomfortable. If AI can contribute to research mathematics, what does that do to the career structure of academic mathematicians? It probably doesn’t eliminate the field — mathematics requires deep intuition, taste, and judgment that models still don’t fully replicate. But it likely changes what skills are most valuable. Knowing how to frame a problem well, how to interpret AI-generated candidate proofs, and how to push models toward productive directions becomes more important than raw computational ability.
We’re already seeing this dynamic in software development, as covered in our piece on AI coding agents doing real science. The pattern — AI as a powerful collaborator that shifts rather than replaces human expertise — seems to be repeating in mathematics.
For Everyone Else
The direct impact on everyday users is indirect but real. Progress in theoretical computer science shapes the algorithms behind every piece of software. Advances in cryptography eventually feed into security standards that protect digital infrastructure. Geometry results can inform materials science and engineering. These are slow pipelines, but they’re real ones.
OpenAI’s broader ambition here also connects to its stated mission in ways that feel more concrete than most announcements. The company has talked at length about AI accelerating scientific progress — this is what that looks like in practice, rather than as an abstract promise. For more context on how OpenAI is framing its long-term goals, our breakdown of OpenAI’s Abundant Intelligence plan is worth reading alongside this announcement.
Key Takeaways
- OpenAI’s models contributed to ten open problems across geometry, cryptography, and complexity theory — domains where progress has historically required years of expert human effort.
- The cryptography results will draw the most immediate scrutiny, given the real-world stakes of foundational security mathematics.
- These aren’t benchmark scores — they’re claimed contributions to mathematical knowledge, which is a significantly higher bar.
- Independent verification by the mathematics community is the next critical step before these results should be taken as settled.
- Google DeepMind’s AlphaProof program is the most direct competitor here, and OpenAI’s announcement escalates that rivalry.
- The practical downstream effects — on security, algorithms, and engineering — are real but will take years to materialize.
Frequently Asked Questions
What exactly did OpenAI’s AI solve?
OpenAI’s models made progress on ten long-standing open problems across geometry, cryptography, and complexity theory in theoretical computer science. These are problems that hadn’t been solved by human researchers, not standard test questions or competition problems. The specific technical details of each result are documented in OpenAI’s research publication.
Has the mathematical community verified these results?
Not yet, at least not publicly. OpenAI has published the results, but formal peer review by independent mathematicians takes time. Some results may hold up immediately; others may require revision or turn out to contain errors. This is normal in mathematics — the verification process is the most important step.
Does this mean AI can now do all of mathematics?
No. Contributing to ten specific problems — however impressive — doesn’t mean AI can tackle arbitrary mathematical research. There are entire classes of problems that require conceptual leaps, aesthetic judgment, and connections between distant fields that current models still struggle with. This is progress, not a finish line.
How does this compare to DeepMind’s math AI research?
Google DeepMind has been working on mathematical AI through projects like AlphaProof and has published results on olympiad-level problem solving. OpenAI’s claim to have contributed to open research problems — rather than competition problems with known answers — would represent a step beyond what DeepMind has publicly demonstrated, though the two programs are hard to compare directly given different research methodologies.
The mathematics community’s response over the next several months will be telling. If these results survive peer scrutiny, they’ll mark a genuine inflection point in what we expect AI to be capable of — not in consumer products or enterprise software, but in the hardest intellectual work humans do. OpenAI seems to be betting that this is the kind of proof point that matters most, both for its credibility and for the longer-term case that AI can accelerate science at a meaningful scale.