OpenAI claims to have crossed a symbolic milestone in mathematics, with external validation to back it up
OpenAI maintains that one of its reasoning models helped solve a geometry conjecture formulated in 1946, a problem described as having remained open for nearly 80 years. The information, reported by TechCrunch in an article titled “OpenAI claims it solved an 80-year-old math problem — for real this time”, marks a return to particularly sensitive ground for the company: that of spectacular announcements about the scientific capabilities of its models, after a previous episode that part of the community considered premature.
The difference this time lies in the word that comes up in every reaction: validation. According to the details relayed by TechCrunch, OpenAI is not merely claiming that a model produced an elegant lead or a promising intuition, but that a result was examined by outside mathematicians. In a field where the slightest logical flaw invalidates an entire proof, that nuance profoundly changes the scope of the announcement. It does not automatically turn a model into an autonomous researcher, but it shifts the debate: we are no longer talking only about generating plausible mathematical text, we are talking about a result submitted to formal verification by specialists.
The subject is all the more explosive because mathematics has long served as a truth test for artificial intelligence. Large language models excel at rephrasing, synthesizing, and imitating reasoning. By contrast, they have often shown their limits as soon as it comes to maintaining a rigorous deductive chain over several steps, without hidden errors or semantic drift. That is precisely why an old geometry conjecture, if it really was solved with the decisive help of an OpenAI model, represents something other than a communications demonstration.
For French-speaking audiences, the stakes go far beyond anecdote. In France as in the rest of Europe, the adoption of AI in research, engineering, and R&D is now facing a central question: when can a model output be trusted? Laboratories, engineering schools, quantitative teams, and industrial companies are not just looking for impressive tools. They are looking for systems whose proposals can be traced, verified, reproduced, and integrated into serious scientific workflows.
This announcement also comes at a time of intense competition around “reasoning models,” models explicitly optimized for solving complex problems. OpenAI, Google DeepMind, Anthropic, xAI, and open-source players alike are all trying to convince the market that they are not merely building conversation machines, but systems capable of assisting, or even accelerating, high-level intellectual work. In this battle, mathematics plays a special role: it offers clearer success criteria than many other disciplines, even if the boundary between assistance and discovery remains delicate to establish.
Caution therefore remains essential. A validated resolution of one specific problem means neither that models “understand” mathematics in the human sense of the term, nor that they suddenly become reliable in all scientific contexts. But if the information reported by TechCrunch is confirmed in its technical details and in its academic reception, then OpenAI may have obtained what generative AI often lacks: a concrete case where the demonstration of value does not rest only on an internal benchmark or a demo video, but on the examination of a result by an expert community.
An old AI dream: moving from assistance to original reasoning
To measure the significance of the announcement, it must be placed in a longer history. Since the beginnings of artificial intelligence, mathematics and logic have occupied a central place in the sector’s imagination. As early as the 1950s and 1960s, AI pioneers saw theorem solving as a privileged testing ground for machines’ ability to reason. The first symbolic systems, long before the era of large language models, had already shown that a computer could explore proof spaces, apply formal rules, and in some cases rediscover known proofs.
But those classical systems were very different from today’s models. They relied on explicit representations, coded rules, and logical search engines. Contemporary large models, by contrast, learn from immense text corpora and produce answers through statistical prediction. Their strength is flexibility; their weakness, for a long time, has been robustness. They can “look” like they are reasoning without guaranteeing the validity of every step. Hence mathematicians’ persistent mistrust of announcements of breakthroughs coming from generative models.
OpenAI knows this tension well. The company first established itself with the general public through ChatGPT, launched at the end of 2022, and then gradually shifted its messaging toward models more capable of planning, problem decomposition, and multi-step reasoning. This evolution responds to a recurring criticism: an impressive conversational assistant is not necessarily a reliable tool for science, law, finance, or engineering. Hence the sector’s growing investment in architectures, training techniques, and evaluation protocols meant to improve the coherence of reasoning.
The previous false start mentioned by TechCrunch also explains why this new announcement is being scrutinized with particular attention. OpenAI had already suggested, in a highly publicized way, that a model had reached a remarkable level on a difficult mathematical problem. But the excitement quickly ran into specialists’ skepticism, as they pointed to the gap between an appealing proposal and an accepted proof. In the academic world, especially in mathematics, media enthusiasm has no demonstrative value. A conjecture does not fall because a company claims to have found a solution; it falls when a proof withstands the meticulous examination of competent peers.
This reminder is essential, because the AI industry often tends to blur several levels of success:
- solving a standardized exercise on a benchmark;
- producing a useful lead for a human researcher;
- generating a correct proof of a known result;
- obtaining a novel result later validated by the community.
These four levels are very different, both scientifically and commercially. The first concerns performance evaluation. The second can already be valuable in practice. The third touches on rigorous formalization. The fourth, by contrast, enters the realm of contributing to research. What OpenAI is claiming, if we stick to TechCrunch’s framing, comes close to that last level, even if the exact formulation of the model’s contribution remains crucial: did it find the structure of the proof on its own, propose a decisive intuition, or accelerate human work already under way?
The question is not secondary. In scientific research, originality is rarely measured by an isolated gesture. A proof is often the product of back-and-forth exchanges, attempts, corrections, intermediate tools, and discussions. AI can play several roles in this process, from the most modest to the most ambitious: calculation assistant, combinatorial exploration engine, counterexample generator, brainstorming partner, or source of unexpected ideas. The potential novelty of OpenAI’s announcement is that the company seems to want its model recognized no longer as a simple productivity accelerator, but as an actor in an original mathematical discovery.
In the European context, this distinction resonates particularly strongly. French and European research institutions have so far adopted a generally cautious line on generative AI: strong interest in task automation, but heightened vigilance regarding scientific reliability. The CNRS, Inria, grandes écoles, and universities are already working on AI tools for assisted proof, formal verification, or literature analysis. But the idea that a closed commercial model could contribute to solving a historic conjecture adds a new dimension: that of potential dependence on proprietary systems in the production of knowledge.
What exactly OpenAI is announcing, and why external validation changes the game
According to TechCrunch, OpenAI claims that one of its reasoning models solved a geometry conjecture open since 1946. The most important point is not merely the age of the problem, which naturally fuels the announcement effect, but the fact that the company insists on validation by external mathematicians. After previous controversies, OpenAI seems to have understood that in fundamental research, authority does not come from the brand, nor from the model’s perceived sophistication, but from independent scrutiny.
In the AI ecosystem, this kind of validation is rare. Companies frequently publish scores on in-house benchmarks or standardized test sets, sometimes with protocols that are difficult to compare from one player to another. The results are then impressive, but often debated: contamination of training data, benchmark-specific optimization, use of external tools, or simply difficulty reproducing the experiment. A mathematical proof, by contrast, theoretically offers cleaner ground. Either it holds, or it breaks. In practice, of course, it is still necessary to determine who produced what, under what conditions, and with what degree of human intervention.
The previous false start makes this new caution strategic. OpenAI can no longer settle for a lab narrative or an anecdote relayed on social media. To convince, it needs names, steps, a minimum of methodological transparency and, above all, experts willing to say publicly that the result is serious. TechCrunch highlights precisely this change in tone: the company is not just trying to impress, it is trying to restore credibility on a subject where the scientific community does not forgive approximations.
This point is crucial for understanding the announcement’s potential effect. If an OpenAI model really did help solve a conjecture that had remained open for decades, then the discussion shifts from “LLMs hallucinate” to “under what conditions can a reasoning model produce scientifically usable results?” That is a much more mature debate, and a much more useful one for industrial and academic players.
The notion of a “reasoning model” also deserves clarification. For about two years now, the sector’s main players have been highlighting systems capable of devoting more computation to solving a problem, generating intermediate steps, exploring several avenues, and correcting certain errors along the way. OpenAI has been one of the most visible promoters of this approach, with messaging centered on models that “think longer” before answering. In practice, this does not guarantee the truth of an output, but it often improves performance on structured tasks, especially in mathematics, programming, and logic.
The problem is that this quantitative improvement does not automatically translate into qualitative reliability. A model may succeed much more often in math olympiads or coding competitions without thereby becoming capable of original research. Benchmarks measure skills on distributions of problems, often well formatted. An open conjecture, by contrast, is out of distribution by definition. It has no available solution in the data, at least in theory, and requires either a new idea or a novel combination of existing ideas. That is where OpenAI’s announcement, if substantiated, takes on particular significance.
There nevertheless remain several gray areas that will have to be clarified in order to assess the exact significance of the result:
- Which precise model was used, and with what settings?
- What was the role of the human researchers in formulating the problem, guiding the process, and verification?
- Is the proof entirely new, or does it rely on a reformulation of existing tools?
- Was the result submitted to peer review, to a detailed preprint, or to informal expert validation?
- Is the approach reproducible by other teams, on other problems?
These questions do not diminish the interest of the announcement; they define its real value. In AI’s recent history, many spectacular demonstrations have lost their shine once subjected to methodological scrutiny. Conversely, advances that seemed modest at first have ended up durably transforming a field because they were robust, reproducible, and usable by others.
For technology companies, the temptation is great to present every victory as proof of the system’s “generality.” But a validated contribution to a geometry conjecture does not imply that a model will tomorrow be able to propose physical theories, therapeutic molecules, or proofs in algebraic topology. The real signal would lie elsewhere: in demonstrating that a model can, within a well-defined framework, generate a mathematical idea solid enough to survive independent human expert review.
Why this announcement revives the debate over original research produced by AI
The heart of the debate is here: is this a case of genuine original scientific production by an AI, or another episode in which the human remains the main author and the model an advanced tool? The answer is not binary, and that is precisely what makes the matter interesting. In contemporary research, contributions are already distributed among humans, software, libraries, calculation assistants, formal proof systems, and simulation infrastructures. The arrival of reasoning models does not eliminate this chain; it enriches it with a new agent capable of proposing structures, analogies, and sometimes nontrivial leads.
The parallel with other announcements in the sector is illuminating. In 2023 and 2024, Google DeepMind communicated extensively about systems such as AlphaGeometry or AlphaProof, designed to solve geometry problems or olympiad-level statements. This work impressed the community, notably because it combined learning and symbolic search, and because it tackled tasks where formal rigor is unavoidable. But here again, the difference between solving competition problems and contributing to an open conjecture remained immense.
Anthropic, for its part, has emphasized the safety, interpretability, and reasoning capabilities of its Claude models, while Google integrated its progress into Gemini with demonstrations increasingly oriented toward code, science, and agents. xAI and some open-source teams also claim gains in reasoning thanks to specific training strategies. But most of these announcements still rest on evaluable tasks, not on open problems validated by peers.
What potentially sets the OpenAI episode apart, then, is less the raw performance than the nature of the test. A conjecture formulated in 1946 belongs to a very different timescale from that of benchmarks. It has resisted generations of mathematicians, increasingly sophisticated tools, and profound transformations of the discipline. If an AI brings a usefully novel solution to it, even with human assistance, that indicates that models can sometimes move beyond the simple regime of banal recombination.
Two symmetrical excesses must nevertheless be avoided. The first would consist in systematically minimizing any AI contribution on the grounds that a human verified or guided the process. That would misunderstand how science actually works, through interactions, corrections, and tooling. The second would be to proclaim that the machine “does science” in the same way as a researcher. That would ignore the essential dimensions of research: choosing problems, deep conceptual understanding, judgment about what is interesting, long-term intuition, and placement within a theoretical tradition.
The right reading is probably somewhere in between. If the external validation is solid, OpenAI has a credible case of AI-assisted scientific co-production. And that alone is enough to change the conversation. Until now, many scientific uses of LLMs have involved documentation, literature synthesis, coding help, or exploration of leads. A validated original mathematical proof would move AI into a more demanding category: that of tools capable, in certain circumstances, of contributing to the creation of new knowledge.
For mathematicians, the reaction will likely depend on the quality of the published material. A proof is not just a result; it is also a style of writing, an architecture, a set of lemmas, an economy of means, sometimes a vision. If the model produced a correct but opaque demonstration, difficult to generalize or interpret, the scientific interest will be real but limited. If, on the contrary, the solution reveals a fertile idea capable of opening other avenues, then the contribution will appear deeper. The history of mathematics is full of important results not only because they solve a problem, but because they invent a method.
This distinction is particularly important for industry. A model that solves an isolated problem thanks to an enormous expenditure of computation does not have the same value as a system that regularly helps researchers formulate new approaches. In other words, the real market question is not simply “Did OpenAI solve a conjecture?” but “Can this kind of success become systematic, reliable, and economically exploitable?”
The answer remains very open. The computing costs of cutting-edge models remain high. Their behavior remains nondeterministic. Their traceability is imperfect. And their use in sensitive scientific environments requires guarantees of confidentiality, reproducibility, and sometimes digital sovereignty that are not trivial, especially in Europe. Yet even with these limits, external mathematical validation gives OpenAI an argument that few players today can claim with such symbolic force.
What this changes, concretely, for the credibility of OpenAI’s reasoning models
The announcement does not make the criticisms directed at OpenAI’s models disappear, but it may change their hierarchy. Until now, the main reservation was simple: a model can appear brilliant while remaining fundamentally unreliable as soon as accuracy really matters. Hallucinations, calculation errors, invented references, or broken reasoning have extensively documented this problem. In this context, the credibility of a reasoning model cannot be based on verbal fluency alone, nor on its successes in public demonstrations.
External validation on an old mathematical problem does not erase that track record, but it brings a new element: it suggests that with the right protocols, the right guardrails, and a suitable domain, the model’s outputs can reach a level of reliability sufficient to interest experts. That is an important nuance. We are not moving from a fallible system to an oracle. We are moving from a system that is “often useful but intrinsically suspect” to a system “potentially capable of rigorous results when inserted into a serious verification process.”
For OpenAI, the stakes are strategic. Since the explosion of ChatGPT, the company has been trying to reposition itself beyond the mass-market conversational assistant. Its stated ambition concerns intellectual productivity, the automation of complex tasks, and, in the longer term, more general forms of decision and research assistance. In that narrative, reasoning models are essential. They must convince the market that they are not merely more talkative or slower to answer, but qualitatively better suited to handling difficult problems.
The validated mathematical case serves precisely that demonstration. It provides a more solid anchor point than benchmarks that are sometimes abstract for the general public and contested by specialists. It also allows OpenAI to respond indirectly to its competitors. Against Google DeepMind, which benefits from strong scientific credibility inherited from its work on AlphaFold, games, geometry, or optimization, OpenAI needs tangible proof that it too can produce results of high academic value. Against Anthropic, which emphasizes reasoning quality and safety, OpenAI can point to a concrete example of a validated result. Against open source, it can remind the market that the race is not only about model accessibility, but about the ability to achieve rare performance in extreme contexts.
But this added credibility remains conditional. It will depend on several factors:
- transparency about the protocol that led to the result;
- actual academic recognition of the proof;
- reproducibility of comparable approaches;
- the frequency of results of the same order in other fields;
- OpenAI’s ability to avoid overselling still-fragile successes.
On this last point, communication will be decisive. The AI industry has often erred through excessive promises, to the point of eroding the trust of researchers and companies. An announcement like this can restore credibility if it is accompanied by methodological humility. It can, on the contrary, revive skepticism if it is presented as proof that models “now do science” without further nuance.
For French-speaking players, this question of credibility is particularly sensitive. The European market is generally more cautious than the American market about adopting opaque technologies in critical uses. French companies exploring AI for R&D, engineering, or proof assistance will not settle for storytelling. They will ask for guarantees on data governance, auditability, regulatory compliance, and integration with formal verification or scientific computing tools already in place.
In that sense, the real significance of the announcement is not merely symbolic. If OpenAI manages to show that its models can be inserted into work chains where every important step is controlled, then its value proposition changes. The model is no longer an inspired text generator; it becomes a potential component of assisted research. That is a major shift, but one that rests less on the supposed magic of AI than on the quality of the validation procedures around it.
French-speaking market, European R&D, and the long-term perspective: the real issue is verification
Seen from France and Europe, the most important lesson of this matter may be less “AI solved an old problem” than “the value of scientific AI depends on its verification ecosystem.” This idea is central for public laboratories, universities, computing centers, engineering offices, and industrial companies engaged in research activities. A model output, however brilliant, has value only if it can be checked, documented, and integrated into a reproducible methodology.
In the European context, several trends reinforce this requirement. First, the rise of regulatory constraints and expectations around accountability. Second, growing sensitivity to questions of technological sovereignty. Finally, the existence of a very strong academic fabric in mathematics, theoretical computer science, and formal verification. France has specific strengths: engineering schools, leading laboratories, traditions in logic, optimization, and scientific computing. For these players, the interest of a model like OpenAI’s is not measured by its media aura, but by its ability to fit into proof tools, computing environments, and publication processes.
Concretely, if the announcement is confirmed, several uses could gain credibility:
- the exploration of conjectures in fields where AI can propose lemmas, counterexamples, or reformulations;
- proof assistance in formal systems such as Lean, Coq, or Isabelle;
- industrial R&D, notably for optimization, modeling, algorithm verification, or structure design;
- advanced training, with tools capable of suggesting richer proof avenues than current educational assistants;
- scientific monitoring, enriched by systems capable of linking distant results and proposing working hypotheses.
But these uses will not emerge automatically. They require investment in evaluation, in human-machine interfaces, and in formalization. They also require distinguishing between fields where error is acceptable and those where it is not. In exploratory research, an AI can be useful even if it is often wrong, provided it stimulates avenues of inquiry. In critical engineering, quantitative finance, or healthcare, the threshold of tolerance for error is obviously much lower.
In the long term, OpenAI’s announcement may therefore matter less for the particular problem it concerns than for the implicit standard it sets. If a company wants to convince the market that a model can contribute to science, it will no longer be enough to display scores or selected examples. It will have to show:
- a novel result;
- independent validation;
- a documented protocol;
- a clear articulation between human and machine;
- a possibility of reproduction or at least serious inspection.
If this standard takes hold, the scientific AI market could enter a more mature phase. Players would no longer be judged only on the size of their models or on the “wow” effect of their demonstrations, but on their ability to produce verifiable results. For the French-speaking market, that would be a rather healthy evolution. It would favor hybrid approaches combining generative models, formal tools, human expertise, and rigorous data governance.
It should also be noted that this dynamic could benefit European players. Europe does not dominate the race for very large generalist models, but it does have strong expertise in fields where verification, proof, and methodological rigor are central. If the future of scientific AI lies in systems better integrated into validation chains, then the continent’s laboratories, specialized software publishers, and deeptech startups have a card to play. Value could shift from a simple race for scale toward a race for operational reliability.
One fundamental question remains: will reasoning models progress steadily until they become commonplace scientific partners, or will these successes remain rare, costly, and difficult to generalize? AI’s recent history invites caution about linear extrapolations. Many impressive capabilities appear discontinuously, with strong dependencies on the domain, the problem format, and the level of human supervision. It is therefore possible that the validated resolution of an old conjecture is an important milestone without being the sign of an immediate shift toward “automated science.”
By contrast, it could herald something subtler and more durable: the emergence of a new research regime, in which models do not replace scientists, but become increasingly powerful instruments of conceptual exploration, provided they are embedded in strict control procedures. For OpenAI, the challenge is to turn a coup into structural credibility. For French and European researchers and companies, the challenge is not to confuse an isolated feat with industrial maturity, while recognizing that external mathematical validation, if it holds, may constitute one of the most serious signals to date that reasoning models are beginning to cross the boundary between sophisticated assistance and usable scientific contribution.
Comments· 2 comments
“Solved” is a strong word here. Could the article link the exact conjecture, the model’s full proof, and an independent verification statement from the mathematicians involved? In geometry, a proof can look persuasive while still relying on an unstated lemma or a gap in a special case.
Those are the right things to ask for. A useful standard would be a publicly readable proof with definitions and intermediate lemmas, plus comments or a publication-quality review from researchers who were not part of the original collaboration. If the claim is genuinely solid, that level of scrutiny should only strengthen it.