AI Aced the 2026 Math Olympiad: What This Means for AI Reasoning.
💡 In July and August 2026, multiple AI models - including Anthropic's Claude Opus 5 and two officially graded Chinese systems - scored a perfect 42 out of 42 on the International Mathematical Olympiad, the world's hardest math competition for high school students. The gold threshold is 29 points. AI can now solve every IMO problem. The real question is what that actually changes.
- Huawei's Celia and Xiaohongshu's dots-note-3.0 were officially graded by IMO organisers and each scored a perfect 42/42 - the first models formally awarded this through the official competition channel.
- Claude Opus 5 (Anthropic, released July 24, 2026, at $5 per million input tokens) also scored 42/42, with four independent no-tool solutions per problem confirmed by human experts.
- The human gold medal threshold is 29 out of 42. AI cleared it by 13 full points, on every problem, including Problems 3 and 6 that human gold medalists routinely leave incomplete.
- Honest caveat: only two models went through official IMO grading. Four others were evaluated using AI agents as graders - a weaker claim that many headlines did not distinguish.
- IMO problems are existing, known-format competition questions. Solving them is not the same as doing novel mathematics, and the benchmark is now saturated.

What Happened at the 2026 International Mathematical Olympiad?
The International Mathematical Olympiad is the oldest and most prestigious math competition for pre-university students. Six problems are set by an international jury each year, covering number theory, geometry, combinatorics, and algebra. Problems 3 and 6 are typically so hard that even the world's best young mathematicians commonly leave them blank or earn only partial credit.
At the 2026 competition, two AI systems were officially submitted for grading through the IMO's formal channel: Huawei's Celia and Xiaohongshu's dots-note-3.0. Both received a perfect 42 out of 42 points, solving all six problems with no human assistance. Separately, Anthropic's AI mathematical reasoning model Claude Opus 5 - released July 24, 2026 - was tested on the same problems by independent researchers and produced four independent proof-style solutions per problem, each confirmed correct by a panel of human experts. A venture capitalist also tested four frontier models including Claude Fable 5 and GPT-5.6 Sol; all scored 42/42 through AI-graded evaluation. The gold medal threshold for human competitors is 29 points. Every AI model cleared it by 13 full points.
Not All 42/42 Scores Are Equal
The headline hides an important distinction: the strength of any 42/42 claim depends on how the solutions were verified.
The two officially graded models - Celia and dots-note-3.0 - went through the same formal process as human competitors: solutions submitted to IMO organisers, graded without human intervention by the same team that grades student work. That is the highest standard of verification available.
The other models scored 42/42 through a different path: a researcher ran the problems, the models produced solutions, and other AI systems acted as judges. Meaningful, but not the same as official scrutiny. In at least one documented case, two AI models gave identical wrong answers to Problem 3, suggesting shared training data rather than genuine independent reasoning - a red flag for anyone assuming every 42/42 reflects infallible performance across all math tasks.
Cost per correct answer also varied considerably: from roughly $20 to $51 per problem across models, which matters for anyone thinking about practical deployment at scale.
Why This Is Still a Genuine Milestone
Even with those caveats, the achievement is real. IMO problems require multi-page proofs, careful logical chaining, and combinations of mathematical ideas well beyond rote computation. Human gold medalists train for years to reliably solve three or four of the six problems. The problems are only released to AI systems after the human competition ends, ruling out shortcuts from prior exposure.
Claude Opus 5 producing four independent solutions per problem - each verified by a human expert panel - is not a lucky guess. It shows that frontier AI mathematical reasoning can now engage with the full machinery of structured proof-writing: stating claims precisely, building arguments step by step, and arriving at a complete, checkable logical conclusion. This level of performance was not achievable two years ago.
What Does This Mean for You?
If you use AI as a math tutor, study partner, or problem-solver, the practical message is clear: on structured mathematical problems with verifiable answers, frontier AI is now operating at world-class level. That changes what you can reasonably ask of it:
- Checking your own solutions for logical errors, at the level of an expert human reviewer
- Explaining every step of a complex proof, not just providing the answer
- Generating multiple distinct approaches to the same problem - useful when understanding matters more than copying
- Working through advanced topics in calculus, number theory, combinatorics, or linear algebra at a level that would normally require a specialist
For researchers and educators, the implications are worth thinking through. Theorem-proving tools that use AI to verify or suggest proofs have become significantly more capable. And if AI can now solve every problem on the hardest student math test, the value of assessing mathematical ability purely through solving those types of problems needs to evolve.
For anyone working with language and technical content: the same reasoning capability that enables this level of structured language processing also underpins AI's growing ability to parse logical dependencies in technical documents, contracts, and patent claims. The architectural connection is real.
What Does AI Getting Perfect IMO Scores NOT Mean?
Solving the 2026 IMO is not the same as doing original mathematics. The problems, as hard as they are, are designed questions with correct answers in a known format, tested on a known population. Generating a new conjecture, finding a proof for an open problem, or recognizing which questions are worth asking in the first place: those are different skills, and AI performance there remains far more limited.
The benchmark is also now saturated. When multiple models all score 42/42, the number stops carrying information. It no longer distinguishes between good models and great ones, or tells you which model is right for your actual task.
The most important thing to internalize in 2026 is the jagged frontier: the same models that solve all six IMO problems have documented failures on much simpler tasks. According to a 2026 analysis, GPT-5.4 reads analog clock faces correctly only 50.1% of the time, roughly the same as guessing. A separate study found that 42 out of 53 frontier AI models gave wrong answers to a straightforward question about whether to walk or drive to a car wash 50 meters away. The models failed because the question required implicit spatial reasoning outside their training distribution.
The lesson is not that AI is overrated overall. It is that AI is very strong in domains where its training has good coverage, and surprisingly weak in edge cases it has not encountered before. The IMO lives in the first category. Everyday practical reasoning sometimes falls in the second. Understanding which is which, for your actual use case, matters more than any single benchmark score.
For a wider look at how AI handles the kind of nuanced language tasks that also require structured reasoning, recent research on the brain's language network offers a useful counterpoint: even as AI matches humans on structured problems, the underlying mechanisms remain different in ways that matter.
What to Watch Next
The frontier for AI reasoning has moved past the IMO. The new tests are harder and more open-ended:
- FrontierMath: original, unpublished mathematical problems designed by research mathematicians, where AI performance is still well below human expert level
- Open-ended proof discovery: can AI generate a proof for a genuinely unsolved problem, independently verifiable by the mathematical community?
- Reliability under novel formats: does the model's reasoning hold when the problem structure changes from what appears in training?
The IMO told us where the bar was in 2024. It no longer tells us where the frontier is in 2026. To understand what AI can actually do in mathematics, you need harder benchmarks, real research tasks, and honest failure-mode analysis, not the competition AI already aced.
FAQ
Did AI actually compete in the 2026 International Mathematical Olympiad?
Two AI systems - Huawei's Celia and Xiaohongshu's dots-note-3.0 - were officially submitted to IMO organisers for grading after the human competition ended and both received a perfect 42/42. Claude Opus 5 and other frontier models were tested independently on the same problems but did not go through the official competition channel. Both types of testing are meaningful, but they are not equivalent.
Does this mean AI is better at math than any human?
At solving structured, competition-style proof problems with clear correct answers, frontier AI models now match or exceed the world's best young mathematicians. That is not the same as being better at mathematics overall. Original research, problem formulation, and mathematical creativity are different skills, and AI performance there remains far more limited. The IMO tests a specific, important slice of mathematical ability, not the full picture.
Can I use AI to help me study mathematics now?
Yes, and this result makes that more credible. Frontier models can explain, check, and solve advanced mathematical problems at world-class level on structured problems. For learning, the most useful approach is not copying the answer but asking the model to explain each step, produce multiple approaches, or identify where your own reasoning went wrong. That is where AI mathematical reasoning adds genuine learning value.
Why does the verification method matter if the scores are all 42/42?
Because AI-graded-by-AI is not the same as graded by official human examiners. When two AI models gave identical wrong answers to Problem 3 in independent testing, it revealed a shared training-data blind spot, not independent reasoning. The official IMO grading process is the only standard that matches how human performance is measured, which is why the distinction between the two verified models and the four independently tested models matters.
What happens now that AI has saturated the IMO benchmark?
The IMO stops being useful as an AI evaluation tool - once every model scores 42/42, the number carries no information. Researchers are moving to harder tests: FrontierMath uses original unpublished problems by research mathematicians, and open-ended proof discovery asks whether AI can make progress on genuinely unsolved mathematics. Those are the benchmarks that will show where the real frontier is.
Sources: Anthropic - Claude Opus 5 release notes (2026); BreezyScroll - AI Scores Perfect 100% at International Mathematical Olympiad (2026); Digital Applied - Four AIs Scored a Perfect 42/42 on IMO 2026. So What? (2026)
About the author
Dao Huy (Lucas) is a professional translator working between English, Vietnamese, Chinese, and French, with more than seven years of experience. He follows advances in AI and mathematical reasoning not as a researcher, but as someone whose work depends on precision, logical structure, and detecting where an automated system is confident versus where it quietly fails - skills that matter as much in a verified proof as in a translated technical contract.
Lucas offers professional English-Vietnamese translation services, including technical, legal, and patent translation. If you need content where both accuracy and human judgment matter, he is happy to provide a quote at daohuy.com.
Written by Dao Huy (Lucas), Vietnamese translator & localization specialist (EN · ZH · FR → Vietnamese). See translation services →
