
☰
AIMO 3 Winners Announced; AIMO Proof Pilot Launched
We are pleased to announce the five winning teams of AIMO 3: Exalted Joseph, varianceofx, SKobayak, TAMU-TACO and yemao ye. These are the #1, #2, #3, #5 and #7 ranked participants, respectively.
Ensuring that all submissions complied with this year’s stricter rules required additional time. This included submissions being released on HuggingFace, in order to more easily disseminate submissions so that others can learn from and build on them.
Congratulations to the AIMO 3 winners and thank you to everyone who contributed to making AIMO 3 such a productive and collaborative competition.
AIMO 3 analysis
This was an unusual competition for two reasons.
First, many submissions focused on a few strong shared notebooks. This was due in part to the exceptional, though hard-to-train, performance of the gpt-oss-120b model, and in part to other strong notebooks released early that shaped later work.
Second, the community produced a huge volume of new open-source artifacts. In particular, this year we added several additional prizes, including the best dataset (Math Corpus Prize) and the best write-ups about participants‘ models and experiments (Write-Up Prizes).
All entries for these additional prizes were public, helping generate approximately one million new data points and dozens of thoughtful papers, analyses, and resources for the broader community. Several submissions matched the quality expected of peer‑reviewed conference papers.
Launching AIMO Proof Pilot
We are delighted to announce a new evaluation: AIMO Proof Pilot.
AIMO Proof Pilot is an invite-only evaluation comprising seven teams: all six Extra Prize winners, plus Andreas Bisiadis, whose notebook served as the foundation for many other notebooks, including many top-performing ones. The community recognized his contribution by upvoting his notebook over 1,000 times, and many Kagglers declaring him the "real champion".
The goal of AIMO Proof Pilot is to explore how well open-source LLMs can produce mathematical proofs in a form similar to the International Mathematical Olympiad, rather than only providing final, numeric answers. This requires human evaluation of model outputs.
AIMO Proof Pilot also introduces another important experimental feature. It will allow only a small set of fully open-source models, for which intermediate checkpoints, training protocols, and training data are publicly available (e.g. models from the OLMo family).
While AIMO 3 allowed open-weight models that conformed to the rules, relatively few contestants were able to train gpt-oss-120b successfully. As a result, many submissions focused on building harnesses, prompt-engineering strategies, and other systems around it, contributing to the closely clustered leaderboard scores. Another limitation of running a competition with open-weight models is that it complicates clear post-mortem analysis. As a result, it is often impossible to determine why models perform well on one problem and not on another.
The evaluation will go live on 26 June 2026, when the outputs from the seven competing teams will be posted on Kaggle. The grading process will continue until 5 July 2026, with human graders from AIMO assessing the outputs, and the Kaggle Community providing supplementary feedback.
The seven participating teams were given one month to prepare for this evaluation. During this time, they were provided with compute through Fields Model grants. As part of this joint initiative, each team in the AIMO Proof Pilot has been given access to three GPU nodes (8x H200s each) by the Japanese National Institute of Informatics. This totals 168 H200 GPUs across the evaluation, plus engineering support for training and computation runs.
More information about AIMO Proof Pilot is on the evaluation's website: https://www.kaggle.com/competitions/ai-mathematical-olympiad-proof-pilot/overview