One family of models, two gold-level results

On October 7, 2026, NVIDIA's Nemotron team published a Hugging Face post reporting that models built from the Nemotron 3 family reached gold-medal-level scores at two of the hardest problem-solving contests in the world: the International Mathematical Olympiad (IMO), which demands rigorous written mathematical proofs, and the International Olympiad in Informatics (IOI), which demands algorithms and code that pass hidden tests under strict time limits.

The headline numbers are striking. At IOI 2026, a fine-tuned system called Nemotron-3-Ultra-CC scored 535.4 out of 600 in a live, prospective run held under the same time, internet-access, and submission constraints as human contestants. That is above the reported gold threshold of 361.12 points and above the reported top human score of 498.27. At IMO 2026, a generate-verify-refine system built from three Nemotron 3 Ultra checkpoints scored 30 out of 42 points, exceeding the official gold-medal threshold of 29, with the submitted proofs graded by official IMO graders.

These are not identical kinds of claims, and the distinction matters. The IMO score rests on official grading of submitted proofs. The IOI score comes from an unofficial run: real conditions, but not an official ranking. And both are results reported by NVIDIA about NVIDIA's own models. This article takes the team's own release as its primary source and separates what is independently checkable from what is not.

An unofficial run at IOI, officially graded proofs at IMO

The IOI run has an important caveat that the post states directly.

"It was an unofficial, unsupervised benchmark and was not included in the official IOI ranking."

NVIDIA Nemotron team authors, Hugging Face post "One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO," published October 7, 2026 (https://huggingface.co/blog/nvidia/nemotron-ioi-and-imo-2026). The post does not describe the run as an official medal. Treat 535.4/600 as a documented performance on that year's problem set under contest-like constraints, not as an entry in the official IOI standings.

The IMO result is stronger on the grading side: the proofs the system submitted were scored by official IMO graders, and 30/42 exceeded the official gold threshold. Even there, the system is a research artifact, not a contestant, so no medal was awarded to anyone. The useful framing is "gold-threshold performance," not "a gold medal."

The recipe behind both results

The most interesting claim in the release is not either score in isolation. It is the recipe. The team describes four steps that produced both specialists from the same base model family:

  1. Start with a strong Nemotron base model.
  1. Curate domain-specific problems and high-quality reasoning traces.
  1. Apply standard post-training methods such as supervised fine-tuning (SFT) and, where useful, reinforcement learning (RL).
  1. Pair the specialist model with an inference loop that generates, evaluates, and improves candidate answers.

For competitive programming, the team curated 22,000 problems and generated synthetic reasoning traces. They trained Nemotron-3-Nano-CC, a small model with 30 billion total parameters and 3 billion active parameters, with both SFT and RL, and Nemotron-3-Ultra-CC, a large model with 550 billion total parameters and 55 billion active parameters, with SFT alone.

The progression on the IOI 2025 problem set illustrates how the pieces add up. According to the post and the team's IOI paper on arXiv, the small Nano model scored 130 points before post-training, 280 after SFT, and 291 after RL. With GenCorrect, an iterative generate-evaluate-refine strategy, it reached 468 points, crossing that year's gold threshold of 438.3. The larger Ultra-CC model reached 502 with the same test-time strategy. These are NVIDIA's own reported training and evaluation numbers; independent parties have the code and data to try to reproduce them, but the reported figures themselves have not been rerun by a third party in the evidence reviewed for this article.

Teaching a model to prove, check, and revise

The IMO project started from Nemotron 3 Ultra and produced two specialists: one trained with SFT and one with RL. The SFT corpus contained 414,890 quality-filtered examples spanning 15,818 unique proof problems. The training data did more than teach final answers; it covered proof generation, refinement, verification, and meta-verification, so the model learned to construct arguments, find gaps, respond to critiques, and judge whether a proof was complete. The RL model was trained on 9,597 proof problems selected near the model's capability frontier.

In the development experiments, the two specialists had complementary strengths. The SFT checkpoint was strongest in the first search round; the RL checkpoint achieved the best overall single-checkpoint result. The final system used both specialists alongside the general-availability model. For each problem, the models generated candidate proofs, scored them, produced critiques, and refined the most promising attempts, with a separate high-compute stage selecting each final submission.

The team emphasizes that the whole system worked in natural language, with no formal prover, no external tools, and no internet access. A medal-winning IMO proof is an argument, not a calculation, so this constraint is what makes the IMO result meaningful as a measure of reasoning rather than of tool use.

Why fine-tuning and test-time compute had to work together

The post draws a conclusion that goes beyond either competition, and it is worth quoting exactly because it is the analytical heart of the release.

"The medals were not produced by fine-tuning alone, and they were not produced by brute-force sampling alone."

NVIDIA Nemotron team authors, same post, October 7, 2026. The team's claim is that the gains came from co-designing the model, the data, and the inference loop.

In plain terms: fine-tuning makes each candidate answer better, and test-time search makes a collection of imperfect candidates much better than any single one. At IOI, GenCorrect converted fine-tuning gains into larger improvements over multiple feedback rounds. At IMO, combining complementary SFT and RL checkpoints was more valuable than drawing more samples from one checkpoint. Neither trick alone produced gold-threshold scores.

A related observation from the paper is that adaptation does not look the same at every scale. For the small Nano model, SFT produced most of the gain and RL added a smaller, consistent improvement. For the much stronger Ultra model, one epoch of SFT was enough to outperform the fully post-trained Nano across IOI, ICPC, and LiveCodeBench Pro.

What was released, and what outsiders can now check

What makes this release unusual is how much is public. The Nemotron Labs IMO 2026 collection on Hugging Face brings together the SFT and RL checkpoints, both training datasets, and Nemotron-IMO-Bench, a new benchmark of 200 olympiad-level problems. The IMO paper on arXiv describes the training approach and the generate-verify-refine system. The NeMo-Skills repository on GitHub includes the IMO inference pipeline, the prompts, the submitted proofs, and a reproducible quickstart. The Ultra-CC model card for competitive programming is public, and the IOI paper covers its training recipe and the GenCorrect methodology, with the evaluation and inference pipeline available in NeMo-Skills.

That is what turns a vendor claim into something closer to shared knowledge. An outside party can, in principle, read the exact prompts, inspect the proofs that official graders scored, and rerun the inference pipeline. The open release also enables a new kind of benchmark: Nemotron-IMO-Bench gives the community 200 fresh olympiad-level problems on which any future system can be evaluated under the same conditions.

What the release does not establish

Honesty requires the counterweight. The open artifacts let outsiders verify the proofs, the prompts, and the inference pipeline. They do not let outsiders verify that the reported training runs were done exactly as described without redoing them, and the headline evaluation numbers, especially the IOI 2025 progression, are the vendor's own reporting. The IOI 2026 run was prospective and live but unofficial, so it cannot be compared against the official ranking. And one documented result about one model family does not establish that open weights generally beat closed systems; it establishes that this recipe, on this base model, on these problem sets, reached gold-threshold scores.

For a nontechnical reader, the deeper point is about the shape of achievement. A medal at the IMO or IOI has normally been the ceiling of human performance, produced by a handful of teenagers each year. Here, the same level of performance is a documented engineering recipe with public checkpoints, public data, and published code. Elite competence is becoming something you can reproduce and inspect rather than only something a closed laboratory claims. Whether and when that reproducibility extends to other demanding domains is an open question, and this release is one strong data point, not a proof of the trend.

The claim, hedged exactly as its authors hedged it

The IOI paper states its claim about the 2026 run with an explicit scope qualifier, which is exactly the right epistemic posture and is quoted here for precision.

"To our knowledge, this is the first AI system to outscore the highest-scoring human contestant on an IOI problem set."

Aleksander Ficek, Sean Narenthiran, Mehrzad Samadi, Somshubra Majumdar, and Boris Ginsburg, "Post-Training Language Models for Gold-Medal Performance in Coding Competitions," arXiv:2609.02849, abstract, v2 revised September 4, 2026 (https://arxiv.org/abs/2609.02849).

That qualifier, "to our knowledge," is doing real work. It is a claim about the authors' own survey, not a certified record. But even hedged, the claim marks a milestone: on one of the two most demanding problem-solving competitions in the world, a system built from openly published parts outscored, on that year's problem set, the best human contestant's reported score, in a run under human contest constraints. The mathematical side required no medal metaphor at all: 30/42 against an official gold threshold of 29, graded by official graders, in natural language only.

The lesson this article draws, and it is opinion rather than fact, is that the release is less about beating teenagers at contests and more about the maturity of a method. Strong base model, curated data, standard post-training, and a generate-verify-refine loop are all public, standard techniques. When their combination reaches the frontier of human competition, with the artifacts published for inspection, the frontier becomes a destination anyone with compute and patience can verify and aim for.