Realtime AI News
NVIDIA Fine-Tunes Nemotron to Gold-Medal Level in Both IOI and IMO
NVIDIA says a recipe built on its Nemotron 3 base model, combining supervised fine-tuning, reinforcement learning, and feedback-driven inference, reached gold-medal level at both IMO 2026 and IOI 2026. The company also released the checkpoints, training data, benchmarks, and inference pipelines on Hugging Face.

NVIDIA researchers published a Hugging Face blog post on October 7 describing how they specialized the Nemotron 3 family with supervised fine-tuning (SFT), reinforcement learning (RL), and feedback-driven inference to build systems that reached gold-medal level at both the International Mathematical Olympiad (IMO 2026) and the International Olympiad in Informatics (IOI 2026).
The headline results: at IOI 2026, Nemotron-3-Ultra-CC combined with SFT and NVIDIA's GenCorrect strategy scored 535.4 out of 600, above the 361.12 gold threshold and above the top human score of 498.27. At IMO 2026, a generate-verify-refine system built from Nemotron 3 Ultra's general, SFT, and RL checkpoints scored 30 out of 42, above the official gold threshold of 29.
The company is explicit about the caveats. The IOI run was a live, prospective run under the same time, internet-access, and submission constraints as human contestants, but it was an unofficial, unsupervised benchmark and was not included in the official IOI ranking. The IMO system's proofs, by contrast, were graded by official IMO graders.
NVIDIA describes the underlying method as a reusable four-part specialization recipe: start from a strong Nemotron base; curate domain-specific problems and high-quality reasoning traces; apply standard post-training such as SFT and, where useful, RL; then pair the specialist with an inference loop that generates, evaluates, and improves candidate answers.
For competitive programming, the team curated 22,000 problems and generated synthetic reasoning traces to train two specialists. Nemotron-3-Nano-CC (30 billion total, 3 billion active parameters) received both SFT and RL, while Nemotron-3-Ultra-CC (550 billion total, 55 billion active) received SFT. On IOI 2025, Nano moved from 130 points before post-training to 280 after SFT and 291 after RL, reaching 468 with GenCorrect and crossing the 438.3 gold threshold.
For mathematics, the SFT corpus contained 414,890 quality-filtered examples across 15,818 unique proof problems, covering proof generation, refinement, verification, and meta-verification, while the RL model trained on 9,597 proof problems near the model's capability frontier. The final system worked entirely in natural language, with no formal prover, external tools, or internet access, scoring full credit on four of six problems.
The blog stresses that the medals came from co-designing the model, the data, and the inference loop rather than from fine-tuning alone or brute-force sampling alone: better specialization gives the inference system better candidates, better critics, and better refinements.
The artifacts are open on Hugging Face. The Nemotron Labs IMO 2026 collection bundles the SFT and RL checkpoints, both training datasets, and Nemotron-IMO-Bench, a new 200-problem olympiad benchmark; the IMO paper and the NeMo-Skills repository provide the inference pipeline, prompts, submitted proofs, and a reproducible quickstart. For competitive programming, the Nemotron-3-Ultra-CC model, the IOI paper, and the GenCorrect methodology are also available.
The broader signal is that easy to fine-tune is being redefined from the checkpoint can be trained to a capable base can be adapted to a demanding domain with a clear, reusable recipe. For the open community, shipping the model, data, and pipeline together means this gold-medal recipe can be reproduced and ported to other specialist domains.
Why it matters
Releasing the models, data, and pipelines together lowers the barrier to reproducing olympiad-grade specialist systems and reinforces the strong base plus specialization plus inference loop path for frontier reasoning.
Nearby Updates
All10/07, 20:45
Google Brings Gemini and Auto Browse to Chrome Users in India
Google has launched its Auto Browse feature in India and extended Gemini in Chrome to Android phones, letting the assistant perform multi-step tasks across websites on both desktop and mobile. Access is limited to Google AI Pro and AI Ultra subscribers for now.
10/07, 19:49
Tencent Reported Among Biggest Backers of DeepSeek's $12B+ Funding Round
A Music Business Worldwide report says Tencent is among the biggest backers of a DeepSeek funding round reported to top $12 billion. The investment follows Tencent Music's earlier integration of DeepSeek into its streaming service, tying China's largest internet group closer to a leading domestic AI lab.
10/07, 21:47
Sui Partners With Alibaba Cloud to Bring Cloud Services Into Sui Agent Payments
Sui has partnered with Alibaba Cloud to integrate cloud services into its Agent Payments system. The move links on-chain agent payments to conventional cloud computing, pointing toward a future where AI agents buy and settle for cloud resources on their own.
10/07, 19:40
Researchers Say an AI Agent Fleet Likely Linked to Tencent Scraped Rival Alibaba's Amap Maps Data
A preliminary report from the research group Swarmchasers says a fleet of AI agents likely running Tencent's Hunyuan (Hy) models spent more than a week pulling data from Alibaba's Amap mapping service, peaking at 1,810 URL query scans on October 4 alone. The report also found 211 scans labeled “claude,” but the researchers concluded the fleet was almost certainly not Claude, since the code matches Chinese models.