An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics
- 类型:arxiv
- 标识:2609.10712
- 链接:https://arxiv.org/abs/2609.10712
- 主分类:engineering
- 形态:method
- TLDR:We study how model post-training and test-time inference design affect natural-language proof generation for hard olympiad mathematics. Starting from Nemotron 3 Ultra, we train two specialist checkpoints using supervised fine-tuning and reinforcement learning, and evaluate checkpoint choice, verification, and refinement. Based on these findings, we present an open-model test-time-compute pipeline. The system operates entirely in natural language, with no formal prover, external tools, or internet access. Three Nemotron 3 Ultra checkpoints - the general-availability model and two post-trained s
- 待LLM分类:否
- 标题中文:通往 IMO 金牌的开放配方:面向奥林匹克数学的 Nemotron 训练
- TLDR中文:我们研究模型后训练与测试时推理设计如何影响面向困难奥林匹克数学的自然语言证明生成。从 Nemotron 3 Ultra 出发,我们使用监督微调与强化学习训练两个专家 checkpoint,并评估 checkpoint 选择、验证与修正。基于这些发现,我们提出一个开放模型的测试时计算流水线。该系统完全在自然语言中运行,无需形式化证明器、外部工具或网络访问。三个 Nemotron 3 Ultra checkpoint——通用发布模型与两个后训练专用模型……
- 来源文件:
- /inbox/tom/_candidates/2026-09-11-agent-rag-longcontext-candidates.json