Smaller Models, Better Rejects: Preference Distillation Scaling
- 类型:arxiv
- 标识:2609.38987
- 链接:https://arxiv.org/abs/2609.38987
- 主分类:llm-infra
- 形态:method
- TLDR:Preference distillation typically treats a teacher response as preferred and the student's own response as rejected. This assumes that self-generated failures are the most informative negatives and that rejects must come from a model at least as large as the student, making generation costly at scale. We find neither assumption holds: across students from 7B to 72B, smaller frozen models generate rejects with less inference compute yet train stronger students than self-generated rejects, before and after sequence-level knowledge distillation, on code generation and mathematical reasoning. To e
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-10-02-agent-rag-longcontext-candidates.json