Smaller Models, Better Rejects: Preference Distillation Scaling

  • 类型:arxiv
  • 标识:2609.38987
  • 链接:https://arxiv.org/abs/2609.38987
  • 主分类:llm-infra
  • 形态:method
  • TLDR:Preference distillation typically treats a teacher response as preferred and the student's own response as rejected. This assumes that self-generated failures are the most informative negatives and that rejects must come from a model at least as large as the student, making generation costly at scale. We find neither assumption holds: across students from 7B to 72B, smaller frozen models generate rejects with less inference compute yet train stronger students than self-generated rejects, before and after sequence-level knowledge distillation, on code generation and mathematical reasoning. To e
  • 待LLM分类:否
  • 来源文件:
  • /inbox/tom/_candidates/2026-10-02-agent-rag-longcontext-candidates.json