Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling

  • 类型:arxiv
  • 标识:2609.19499
  • 链接:https://arxiv.org/abs/2609.19499
  • 主分类:llm-infra
  • 形态:method
  • TLDR:Test-time scaling can improve large language model reasoning by generating and combining multiple candidate responses. In sampling-based methods, the inference budget is often described by the number of generated candidates, N. However, N tells us how many candidates are generated, not how they are executed. The same candidate budget can be produced in one batched generation call or split across several sequential calls with smaller batch sizes. We first study the effect of increasing N on reasoning accuracy using Phi-3-mini and Qwen2.5-1.5B on 500 GSM8K prompts. As expected, increasing N from
  • 待LLM分类:否
  • 标题中文:样本数量远远不够:候选生成策略决定 LLM 测试时扩展的能耗与性能
  • TLDR中文:测试时扩展可通过生成并组合多个候选响应来提升大语言模型的推理能力。在基于采样的方法中,推理预算通常用生成的候选数量 N 来描述。然而 N 只能告诉我们生成了多少候选,而无法反映它们是如何被执行的。相同的候选预算可以通过一次批量生成调用完成,也可以拆分为若干次批量较小的顺序调用完成。我们首先基于 Phi-3-mini 和 Qwen2.5-1.5B 在 500 条 GSM8K 提示上研究增大 N 对推理准确率的影响。正如预期的那样,将 N 从
  • 来源文件
  • /inbox/tom/_candidates/2026-09-19-agent-rag-longcontext-candidates.json