How Fast Can Reward Models Score? A Systems Study of C++ and PyTorch Inference Runtimes for RLHF
- 类型:arxiv
- 标识:2607.19712
- 链接:https://arxiv.org/abs/2607.19712
- 主分类:llm-infra
- 形态:method
- 被引:0
- 被引来源:Semantic Scholar + OpenAlex
- S2被引:0
- OpenAlex被引:0
- 影响力被引:0
- TLDR:A native C++ inference engine on ONNX Runtime that beat every baseline, confidence intervals didn't even overlap, and batching strategy mattered more than either the language or the runtime choice, more than the authors expected.
- OpenAlex ID:W7170171075
- OpenAlex DOI:10.48550/arxiv.2607.19712
- DOI:10.48550/arxiv.2607.19712
- DOI来源:OpenAlex
- 开放获取:green
- 开放获取链接:https://doi.org/10.48550/arxiv.2607.19712
- OpenAlex更新:2026-08-25
- 待LLM分类:否
- 标题中文:奖励模型能跑多快?RLHF 中 C++ 与 PyTorch 推理运行时的系统研究
- TLDR中文:基于 ONNX Runtime 的原生 C++ 推理引擎击败了所有 baseline,置信区间甚至不重叠;batch 策略的影响比语言或运行时的选择更关键,超出作者预期。
- 来源文件:
- /inbox/tom/_candidates/2026-07-30-agent-rag-longcontext-candidates.json
- [S2 enrich]
- [OpenAlex backfill]