From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix
- 类型:arxiv
- 标识:2609.01572
- 链接:https://arxiv.org/abs/2609.01572
- 主分类:engineering
- 形态:benchmark
- TLDR:Data-residency constraints force enterprises to self-host LLMs, but continuous adoption of newer models without decommissioning their predecessors expands the serving fleet, fragmenting a finite GPU pool. We consolidate traffic from over 200 internal applications onto a single model by closing quality gaps identified through production error analysis along three axes: instruction following, function-calling, and internal task distribution. Quality is tracked by offline benchmarks stratified to production traffic and scored by deterministic verifiers or calibrated LLM judges. Rather than optimi
- 待LLM分类:否
- 标题中文:从生产流量到后训练:构建覆盖企业请求组合的自托管 LLM
- TLDR中文:数据驻留约束迫使企业自托管 LLM,但不断引入新模型而不下线旧模型会扩张服务集群,分散有限的 GPU 池。我们通过沿指令遵循、函数调用和内部任务分布三个维度,针对生产错误分析所发现的质量差距,将 200 多个内部应用的流量整合到单一模型上。质量通过按生产流量分层的离线基准进行跟踪,并由确定性验证器或经过校准的 LLM 评判器打分。不同于针对
- 来源文件:
- /inbox/tom/_candidates/2026-09-02-agent-rag-longcontext-candidates.json