From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix

  • 类型:arxiv
  • 标识:2609.01572
  • 链接:https://arxiv.org/abs/2609.01572
  • 主分类:engineering
  • 形态:benchmark
  • TLDR:Data-residency constraints force enterprises to self-host LLMs, but continuous adoption of newer models without decommissioning their predecessors expands the serving fleet, fragmenting a finite GPU pool. We consolidate traffic from over 200 internal applications onto a single model by closing quality gaps identified through production error analysis along three axes: instruction following, function-calling, and internal task distribution. Quality is tracked by offline benchmarks stratified to production traffic and scored by deterministic verifiers or calibrated LLM judges. Rather than optimi
  • 待LLM分类:否
  • 标题中文:从生产流量到后训练:构建覆盖企业请求组合的自托管 LLM
  • TLDR中文:数据驻留约束迫使企业自托管 LLM,但不断引入新模型而不下线旧模型会扩张服务集群,分散有限的 GPU 池。我们通过沿指令遵循、函数调用和内部任务分布三个维度,针对生产错误分析所发现的质量差距,将 200 多个内部应用的流量整合到单一模型上。质量通过按生产流量分层的离线基准进行跟踪,并由确定性验证器或经过校准的 LLM 评判器打分。不同于针对
  • 来源文件
  • /inbox/tom/_candidates/2026-09-02-agent-rag-longcontext-candidates.json