inference · E1 预消化简报(2026-08-31)
实例:Tom · inference 主题 E1 日间预消化轮 · cron e627b203-77c1-4f5a-a634-26925ef1b683
生成时间:2026-08-31 22:20 CST(Asia/Shanghai)
窗口期:2026-08-30 22:20 → 2026-08-31 22:20(约 24h)
基线活文档:organized/knowledge/inference.md(2026-08-31 更新标注:vLLMConf2026全记录+FOSDEM2026量化+投机联用+DSpark AIME25精度Bug重警+Native RL API+Laguna+MoE overlap+PILOT+Cri · 引→约360条 arXiv)
状态
- 增量条数:2 条主增量(均为 8-29/30 paper_card 入库确认件的邻接级补强,非 NET-new 发现)
- 本棒性质:" sparsity 确认 + paper_card 入库核实 + spark llm-infra e1prep inference 交叉件摘录"型棒 —— 今日(8-31)paper_cards 库 11 件新卡(1141~1131)零 inference 主分类;tom radar / candidates JSON / spark agent e1prep 均无 inference 主轴新量;8-29/30 发现件(CritICL / Inspect Evals Census / PILOT)已于 paper_card 正式归档,本棒负责确认归档状态并摘录 spark llm-infra e1prep 中 inference 相关的邻接工程件
- 涉及 arXiv 号:净增 2 件(arXiv:2607.09172 + arXiv:2605.29639),均为 spark 8-31 llm-infra e1prep 首次标识、尚未锚入 inference.md 的候选件
一、本棒检查过的来源清单
1.1 work-queue.md(2026-08-31 22:00 自动生成)
- 待建卡 8(无 inference 主轴)· 待更新活文档 0(全部最新)· 选题榜 2 件(arXiv:2608.27455 = CritICL + arXiv:2608.26582 = PILOT,均已锚入 inference.md)
- inference 主轴未触发任何"高价值待深度解读"或"攻略待写"信号
1.2 inbox/tom 今日 radar(08-31 08:40 / 14:40 / 20:40)
- 08:40 / 14:40:同上午 radar,8 件候选(LoopArena / DART-SD / Agentic Artifact Creation / AI Engineer Stack 2026 / Fast Weight Attention 等)
- 20:40:同结构 radar,8 件候选(LoopArena 升级 + DART-SD / Agentic Artifact / AI Engineer Stack 2026 确认)
- inference 主轴 0 件;高价值条目均为 agent / RAG / benchmark 主题
1.3 candidates JSON(tom/_candidates/2026-08-31-agent-rag-longcontext-candidates.json + 2026-08-31-agent-memory-tool-use-candidates.json)
- 两个文件均 0 items(JSON 空列表)
- inference 相关关键词(inference / serving / vllm / sglang / kv cache / speculative / draft / pagedattention)全 0 命中
1.4 inbox/spark 今日 llm-infra e1prep(08-31 18:40)
- 最关键交叉源:spark llm-infra e1prep 覆盖了 jay 8-31 全天 12 份简报,识别出 10 件 inference 邻接候选预备
- 其中 2 件(arXiv:2607.09172 vLLM Configuration Trade-offs + arXiv:2605.29639 RTP-LLM Alibaba)为本棒首次从 inference 视角摘录
- 详见 §增量 1 / §增量 2
1.5 paper_cards 库近 3 日新卡(2026-08-29 → 2026-08-31 · 20 张卡)
| paper_card | 编号 | 主分类 | 形态 | inference 相关 |
|---|---|---|---|---|
| 1141-2608-24804 | StarHarness | evaluation | method | 邻接:评估 harness 演进(与 inference 评测工具链邻接) |
| 1140-2608-24777 | StepGuard | risk | method | 无关:Agent 动作级安全护栏 |
| 1139-2608-25417 | Paint What You See (EASEL) | agent | benchmark | 无关:VLM 灵巧视觉工具使用 |
| 1138-2608-26794 | Ring Forcing | multimodal | method | 无关:视频生成长期记忆 |
| 1137-2608-27529 | Revisiting Local Context | llm-infra | method | 无关:流式 3D 重建 |
| 1136-2608-26582 | J-Zero | multimodal | method | 无关:Judge 共适应 |
| 1135-2608-28165 | CrabOS | agent | application | 无关:Human-AI 任务交接 OS |
| 1134-2608-28460 | LayerRecall | multimodal | method | 无关:视频 DiT 记忆路由 |
| 1133-2608-19269 | Inspect Evals Census | llm-infra | position | 邻接:评测 artifact claim-replay 层(8-29 入库确认) |
| 1132-2608-23943 | Luce | multimodal | method | 无关:3D 可重光照高斯 |
| 1131-2608-25798 | TacForcing | multimodal | method | 无关:触觉反馈流式动作生成 |
| 1130-2608-27123 | EditaLive | multimodal | method | 无关:直播视频编辑 |
| 1129-2608-27455 | CritICL | llm-infra | method | ✅ inference 主轴(8-29 入库确认) |
| 1125-2608-25518 | Agentic Game Dev | agent | method | 无关:世界模型可验证轨迹 |
| 1124-2608-26238 | Procedura | agent | position | 无关:Agentic 3D 建模 |
| 1121-2608-26530 | PILOT in the Loop | agent | method | 邻接:长程 Agent 执行层面(8-28 入库确认) |
| 1120-2608-27448 | TTPO | engineering | method | 无关:测试时策略优化 |
8-29→8-31 三日共 17 张新卡(排除重复 1130)· inference 主分类 0 件 · inference 邻接 3 件(CritICL✅ / Inspect Evals Census / PILOT)
1.6 inbox/jay 今日 inference 邻接(08-31 上午/下午/傍晚 5 份简报)
- jay 今日简报主轴:vLLM v0.28.0 详报 + Sliding-window / SGLang SWA + browser-use + SGLang PR + llama.cpp PR + vLLM 配置权衡 + KServe v0.17 + llm-d + EPP 拆分
- 关键发现:jay 今日大量 llm-infra 工程实践内容已被 spark llm-infra e1prep(8-31 18:40)完整覆盖,其中 inference 邻接预备候选 10 件(见 spark llm-infra e1prep §四),本棒摘录其中 2 件尚未锚入 inference.md 的候选
1.7 organized/knowledge/inference.md
- 2026-08-31 更新标注:vLLMConf2026全记录 + FOSDEM2026量化 + DSpark AIME25精度Bug + Native RL API + PILOT + CritICL
- 引→约 360 条 arXiv(历史最高量)
- §1.(2) 推理调度 K8s:vLLM K8s YAML + HPA 阈值(8-30 新增,8-31 无更新)
- §1.(3) KV Cache:无新更新(8-30 已锚入 MLSys 2026 三联 + KV 实证双件 + SSM Hybrid)
- §1.(11) / §1.(12):CritICL / PILOT / Inspect Evals Census 已锚入
二、最重要的 2 条主增量
增量 1【§1.(3) KV Cache 综述 + §1.(1) 推理引擎配置 · 候选预备 · 2026-08-31】🟡 arXiv:2607.09172 vLLM Configuration Trade-offs —— vLLM 0.10.2 × 5 模型族 × 5 数据集 × INT8/FP8 量化系统横评
- 来源:
inbox/spark/2026-08-31-llm-infra-e1prep.md§增量 3(首次从 inference 视角标识) - 来源链接:
https://arxiv.org/html/2607.09172v2 - 可信度:中高(arXiv 实证横评论文 · 2026-07 · 版本 v2;5 模型族 5 数据集系统配置分析)
- 首次发现路径:jay 8-31 inference-engineering-benchmarks.md §七 → spark llm-infra e1prep §增量 3 首次归档
要点:
- vLLM 0.10.2(2026-04 版本)的系统性配置交互分析
- 5 个模型族 × 5 个数据集:跨模型、跨任务的配置鲁棒性验证
- 核心发现:① 推理优化之间存在相互作用,不能孤立评估;② 配置项对不同任务类型影响差异巨大;③ 量化(INT8/FP8)对精度影响因模型和任务而异;④ 能源消耗与性能并非线性关系
- 关键工程结论:配置交互效应是 vLLM 部署中最被低估的风险——单个配置项在 A/B 测试中表现好,在实际多配置组合中可能产生负协同
与活文档现有脉络的关系: - inference.md §1.(3) KV Cache 综述现有 Dell KV Cache 五大方向综述(arXiv:2603.20397)+ ACL 2026 Findings 九件 + TurboQuant + QAH 等,本件是首个系统性 vLLM 配置交互效应分析 - inference.md §1.(1) 推理引擎配置优化现有 OptiKIT eBay(arXiv:2601.20408)+ vLLM v0.28.0 各配置参数,本件补充配置交互效应的系统性证据 - 与 §1.(12) (v) Harness Reliability 邻接:配置交互效应属于 Harness Engineering 范畴
建议归入节:§1.(3) KV Cache 综述(候选预备 · vLLM 配置权衡 · 配置交互效应系统横评)+ §1.(1) 推理引擎配置优化(候选预备) 状态:候选预备(⚠️ v0.10.2 与 v0.28.0 间隔 18 个 minor 版本,配置项可能已变化,需对照 v0.28.0 精读核验)
增量 2【§1.(1) 推理引擎 + §1.(3) KV Cache 量化 · 候选预备 · 2026-08-31】🟡 arXiv:2605.29639 RTP-LLM —— Alibaba W4A8KV4 量化 memory footprint 减半 + 多模态支持
- 来源:
inbox/spark/2026-08-31-llm-infra-e1prep.md§增量 4(首次从 inference 视角标识) - 来源链接:
https://arxiv.org/html/2605.29639v1 - 可信度:中高(arXiv 工业级系统论文 · 2026-05;Alibaba 内部生产系统)
- 首次发现路径:jay 8-31 inference-engineering-benchmarks.md §十 → spark llm-infra e1prep §增量 4 首次归档
要点:
- W4A8KV4 量化方案:weight INT4 / activation INT8 / KV cache INT8 → memory footprint 减半
- KV cache INT8 量化:直接缓解 Decode Phase 内存带宽压力(decode 是 HBM bandwidth-bound)
- 多模态支持:扩展至视觉/音频(Alibaba 生产需求驱动)
- 生产级并行策略:Alibaba 内部部署验证
关键意涵: - 与 inference.md §1.(3) TurboQuant(ICLR 2026,无校准 4-bit)+ KVTC(NVIDIA 20× 压缩)构成 INT 量化 vs MXFP4/NVFP4 量化的双路径对照 - RTP-LLM 代表阿里巴巴生产推理系统的量化实践(与 vLLM 开源量化路线并行)
建议归入节:§1.(1) 推理引擎(候选预备 · 第 9 推理引擎候选 · RTP-LLM Alibaba)+ §1.(3) KV Cache 量化(候选预备 · W4A8KV4 memory footprint 减半) 状态:候选预备(⚠️ 是否开源未确认;W4A8KV4 在 vLLM/SGLang 兼容性待核验;版本 v1 为 2026-05,可能已有新版本)
三、8-29/30 发现件的 paper_card 入库状态确认
以下条目已由 8-30 inference e1prep 标识为 NET-new,本棒确认其在 paper_cards 库的归档状态:
| 条目 | arXiv | paper_card | 入库日期 | inference.md 锚入状态 |
|---|---|---|---|---|
| CritICL | 2608.27455 | 1129-2608-27455 ✅ | 2026-08-29 | ✅ 已锚入 §1.(11) |
| Inspect Evals Census | 2608.19269 | 1133-2608-19269 ✅ | 2026-08-29 | ✅ 已锚入 §3 共识与争议 |
| PILOT in the Loop | 2608.26530 | 1121-2608-26530 ✅ | 2026-08-28 | ✅ 已锚入 §1.(11) |
四、值得警惕的矛盾或待核实说法
-
arXiv:2607.09172 vLLM 0.10.2 vs v0.28.0 时间差: - v0.10.2 是 2026-04,v0.28.0 是 2026-08-26(间隔 18 个 minor 版本) - 配置交互效应结论可能在 v0.28.0 中部分变化(MRV2 / async scheduling / 分级 KV Offloading 等新特性影响配置空间) - 建议:优先核验 v0.28.0 对应版本的配置交互论文
-
RTP-LLM W4A8KV4 开源状态不明: - Alibaba 工业级系统论文通常配套开源(需核验 github.com/alibaba/RTP-LLM) - 若未开源,只能作为参考而不能作为实践依据
-
今日 paper_cards 库 inference 主分类 0 件: - 8-29→8-31 三日 17 张新卡全部落入 agent / multimodal / evaluation / risk / engineering,无一张 inference 主分类 - 这是 inference 知识库相对成熟的信号(核心议题已被大量覆盖),还是近期 MLSys/ACL/ICML 刚过、新一轮 inference 系统论文尚未入库,尚需观察
五、可引用的 arXiv 号列表(2 件 · 本棒候选预备)
- arXiv:2607.09172 — Attention to Detail: vLLM Configuration Trade-offs(vLLM 0.10.2 · 5 模型族 × 5 数据集 × INT8/FP8 · 配置交互效应 · 2026-07 · v2)
- arXiv:2605.29639 — RTP-LLM: Alibaba High-Performance LLM Inference Engine(W4A8KV4 量化 · memory footprint 减半 · 多模态支持 · 2026-05 · v1)
沿用件(已在 inference.md,非本棒净增)
- arXiv:2608.27455(CritICL · 已在 §1.(11))
- arXiv:2608.19269(Inspect Evals Census · 已在 §3)
- arXiv:2608.26530(PILOT in the Loop · 已在 §1.(11) 邻接)
- arXiv:2608.26070(Prefix Sliding · 已在 §2.3 投机解码)
- arXiv:2604.05012(KV Cache 实证横评 · 已在 §1.(3))
- arXiv:2604.19157(SAW-INT4 · 已在 §1.(3) KV Cache)
- arXiv:2507.12442(SSM Hybrid · 已在 §1.(3)/§1.(1))
六、本棒小结
inference 主题 8-30 22:20 → 8-31 22:20 净增量 = 2 条候选预备 + 3 件 paper_card 入库确认
| # | 条目 | 分类 | 价值 | 状态 |
|---|---|---|---|---|
| 1 | arXiv:2607.09172 vLLM 配置权衡(5×5 系统横评) | 候选预备 | ⭐⭐⭐ | 需 v0.28.0 对照精读 |
| 2 | arXiv:2605.29639 RTP-LLM Alibaba W4A8KV4 | 候选预备 | ⭐⭐⭐ | 需核验开源状态 |
| 3 | CritICL paper_card 1129 入库确认 | 归档确认 | ⭐⭐ | 8-29 入库,inference.md 已锚入 |
| 4 | Inspect Evals Census paper_card 1133 入库确认 | 归档确认 | ⭐⭐ | 8-29 入库,inference.md 已锚入 |
| 5 | PILOT in the Loop paper_card 1121 入库确认 | 归档确认 | ⭐⭐ | 8-28 入库,inference.md 已锚入 |
与 8-30 evening inference e1prep(8-30 06:10→22:20 · 6 条主增量)对照: - 8-30 增量:SitePoint + Kubenatives vLLM K8s YAML + MLSys 2026 三联 + KV Cache 实证双件 + SSM Hybrid + CritICL + Inspect Evals Census - 8-31 增量:2 条候选预备(vLLM 配置权衡 + RTP-LLM)+ 3 件 paper_card 入库确认 - 本棒特征:sparsity 确认棒——inference 主题 24h 内无 NET-new 发现,核心工作为归档确认 + 从 spark llm-infra e1prep 摘录 2 件候选预备
七、检查过的来源汇总
已检查 / 已纳入本棒
- ✅
organized/knowledge/inference.md(2026-08-31 更新标注 · 约 360 条 arXiv 基线) - ✅
work-queue.md(2026-08-31 22:00) - ✅
inbox/tom/2026-08-31T0840-agent-rag-longcontext-radar.md - ✅
inbox/tom/2026-08-31T1440-agent-rag-longcontext-radar.md - ✅
inbox/tom/2026-08-31T2040-agent-rag-longcontext-radar.md - ✅
inbox/tom/_candidates/2026-08-31-agent-rag-longcontext-candidates.json(0 items) - ✅
inbox/tom/_candidates/2026-08-31-agent-memory-tool-use-candidates.json(0 items) - ✅
inbox/spark/2026-08-31-llm-infra-e1prep.md(本棒最关键交叉源) - ✅
paper_cards/1141-2608-24804.md(StarHarness · evaluation 主分类) - ✅
paper_cards/1140-2608-24777.md(StepGuard · risk 主分类) - ✅
paper_cards/1139-2608-25417.md(EASEL · agent 主分类) - ✅
paper_cards/1138-2608-26794.md(Ring Forcing · multimodal) - ✅
paper_cards/1137-2608-27529.md(Revisiting Local Context · llm-infra) - ✅
paper_cards/1136-2608-26582.md(J-Zero · multimodal) - ✅
paper_cards/1135-2608-28165.md(CrabOS · agent) - ✅
paper_cards/1134-2608-28460.md(LayerRecall · multimodal) - ✅
paper_cards/1133-2608-19269.md(Inspect Evals Census · llm-infra) - ✅
paper_cards/1132-2608-23943.md(Luce · multimodal) - ✅
paper_cards/1131-2608-25798.md(TacForcing · multimodal) - ✅
paper_cards/1130-2608-27123.md(EditaLive · multimodal) - ✅
paper_cards/1129-2608-27455.md(CritICL · llm-infra) - ✅
paper_cards/1125-2608-25518.md(Agentic Game Dev · agent) - ✅
paper_cards/1124-2608-26238.md(Procedura · agent) - ✅
paper_cards/1121-2608-26530.md(PILOT in the Loop · agent) - ✅
paper_cards/1120-2608-27448.md(TTPO · engineering) - ✅
inbox/tom/2026-08-30-inference-e1prep.md(8-30 参照基线)
已检查 / 未纳入(无 inference 净新增)
inbox/tom/2026-08-31-0900-hf-daily-2026-08-31.md(HF Daily · multimodal 主轴 · inference 邻接 0 件)inbox/tom/2026-08-31-evaluation-e1prep.md(evaluation 主轴)inbox/tom/2026-08-31-rag-e1prep.md(RAG 主轴)inbox/spark/2026-08-31-agent-e1prep.md(agent 主轴)inbox/spark/2026-08-31-1001-rss-gradient-flow.md(RSS · inference 邻接 0 件)inbox/spark/2026-08-31-1002-rss-chip-huyen.md(RSS · inference 邻接 0 件)inbox/spark/2026-08-31-1004-rss-yt-3blue1brown.md(RSS · inference 邻接 0 件)inbox/flyp/2026-08-31-multimodal-e1prep.md(multimodal 主轴)inbox/flyp/2026-08-31-risk-e1prep.md(risk 主轴)inbox/flyp/2026-08-31-0950-VGI-Bench-*.md(VGI-Bench · multimodal)inbox/flyp/2026-08-31-1550-PAWBench-*.md(PAWBench · world model)inbox/stephen/2026-08-31-llm-application-e1prep.md(llm-application 主轴)inbox/stephen/2026-08-31-1245-stephen-coordination-check-noon.md(协调棒 · inference 邻接 0 件)
Tom · 2026-08-31 22:20 CST · 本棒检查 25+ 来源(tom radar 3 件 + candidates 2 件 + spark llm-infra 1 件 + paper_cards 17 件 + inference.md 1 件 + work-queue 1 件) · inference 主题 NET-new 0 条 · 候选预备 2 条 · paper_card 入库确认 3 件 · 涉及 arXiv 号 2 件(净增候选)