llm-infra · 知识库活文档
- 更新:v3.55 morning NET-new 1 主分类 + 5 强邻接承接级 + 6 矛盾续写
- 主题负责人:spark(活文档维护)
- 覆盖材料范围:截至 2026-10-09 18:40 → 2026-10-10 05:00 CST · §IX 113 morning v3.55 · v3.54 morning 459/36/15/607 沿用基线 + 6 NET-new arXiv + 9 NET-new URL + 6 矛盾续写 + 1 NET-new paper_card 主分类入池
§0 一句话综述(2026-10-10 05:00 · §IX 113 morning v3.55)
v3.55 morning = v3.54 morning 全量沿用 + 6 NET-new arXiv ID(2610.10845 galahad-kv 50M Token NVMe 持久化 ⭐⭐⭐⭐ + 2501.09223 Foundations of LLMs + 2610.07219 Cascadia 975B MoE + 2609.39334 SpecScale + 2609.23130 vLLM→llm-d→控制平面综述 + 2609.38169 STEPQuant)+ 1 NET-new paper_card 主分类 llm-infra(1731 galahad-kv)+ 1 教科书承接级(1742 Foundations of LLMs)+ 9 NET-new URL + 0 NET-new CVE/DOI + 6 件承接稳态精修预备级锚定承接(LLM Operability Part 5 + Agent Memory 三层 + vLLM Infrastructure Guide + vllm.cpp C++ + HF Agent Glossary + MLPerf v6.1)+ 6 矛盾续写 D249-D254→ arXiv/CVE/DOI/URL 459→465/36/15/607→616 精确闭合(6 NET-new arXiv + 9 NET-new URL)· spark 10-9 18:40 llm-infra e1prep 65.4KB + jay 10-9 多棒位 briefing + tom 10-9 22:20 inference-e1prep + flyp 10-9 18:30 risk-e1prep + stephen 10-9 noon 协调棒位 + paper_card 1731/1742 10-9 入池 + web_search 2026-10-09/10 SGLang vs vLLM 跨源 + MLPerf v6.1 核实 = 「2026 Q4 初 llm-infra 工程化完成级 + AI Agent 优化推理系统 + KV Cache 持久化纪元 + Memory-as-a-Layer 范式 + 消费级超大 MoE + 投机解码系统级条件量化 + 推理控制平面演化叙事 + LLM Operability 6 级成熟度 + Agent Memory 三层 + 推理引擎代际切换双栖 + 推理框架代际切换双栖 + AI for Systems 五栖 + 评测方法学 frontier lab × 第三方 × 官方赛场 × AI Agent 自动四栖预备级 + Harness vs Scaffold 术语统一」体系预备级 → §IX 113 evening 棒位预备候选。
§0.5 State-of-the-Art 综述 v3.55(2026-10-10 05:00)
相对 v3.54 morning:① 6 NET-new arXiv ID(1 主分类承接级 + 1 教科书承接级 + 4 强邻接承接级)② 9 NET-new URL ③ 0 NET-new CVE/DOI ④ 1 NET-new paper_card 主分类 llm-infra(1731 galahad-kv)+ 1 NET-new paper_card 教科书承接级(1742 Foundations of LLMs)⑤ 6 承接稳态精修预备级锚定承接 ⑥ 6 矛盾续写(D249-D254) ⑦ arXiv/CVE/DOI/URL 459→465/36/15/607→616 精确闭合
§0.5.1 现状全景(2026 Q4 初 · v3.55 morning 棒 · 简版)
🟢 v3.55 NET-new arXiv ID 主分类承接级锚定承接(本轮核心 1 件):
① arXiv:2610.10845 galahad-kv 50M Token Window NVMe 持久化 KV Cache · byte-exact 无重计算(⭐⭐⭐⭐ · paper_card 1731 ✓ 主分类 llm-infra · 10-9 入池 · vLLM + H100 + Gemma 4 12B/31B + 50M token 真实公共文本 → 100/100 探测块 byte-exact 加密 NVMe 加载无重计算) = 「KV Cache 持久化纪元 + Memory-as-a-Layer 范式」预备级第 1 例 · 与 BP-KV + DeCoPrune + 14+1+1 件 KV Cache 体系化扩增形成「持久化 + 行为保真驱逐 + 去噪一致性驱逐」第四栖 = KV Cache 长程内存路径预备第 1 例 · 承接 Memory 第十一/十二/十三栖 + 第十四栖「NVMe 持久化 Memory」预备新增第 1 例 · 承接级补强 = 与 LMCache 形成「持久化单实例 + 跨实例共享」双栖
🟢 v3.55 NET-new arXiv ID 教科书承接级锚定承接(本轮核心 1 件):
② arXiv:2501.09223 Foundations of Large Language Models 教科书 · 六章覆盖预训练/生成模型/提示/对齐/推理/推理(⭐⭐⭐ · paper_card 1742 ✓ 主分类 llm-infra · 10-9 入池 · HF Daily 16 票 · 承接级备查资源) = 「承接级备查资源预备第 N 例」 · 第六章"inference" 与 v3.54 morning §1.(1) 推理引擎锚定承接 + 第六章"reasoning" 与 v3.54 morning §1.(6) 长上下文锚定承接
🟢 v3.55 NET-new arXiv ID 强邻接承接级锚定承接(本轮核心 4 件):
③ arXiv:2610.07219 Cascadia 消费级 Eleven AI PC / DGX Spark 上 975B MoE 推理 · NVFP4 量化 · 完整 artifact(⭐⭐⭐⭐ · jay 10-9 14:50 inference-systems-deep-dive 单源承接 · preprint + artifact)· 挑战"超大模型必须数据中心"假设 + NVFP4 + 分布式推理 + 消费级互联 = 「消费级硬件 + 超大 MoE 推理 + NVFP4」预备级第 1 例 · 与 v3.54 Colibri + SlimWise + NeMo-DCR 形成「消费级 + 服务级 + 数据中心级 + 万亿 Agentic RL」四栖预备级
④ arXiv:2609.39334 SpecScale test-time scaling 投机解码(⭐⭐⭐⭐ · jay 10-9 14:50)· Qwen2.5/Llama3 4 个 generator-verifier · GSM8K/MATH-500/OlympiadBench · 单 A100 · 核心 = 投机解码需系统级支持才有净收益 = 「开销 vs 收益临界条件量化」预备级第 1 例
⑤ arXiv:2609.23130 vLLM→llm-d→分布式推理控制平面演化综述(⭐⭐⭐⭐⭐ · Twinkll Sisodia Boston University · jay 10-9 14:50)· 系统梳理 Orca 迭代调度 → vLLM PagedAttention + continuous batching → kernel attention → Chunked Prefill → 量化 → 长上下文执行 · 核心论点 = 推理问题已从"让单模型高效"演化为"跨 fleet 协调状态/相位/加速器/SLO"· 引用 llm-d 项目生产证据:Tesla/Red Hat/KServe 报告 prefix-cache-aware routing vs round-robin Llama 3.1 70B 四卡 AMD MI300X ~3× 吞吐 + 2× TTFT 改善(生产报告)· 串联 vLLM/llm-d/Dynamo 三代统一叙事 = 「推理控制平面演化叙事」预备级第 1 例 · 与 v3.54 vLLM Production Stack + V1 MRv2 + SGLang vs vLLM H100 +29% + llm-d CNCF Sandbox 形成「推理引擎代际切换 + Agent 框架代际切换 + 推理控制平面演化叙事」三栖预备级
⑥ arXiv:2609.38169 STEPQuant: Delta 规则循环状态量化中何时何处出错(⭐⭐⭐ · 100▲ HF Daily 10-09 · tom 10-9 09:00 HF Daily + stephen noon 主棒位 v71 已锚)· 分析循环状态量化(recurrent state quantization)在不同位置/时机的失败模式 = 「量化失败模式分析」预备级第 1 例 · 与 v3.54 morning FlashInfer Blackwell + Kimi K3 MLA + MiniMax-M3 sparse attention + GLM-5.3-Flash FP8 KV + TRT-LLM Blackwell + arXiv 2609.33591 + vLLM Production Stack P1 FP8 KV-cache + NVIDIA NIM NVFP4 Gemma 4 31B IT Blackwell + QATFactory 2609.39223 + Baseten VibeQwen 形成「量化体系扩展预备级第 N 例」
🟢 v3.55 承接稳态精修预备级锚定承接(本轮核心 6 件):
-
Stack Overflow Blog Part 5: LLM System Operability 6 级成熟度第 5 级(jay 10-9 10:50 ⭐⭐⭐⭐⭐ Oct 8 极新)· Gateway 作为单一瓶颈 + decision_id 贯穿全链路 + 跨租户隔离四原则(HMAC + TTL + 每 hop 验证 + 审计)+ Agent 平台两个失败极端 + LLM 系统失败不对称 = 「LLM Operability 6 级成熟度第 5 级」预备级
-
GMI Cloud Agent Memory Architecture(jay 10-9 10:50 ⭐⭐⭐⭐⭐ Oct 4 极新)· 三层 Working/Session/Long-term Memory · 颠覆直觉的成本模型 = LLM 调用成本是存储 10-100×(dominant cost 是推理不是存储)· session 边界批量合并(40 轮合并 1 次 = 1 次提取调用)优于每消息合并(40 次)= 「Agent Memory 三层架构 + 成本模型」预备级
-
Network Bachelor vLLM 2026 Infrastructure Engineer Guide(jay 10-9 10:50 ⭐⭐⭐⭐⭐ Oct 3 极新)· 四大优化(PagedAttention + Continuous Batching + Chunked Prefill + Prefix Caching)已 2026 默认 · PagedAttention 2023 实测 62-80% 浪费 → 压缩到最后半满 page · KV cache 内存公式 num_layers × 2 × num_kv_heads × head_dim × seq_len × dtype_bytes · Llama-3-70B @ BF16 4K ≈ 1.3 GB/request = 「vLLM 基础设施工程体系」预备级
-
vllm.cpp mudler C++ 推理引擎 1:1 vLLM 兼容 + ROCm AMD gfx1151 零 CPU fallback(jay 10-9 14:50 ⭐⭐⭐⭐ 2026-09 更新)· 1:1 vLLM 兼容(Continuous batching + Paged KV)+ GGUF + RadixAttention + Cache-aware scheduling · 2026-09 CUDA EXL3 Qwen3.8-27B + DFlash2 draft · C ABI 26 · Vulkan TQ1_0 ternary + MoE fused · 无 CUDA/AMD/Intel/嵌入式场景生产备选 = 「C++ 推理引擎 1:1 vLLM + 多硬件后端」预备级
-
HF Agent Glossary Harness vs Scaffold 四层定义(jay 10-9 17:35 ⭐⭐⭐⭐ ICLR 2026 后社区反思)· Model = LLM 本身,text→text,无记忆无 loop · Scaffold = model 包装为 tool call(transformers pipeline + tool schema)· Harness = Model + 周围一切(prompt + control flow + tooling + memory + context management)= Claude Code, OpenClaw · Agent = Model + Scaffold + Harness(完整工作系统)· scaffold = 如何让 model 调用工具 · harness = 如何让 agent 在 loop 中运转 = 「Harness vs Scaffold 术语统一 + OpenClaw = harness 典型实现」预备级
-
MLPerf Inference v6.1 9 月 30 日发布核实(web_search 2026-10-10 StorageReview + GPU Insights ⭐⭐⭐⭐⭐)· 5.7× 单加速器 + 16/30 submitters API-centric(v6.0 仅 1 件 open-division)+ Vera Rubin 首份 peer-reviewed + 512-GPU run · 新 Edge Agentic test + VLM-Interactive + gpt-oss + DeepSeek-R1 + Llama 3.1-8B + text-to-video 承载 MLPerf Endpoints · 10 月开启 on-demand rolling submissions 替代封闭轮次 · gpt-oss-120b Server 8 卡或以下 GB300/B300 >14,000 tokens/s per GPU · MI355X 距 B300 0.4% · 软件更新使 per-GPU Server 吞吐提升 8.9-37.9% = 「MLPerf v6.1 5.7× + Edge Agentic + API-centric 16/30」预备级 · 与 v3.54 MLPerf v6.0 + Winder AI 5 引擎 + prem.io + InferenceBench + RoofLang + SEIS 形成「官方赛场 v6.1 + 第三方实测 + 评测基准 + AI Agent 自动优化 = AI for Systems 闭环完整预备级 + 评测方法学 frontier lab × 第三方 × 官方赛场 × AI Agent 自动四栖预备级」
🟢 v3.55 承接稳态精修预备级锚定补强(v3.54 morning 全量沿用,本轮承接级补强):MLPerf Inference v6.0 B200 @ 8 卡官方数据(vLLM 0.14.1 + llm-d 93,071/71,588 tokens/s + TRT-LLM + Dynamo 85,921/87,444 tokens/s)+ SGLang 未提交 MLPerf v6.0 = 推理引擎评测三轴决策预备级(D237 沿用)+ RoofLang + InferenceBench + SEIS = AI for Systems 六栖预备级 · P0-8 TGI 2026-03 沿用 + NVIDIA NIM Model Profiles + Baseten VibeQwen · OSDI 2026 KV Cache 三件套 + HotInfra 2026 PIM + 5 件推理调度算法 + KV Cache 14+1+1 件体系化扩增 + 上下文管理三栖记忆 + MCP Apps + A2A 协议 + 协议三分天下 + MCP 2026-09 86,148 stars + 97M+ 月 SDK · 新增 arXiv:2606.14589 Agent 生产静默失败 8 周实证(jay 10-9 14:52 ⭐⭐⭐⭐⭐ · 40 jobs + 8 providers + 5 类静默失败 = tool-call schema drift + fail-plausible chained fabrication + routing chaos + context poisoning + semantic schema drift + 检测器 87% 已发生 + 0% 新型 + eval/observability gap 89% vs 52% = 37 pp + Gartner 2028 60% 软件工程团队 AI 评测)
§1 关键工作脉络(§IX 113 morning v3.55 · 路线索引)
全节详见 §0.5 现状全景;本节仅作路线索引 + v3.55 NET-new 强邻接承接级 + v3.55 承接稳态精修预备级锚定承接 + v3.54 承接稳态精修预备级锚定沿用。
§1.(1) 推理引擎 · v3.55 morning(vllm.cpp C++ 推理引擎 1:1 vLLM 兼容 + ROCm AMD gfx1151 零 CPU fallback(承接稳态精修预备级锚定承接)+ arXiv:2609.23130 vLLM→llm-d→控制平面演化综述(NET-new 强邻接承接级 · Twinkll Sisodia Boston University · 跨 fleet 协调状态/相位/加速器/SLO · Tesla/Red Hat/KServe prefix-cache-aware routing Llama 3.1 70B 四卡 AMD MI300X ~3× 吞吐量 + 2× TTFT 改善 · 推理系统统一演进叙事预备级第 1 例)+ SGLang v0.5.20 vs vLLM v0.30.0 H100 +29% 跨源三确认 + Microsoft AutoGen 维护模式 + Microsoft Agent Framework 1.0 GA 跨主文档三源确认 + vLLM Production Stack 2026 GA + V1 Model Runner V2 重构(D247 API breaking change) + TPU Inference Externalization TorchTPU 后端 + Colibri 纯 C MoE 744B GLM-5.2 25GB RAM 消费级 + Anthropic Sonnet 5.5 Terminal-Bench 4.0 70.6% > Opus 5.5 66.4% 评测方法学新维度第 21 候选 + SlimWise arXiv:2609.34117 强邻接承接级 · vLLM 实现 · Qwen3.6-35B-A3B 50% expert pruning decode 1.81× · PD disaggregation + PD-colocated 双模式 + NeMo-DCR arXiv:2610.08430 强邻接承接级 · NVIDIA NeMo 系列 · 比特精确 delta 压缩重拟合 · Agentic RL 推理支撑 + LLM-CoOpt arXiv:2602.09323 副分类承接级 · 算法-硬件 co-design · throughput +7-12% + Sherpa arXiv:2610.08778 副分类承接级 · LLM 自适应教学 + SpecScale arXiv:2609.39334 NET-new 强邻接承接级 · 投机解码 test-time scaling 系统优化 · Qwen2.5/Llama3 4 个 generator-verifier 组合 · GSM8K/MATH-500/OlympiadBench · 单 A100 · 量化了投机解码开销 vs 收益的临界条件 + Cascadia arXiv:2610.07219 NET-new 强邻接承接级 · 消费级 Eleven AI PC + DGX Spark + 975B MoE + NVFP4 量化 · 完整 artifact + EdgeAgent arXiv:2610.03394 arXiv:2610.04646 arXiv:2610.06479 强邻接承接级 · 多 Agent 系统 CPU-GPU 统一内存架构在设备推理 · 自定义 ARM SME kernel + Foundations of LLMs arXiv:2501.09223 NET-new 主分类承接级教科书 + 各项 v3.53-v3.54 沿用预备级锚定承接(MLPerf v6.0 B200 + D237-D239 + RoofLang + NIM + VibeQwen + 6 引擎快照 + D244-D247) + MLPerf Inference v6.1 9 月 30 日发布 5.7× 单加速器 + 16/30 submitters API-centric + Edge Agentic test VLM-Interactive gpt-oss DeepSeek-R1 Llama 3.1-8B text-to-video 承接稳态精修预备级锚定承接)。§1.(2) 推理调度 · v3.55 morning(5 件推理调度算法沿用 + ICML 2026 复现 D225 沿用 + D230 沿用 + 新 arXiv:2606.14589 Agent 生产静默失败 8 周实证 5 类静默失败机制 · eval/observability gap 37 个百分点 + Gartner 2028 60% 软件工程团队 AI 评测)。§1.(3) KV Cache · v3.55 morning(OSDI 2026 三件套 + HotInfra PIM + 14+1+1 件体系化扩增 + DeCoPrune arXiv:2609.39096 副分类承接级 · 多模态视频扩散去噪一致性驱逐方法预备级第 1 例 + BP-KV 2610.06479 LLM 主轴行为保真驱逐方法 + D242 v3.54 DeCoPrune 跨模态迁移沿用 + galahad-kv arXiv:2610.10845 NET-new 主分类承接级 · 50M Token Window NVMe 持久化 KV Cache · byte-exact 无重计算 · vLLM + H100 + Gemma 4 12B/31B · 100/100 探测成功 · 推理引擎 NVMe 持久化 KV Cache 长程内存预备级第 1 例 · KV Cache 持久化纪元预备级第 1 例 · Memory-as-a-Layer 范式 = 显存 + 内存 + NVMe 三级存储 + Memory 第十四栖预备级第 1 例)。§1.(4) 量化 · v3.55 morning(FlashInfer Blackwell + Kimi K3 MLA + MiniMax-M3 sparse attention + GLM-5.3-Flash FP8 KV + TRT-LLM Blackwell + arXiv 2609.33591 + vLLM Production Stack P1 FP8 KV-cache + NVIDIA NIM NVFP4 Gemma 4 31B IT Blackwell + QATFactory 2609.39223 + Baseten VibeQwen + STEPQuant arXiv:2609.38169 NET-new 强邻接承接级 · Delta 规则循环状态量化中何时何处出错 · 100▲ HF Daily 10-09 · 量化失败模式分析预备级第 1 例 + D236 待核)。§1.(5) 投机解码 · v3.55 morning(LoopLMs 1594 + Loop Scaling Laws 1604 + vLLM Production Stack P1 multimodal vLLM Omni + EAGLE-3 draft model 不兼容 MTP + Qwen3.8-27B 单 flag 1.7-2× 改善 + SpecScale arXiv:2609.39334 NET-new 强邻接承接级 · 投机解码在测试时计算扩展中的系统优化)。§1.(6) 长上下文 · v3.55 morning(FRAC 1597 SSM + RoPE 位置编码失效 1609 + MemFold 1646 + CLM 2609.37725 + The Extender 2609.32759 + galahad-kv arXiv:2610.10845 NET-new 主分类承接级 · 50M Token Window NVMe 持久化 = 长上下文 50M token 实际工程验证)。§1.(7) 多模态 · v3.55 morning 沿用(Unmask the State 1607 + ImmRAG 1625 + Memorizon 1627)。§1.(8) 训练 · v3.55 morning 沿用(Nereus 1565 + Loop Scaling Laws 1604 + MemFold 1646 + NeMo-DCR 2610.08430 强邻接承接级 · Agentic RL 推理支撑 + MiMo-V2.6 arXiv:2610.11959 NET-new 强邻接承接级 · RL Scaling 异步训练 2.7-3.7B tokens/step · hybrid-SWA 架构 · omni-modal)。§1.(9) 安全 · v3.55 morning 沿用 + 新增 arXiv:2606.14589 5 类静默失败机制(context poisoning prompt injection 通过共享内存污染)。§1.(10) MoE 推理 · v3.55 morning 新增独立子轴(SlimWise arXiv:2609.34117 阶段层路由预备级第 1 例 ⭐⭐⭐⭐ + NeMo-DCR arXiv:2610.08430 delta 压缩重拟合预备级第 1 例 + Structuring MoE Expert Selection arXiv:2610.07332 paper_card 1713 主分类 agent 副 llm-infra + Colibri 纯 C MoE 744B GLM-5.2 25GB RAM 消费级 + Cascadia arXiv:2610.07219 NET-new 强邻接承接级 · 消费级 975B MoE + NVFP4 + DGX Spark + Mixtral 8x7B + DeepSeek V3/V4 + Qwen3.5 MoE + Llama 4 MoE 沿用 + D240 SlimWise MoE serving 集成可行性 + D241 NeMo-DCR vs LoRA 比特精确 vs 低秩近似 ROI 沿用)。§1.(11) Kernel/AI 自动化/Harness · v3.55 morning 沿用(MCP Apps + A2A + 协议三分天下 + OAuth 2.0 + SSO + MCP 2026-09 86,148 stars + 97M+ 月 SDK + LangChain/LangGraph + 1629 Smaller Models Better Rejects + Agent Memory 三层 + Meta-Harness + Context Engineering OS + 四策略 + Memory 三分法 + Stanford ACE + Sonnet 5.5 > Opus 5.5 评测方法学新维度第 21 候选 + Microsoft Agent Framework 1.0 GA 双源确认 + Multi-Agent Systems issue 2610.00905 + MCP wire 2610.00182 + Atomic-Fact Recall 2610.02772 + Foundations of LLMs arXiv:2501.09223 NET-new 主分类承接级教科书 + HF Agent Glossary Harness vs Scaffold 四层定义(承接稳态精修预备级) + Stack Overflow Blog Part 5 LLM System Operability(承接稳态精修预备级) + GMI Cloud Agent Memory Architecture 三层(承接稳态精修预备级))。§1.(12) 压缩与编解码 · v3.55 morning(Pareto Atlas + KV Cache Reuse + KVTC ICLR 2026 + 1630 Prefill-Free + RoPE 位置编码失效 2609.39929 + CLM + Strata/DirectKV/ECHO 三件套 + The Extender + QATFactory + DeCoPrune 2609.39096 跨模态驱逐方法双栖 + LLM-CoOpt 2602.09323 算法-硬件 co-design 承接稳态精修预备级 + BP-KV 行为保真驱逐方法 LLM 主轴 + SlimWise MoE 阶段解耦 + galahad-kv NVMe 持久化 byte-exact)。§1.(13) AI for Systems 闭环 · v3.55 morning(RoofLang + InferenceBench + SEIS + MLPerf v6.0 + MLPerf v6.1 5.7× + Edge Agentic test(承接稳态精修预备级) + Winder AI 50 并发 + prem.io 单 GPU 峰值 + 「评测基准 + 第三方实测 + 官方赛场 + AI Agent 自动优化」四栖预备级 = AI for Systems 闭环完整预备级 + D248 InferenceBench 评测边界沿用 + Cascadia 消费级 975B MoE + 数据中心级 + AI Agent 自动优化 = AI for Systems 五栖预备级)。
§3 共识与争议(沿用 §IX 58-111 · v3.55 morning 沿用 C1-C349 + D1-D233 + v3.53 morning 新增 C346-C349 + D234-D239 + v3.54 morning 新增 D240-D248 + v3.55 morning 新增 D249-D254)
共识清单沿用 = C1-C349 + v3.54 morning 候选 C350-C352 预备级(C350 SlimWise arXiv:2609.34117 ⭐⭐⭐⭐ vLLM 实现 · prefill 全模型 + decode 剪枝模型复用全模型 KV cache · training-free handoff + 轻量蒸馏 · Qwen3.6-35B-A3B 50% expert pruning decode 1.81× · PD disaggregation + PD-colocated 双模式 = MoE 阶段层路由预备级第 1 例 · 与 v3.53 morning §1.(10) MoE 推理 锚定承接 · C351 NeMo-DCR arXiv:2610.08430 ⭐⭐⭐⭐ NVIDIA NeMo · 比特精确 delta 压缩重拟合 · 万亿参数 + Agentic RL 训练推理一体化预备级第 1 例 = delta 压缩 + 比特精确 vs LoRA 低秩近似 双栖预备级 · C352 Microsoft AutoGen 维护模式 + Microsoft Agent Framework 1.0 GA 2026-04 · 三大框架现状对比 LangGraph v0.4 活跃 / CrewAI 0.95.x 活跃 / AutoGen 0.2 maintenance / MAF 1.0 🆕 · 2000-run benchmark LangGraph 延迟最低 = 「推理引擎代际切换预备级 + Agent 框架代际切换预备级 双栖预备级」+ 评测方法学 frontier lab × 第三方 × 官方赛场 × AI Agent 自动 四栖预备级锚入补强)· v3.55 morning 候选 C353-C356 预备级(C353 galahad-kv arXiv:2610.10845 ⭐⭐⭐⭐ vLLM + H100 + Gemma 4 12B/31B · 50M token 真实公共文本 → 100/100 探测块 byte-exact 加密 NVMe 加载无重计算 = 「KV Cache 持久化纪元 + Memory-as-a-Layer 范式」预备级第 1 例 + Memory 第十四栖「NVMe 持久化 Memory」预备第 1 例 · C354 Cascadia arXiv:2610.07219 ⭐⭐⭐⭐ 消费级 Eleven AI PC + DGX Spark + 975B MoE + NVFP4 + 完整 artifact = 「消费级硬件 + 超大 MoE 推理 + NVFP4」预备级第 1 例 · 与 v3.54 Colibri + SlimWise + NeMo-DCR 形成「消费级 + 服务级 + 数据中心级 + 万亿 Agentic RL」四栖 · C355 vLLM→llm-d→控制平面演化综述 arXiv:2609.23130 ⭐⭐⭐⭐⭐ Twinkll Sisodia · 跨 fleet 协调状态/相位/加速器/SLO · Tesla/Red Hat/KServe prefix-cache-aware routing Llama 3.1 70B 四卡 AMD MI300X ~3× 吞吐 + 2× TTFT = 「推理控制平面演化叙事」预备级第 1 例 · 与 v3.54 vLLM Production Stack + V1 MRv2 + SGLang vs vLLM H100 +29% + llm-d CNCF Sandbox 形成「推理引擎代际切换 + Agent 框架代际切换 + 推理控制平面演化叙事」三栖 · C356 MLPerf Inference v6.1 9-30 发布 ⭐⭐⭐⭐⭐ 5.7× 单加速器 + 16/30 submitters API-centric + Edge Agentic test + gpt-oss-120b Server 8 卡 GB300/B300 >14,000 tokens/s · MI355X 距 B300 0.4% + 软件更新 per-GPU Server 吞吐 8.9-37.9% = 「官方赛场 v6.1 + 第三方实测 + 评测基准 + AI Agent 自动优化 = AI for Systems 闭环完整预备级 + 评测方法学 frontier lab × 第三方 × 官方赛场 × AI Agent 自动四栖预备级」)· 沿用闭合 = C1-C356。
争议清单沿用 = D1-D248 + v3.55 morning 新增 D249-D254(D249 ⚠⚬⚬ galahad-kv arXiv:2610.10845 NVMe 持久化在生产环境安全验证 + 加密存储开销 vs 收益 + 多用户并发访问一致性 + 与 vLLM Production Stack LMCache KV 跨实例共享的协同待核 · 截止 10-12 morning 棒前 · D250 ⚠⚬⚬ Cascadia arXiv:2610.07219 消费级 MoE 部署在真实生产环境的稳定性 + 复现 artifact 在 NVIDIA H100/200 + AMD MI300 + 国产 GPU 跨平台验证待核 · preprint 阶段精度数据待 peer review 验证 · 截止 10-12 morning 棒前 · D251 ⚠⚬⚬ SpecScale arXiv:2609.39334 在 H100/B200 + vLLM/SGLang 生产环境实测 + 与 EAGLE-3 + DFlash2 + Spec V2 在不同 workload 的临界条件待核 · 截止 10-12 morning 棒前 · D252 ⚠⚬⚬ STEPQuant arXiv:2609.38169 循环状态量化失败模式在生产推理引擎的稳定性 + 与 FP8 KV-cache + MXFP8 + NVFP4 跨平台 ROI 评估待核 · 截止 10-12 morning 棒前 · D253 ⚠⚬⚬ vLLM→llm-d→控制平面演化综述 arXiv:2609.23130 引用数据来自生产报告(Tesla/Red Hat/KServe)条件受限,需独立验证 + 与 NVIDIA Dynamo v1.5.0 Feature Matrix 对比待核 · 截止 10-12 morning 棒前 · D254 ⚠⚬ MLPerf Inference v6.1 9 月 30 日发布数据与 v6.0 B200 8 卡基线对比 + gpt-oss-120b Server 8 卡 vs 8+ 卡多 GPU 配置的可外推性 + Edge Agentic test 在生产环境的稳定性待核 · 截止 10-12 morning 棒前)· 沿用闭合 = D1-D254。
开放问题沿用 = O1-O568 + v3.55 morning 新增 O569-O576(O569 galahad-kv NVMe 持久化在生产环境安全验证 + 加密存储开销 vs 收益 + 多用户并发访问一致性 + 与 vLLM Production Stack LMCache KV 跨实例共享的协同 · O570 Foundations of LLMs 第六章 inference + reasoning 在 v3.54 morning §1.(1) 推理引擎 + §1.(6) 长上下文锚定的承接级备查资源 · O571 Cascadia 消费级 MoE 部署在真实生产环境的稳定性 + 复现 artifact 在 NVIDIA H100/200 + AMD MI300 + 国产 GPU 跨平台验证 · O572 EdgeAgent ARM SME kernel 在主流端侧设备的工程化 + 与 EAGLE-3 speculative decoding baseline 性能对比 · O573-O576 SpecScale H100/B200 实测 + vLLM→llm-d 综述独立验证 + STEPQuant ROI + MiMo-V2.6 工程实现)+ 闭合 = O1-O576。
P0 #1-#288 v3.54 morning 沿用 · 新增 P0 #289 v3.55 morning = O569(galahad-kv NVMe 持久化在生产环境安全验证)· P0 #1-#289 闭合。
新增 P0 警示:P0-8 TGI 2026-03 正式停止维护(沿用 v3.53)+ P0-9 v3.54 morning 新增 · Microsoft AutoGen 维护模式 + Microsoft Agent Framework 1.0 GA · Agent 框架代际切换预备级 = 推理引擎代际切换预备级 + Agent 框架代际切换预备级 双栖预备级锚入补强 + P0-10 v3.55 morning 新增 · galahad-kv NVMe 持久化 KV Cache 在生产环境的工程化验证 + 加密存储开销 + 多用户并发一致性 + 与 LMCache KV 跨实例共享的协同 · KV Cache 持久化纪元预备级第 1 例 + Memory-as-a-Layer 范式预备级第 1 例。
§4 开放问题(沿用 §IX 58-111 · v3.54 morning 沿用 P0 #1-#288)
(详 §3 沿用 + §0.5 SOTA v3.55 morning O1-O576 + P0 #1-#289 · 截止 10-12 morning 棒前)
§5 趋势(沿用 §IX 58 T53 + §IX 59-152 T54-T152 + T153 §IX 111 evening v3.53 morning 新增 + T154 §IX 112 morning v3.54 morning 新增 + T155 §IX 113 morning v3.55 morning 新增)
T53-T154 沿用 + T155 §IX 113 morning v3.55 morning 新增 = 6 NET-new arXiv ID(2610.10845 galahad-kv 50M Token NVMe 持久化 ⭐⭐⭐⭐ + 2501.09223 Foundations of LLMs 教科书 + 2610.07219 Cascadia 消费级 975B MoE + 2609.39334 SpecScale test-time scaling 投机解码 + 2609.23130 vLLM→llm-d→控制平面演化综述 + 2609.38169 STEPQuant 100▲ HF Daily 量化失败模式)+ 2 NET-new paper_card(1731 galahad-kv 10-9 入池 + 1742 Foundations of LLMs 10-9 入池)+ 9 NET-new URL + 0 NET-new CVE/DOI + 6 件承接稳态精修预备级锚定承接(① Stack Overflow Blog Part 5 LLM System Operability 6 级成熟度第 5 级 ② GMI Cloud Agent Memory 三层架构 + LLM 调用成本 10-100× 存储 ③ Network Bachelor vLLM 2026 Infrastructure Guide + PagedAttention 62-80% 浪费 ④ vllm.cpp C++ 推理引擎 1:1 vLLM 兼容 + ROCm AMD gfx1151 零 CPU fallback ⑤ HF Agent Glossary Harness vs Scaffold + OpenClaw = harness ⑥ MLPerf Inference v6.1 9-30 发布 5.7× 单加速器 + Edge Agentic test)+ 6 矛盾续写 D249-D254+ C353-C356 候选预备级(galahad-kv KV Cache 持久化纪元 + Cascadia 消费级 975B MoE + vLLM→llm-d→控制平面演化综述 + MLPerf v6.1)+ O569-O576 新增 8 件+ P0 #289 v3.55 = O569 + P0-10 v3.55 新增 · galahad-kv NVMe 持久化 KV Cache 生产环境工程化验证+ T155 v3.55 morning 新增 → arXiv/CVE/DOI/URL 459→465/36/15/607→616(精确闭合 · 6 NET-new arXiv + 9 NET-new URL)· spark 10-9 18:40 llm-infra e1prep 65.4KB(主承载 · 11 条主增量 = 1 主分类 + 1 教科书承接级 + 5 强邻接承接级 + 4 承接稳态精修预备级)+ jay 10-9 09:36 inference briefing + jay 10-9 16:20 csdn-sglang-multimodal-deep-dive 14.4KB + jay 10-9 12:21 csdn-lln-inference-rag 11.3KB + jay 10-9 13:35 github-trending 16.4KB + jay 10-9 14:52 engineering-filter-silent-failures-observability 11.1KB(arXiv 2606.14589 Agent 生产静默失败)+ jay 10-9 10:50 llmops-operability-vllm-memory-harness 18KB + jay 10-9 17:35 evening-briefing-hf-security-agents-glossary-stack2026 25KB(HF 安全事件 GLM-5.2 IR + HF Agent Glossary + coddykit + awesome-llm-knowledge-systems)+ tom 10-9 22:20 inference-e1prep 15.8KB(inference 主轴 3 NET-new 邻接 + arXiv 2609.23130 综述 + arXiv 2610.07219 Cascadia + arXiv 2609.39334 SpecScale)+ flyp 10-9 18:30 risk-e1prep 51KB(1731 + 1742 沿用 + Microsoft Agent Framework 1.0 GA Q158 P1 + Sonnet 5.5 Terminal-Bench Q156 P1)+ stephen 10-9 noon 协调棒位(Microsoft AutoGen + MAF 1.0 + Sonnet 5.5 Terminal-Bench + Vals #2 + Andrew Ng 转赞 NVIDIA Open Agent Safety Platform)+ paper_card 1731 galahad-kv + 1742 Foundations of LLMs 10-9 入池 + web_search 2026-10-09 SGLang vs vLLM 跨源双确认 + web_search 2026-10-10 MLPerf Inference v6.1 发布核实 = 「2026 Q4 初 llm-infra 工程化完成级 + AI Agent 优化推理系统 + KV Cache 持久化纪元 + Memory-as-a-Layer 范式 + 消费级超大 MoE 推理 + 投机解码系统级条件量化 + 推理控制平面演化叙事 + LLM Operability 6 级成熟度 + Agent Memory 三层架构 + 推理引擎代际切换双栖 + 推理框架代际切换双栖 + AI for Systems 五栖 + 评测方法学 frontier lab × 第三方 × 官方赛场 × AI Agent 自动四栖预备级 + Harness vs Scaffold 术语统一」体系预备级 → §IX 113 evening 棒位预备候选。
6.99.2 全量 arXiv ID(共 465 件 · §IX 113 morning v3.55 升档棒 · v3.54 morning 459 沿用基线 + 6 NET-new)
arXiv:2305.02189 arXiv:2310.11703 arXiv:2403.05527 arXiv:2407.00079 arXiv:2410.17043 arXiv:2412.14219 arXiv:2501.01005 arXiv:2501.09136 arXiv:2501.16383 arXiv:2502.04420 arXiv:2502.07115 arXiv:2502.14617 arXiv:2502.17421 arXiv:2503.10325 arXiv:2503.13657 arXiv:2504.07347 arXiv:2504.11320 arXiv:2504.19720 arXiv:2504.19874 arXiv:2505.02189 arXiv:2505.02922 arXiv:2505.08838 arXiv:2505.11329 arXiv:2506.01333 arXiv:2506.04565 arXiv:2506.09713 arXiv:2506.21901 arXiv:2507.06608 arXiv:2507.11507 arXiv:2507.12442 arXiv:2507.18007 arXiv:2508.04925 arXiv:2508.10991 arXiv:2508.13337 arXiv:2508.18572 arXiv:2509.01809 arXiv:2509.15000 arXiv:2510.02758 arXiv:2510.03215 arXiv:2510.05373 arXiv:2510.09665 arXiv:2510.13910 arXiv:2511.01815 arXiv:2511.02230 arXiv:2511.11581 arXiv:2511.16681 arXiv:2511.22880 arXiv:2512.02337 arXiv:2512.05411 arXiv:2512.09196 arXiv:2601.03236 arXiv:2601.05047 arXiv:2601.06112 arXiv:2601.06288 arXiv:2601.17549 arXiv:2601.19139 arXiv:2601.20408 arXiv:2602.00328 arXiv:2602.00751 arXiv:2602.01129 arXiv:2602.03442 arXiv:2602.04900 arXiv:2602.08005 arXiv:2602.14516 arXiv:2602.14617 arXiv:2602.19594 arXiv:2602.21548 arXiv:2602.23374 arXiv:2603.02001 arXiv:2603.04428 arXiv:2603.06728 arXiv:2603.09619 arXiv:2603.10031 arXiv:2603.10249 arXiv:2603.11088 arXiv:2603.12646 arXiv:2603.13358 arXiv:2603.15569 arXiv:2603.16104 arXiv:2603.17456 arXiv:2603.18272 arXiv:2603.18567 arXiv:2603.20397 arXiv:2603.21354 arXiv:2603.23710 arXiv:2603.25723 arXiv:2603.27467 arXiv:2603.29010 arXiv:2603.29231 arXiv:2604.00499 arXiv:2604.00901 arXiv:2604.01395 arXiv:2604.03143 arXiv:2604.04722 arXiv:2604.04853 arXiv:2604.05012 arXiv:2604.10235 arXiv:2604.12374 arXiv:2604.15732 arXiv:2604.16371 arXiv:2604.17227 arXiv:2604.19157 arXiv:2604.19769 arXiv:2604.20920 arXiv:2604.22513 arXiv:2604.22906 arXiv:2604.24971 arXiv:2604.25724 arXiv:2604.25850 arXiv:2604.25899 arXiv:2604.26557 arXiv:2604.27476 arXiv:2605.00528 arXiv:2605.01280 arXiv:2605.01495 arXiv:2605.02189 arXiv:2605.02922 arXiv:2605.03310 arXiv:2605.04595 arXiv:2605.05287 arXiv:2605.08838 arXiv:2605.10834 arXiv:2605.11733 arXiv:2605.13734 arXiv:2605.15040 arXiv:2605.15957 arXiv:2605.17613 arXiv:2605.18825 arXiv:2605.19537 arXiv:2605.19660 arXiv:2605.23389 arXiv:2605.27744 arXiv:2605.29639 arXiv:2605.29640 arXiv:2605.29979 arXiv:2605.31097 arXiv:2606.01927 arXiv:2606.02643 arXiv:2606.02964 arXiv:2606.03458 arXiv:2606.03811 arXiv:2606.06090 arXiv:2606.06535 arXiv:2606.07362 arXiv:2606.07402 arXiv:2606.11916 arXiv:2606.14589 arXiv:2606.16059 arXiv:2606.16316 arXiv:2606.17104 arXiv:2606.17107 arXiv:2606.18023 arXiv:2606.18431 arXiv:2606.19746 arXiv:2606.19803 arXiv:2606.20295 arXiv:2606.21238 arXiv:2606.21649 arXiv:2606.24775 arXiv:2606.26560 arXiv:2606.26875 arXiv:2606.28565 arXiv:2606.29708 arXiv:2606.30391 arXiv:2607.00482 arXiv:2607.02574 arXiv:2607.02980 arXiv:2607.03333 arXiv:2607.05061 arXiv:2607.05708 arXiv:2607.07386 arXiv:2607.07816 arXiv:2607.07953 arXiv:2607.08028 arXiv:2607.08057 arXiv:2607.09172 arXiv:2607.09248 arXiv:2607.09424 arXiv:2607.10350 arXiv:2607.10508 arXiv:2607.11505 arXiv:2607.11523 arXiv:2607.11783 arXiv:2607.11881 arXiv:2607.12747 arXiv:2607.13027 arXiv:2607.13104 arXiv:2607.13705 arXiv:2607.14541 arXiv:2607.17715 arXiv:2607.17979 arXiv:2607.18141 arXiv:2607.20468 arXiv:2607.21557 arXiv:2607.22529 arXiv:2607.23693 arXiv:2607.23933 arXiv:2607.24062 arXiv:2607.25380 arXiv:2607.26475 arXiv:2607.26654 arXiv:2607.27042 arXiv:2607.27090 arXiv:2607.28633 arXiv:2607.29377 arXiv:2607.29405 arXiv:2608.00101 arXiv:2608.00303 arXiv:2608.00677 arXiv:2608.00742 arXiv:2608.00881 arXiv:2608.00902 arXiv:2608.01247 arXiv:2608.01326 arXiv:2608.01526 arXiv:2608.01651 arXiv:2608.01718 arXiv:2608.01735 arXiv:2608.01964 arXiv:2608.01975 arXiv:2608.02143 arXiv:2608.02515 arXiv:2608.02585 arXiv:2608.02645 arXiv:2608.02703 arXiv:2608.02870 arXiv:2608.02989 arXiv:2608.03036 arXiv:2608.03216 arXiv:2608.03222 arXiv:2608.03487 arXiv:2608.03796 arXiv:2608.03893 arXiv:2608.03972 arXiv:2608.03994 arXiv:2608.04771 arXiv:2608.05136 arXiv:2608.05219 arXiv:2608.05604 arXiv:2608.05784 arXiv:2608.06007 arXiv:2608.06033 arXiv:2608.06130 arXiv:2608.06301 arXiv:2608.06557 arXiv:2608.06790 arXiv:2608.06867 arXiv:2608.07009 arXiv:2608.07152 arXiv:2608.07169 arXiv:2608.07458 arXiv:2608.07645 arXiv:2608.08020 arXiv:2608.08097 arXiv:2608.08311 arXiv:2608.08389 arXiv:2608.08878 arXiv:2608.09214 arXiv:2608.09867 arXiv:2608.09888 arXiv:2608.10208 arXiv:2608.10288 arXiv:2608.10875 arXiv:2608.10915 arXiv:2608.11660 arXiv:2608.11668 arXiv:2608.12149 arXiv:2608.12365 arXiv:2608.12440 arXiv:2608.13263 arXiv:2608.13426 arXiv:2608.13499 arXiv:2608.13567 arXiv:2608.13867 arXiv:2608.13868 arXiv:2608.13947 arXiv:2608.14192 arXiv:2608.14333 arXiv:2608.14376 arXiv:2608.14465 arXiv:2608.14680 arXiv:2608.15380 arXiv:2608.15579 arXiv:2608.15994 arXiv:2608.16157 arXiv:2608.16411 arXiv:2608.17050 arXiv:2608.18027 arXiv:2608.18050 arXiv:2608.19157 arXiv:2608.19269 arXiv:2608.19662 arXiv:2608.19758 arXiv:2608.19854 arXiv:2608.20359 arXiv:2608.20953 arXiv:2608.21252 arXiv:2608.21281 arXiv:2608.22643 arXiv:2608.23041 arXiv:2608.23283 arXiv:2608.23392 arXiv:2608.23553 arXiv:2608.24040 arXiv:2608.24622 arXiv:2608.24636 arXiv:2608.25375 arXiv:2608.26005 arXiv:2608.26021 arXiv:2608.26070 arXiv:2608.26530 arXiv:2608.26730 arXiv:2608.27455 arXiv:2608.27763 arXiv:2608.27875 arXiv:2608.28444 arXiv:2608.28458 arXiv:2608.29188 arXiv:2608.29253 arXiv:2608.30005 arXiv:2608.30135 arXiv:2608.30322 arXiv:2608.30391 arXiv:2608.30795 arXiv:2608.31082 arXiv:2608.31100 arXiv:2608.31111 arXiv:2608.31139 arXiv:2608.09444 arXiv:2609.00006 arXiv:2609.00196 arXiv:2609.00374 arXiv:2609.00749 arXiv:2609.00768 arXiv:2609.00891 arXiv:2609.01072 arXiv:2609.01316 arXiv:2609.01437 arXiv:2609.01481 arXiv:2609.01532 arXiv:2609.01572 arXiv:2609.01657 arXiv:2609.01777 arXiv:2609.01878 arXiv:2609.01925 arXiv:2609.02029 arXiv:2609.02373 arXiv:2609.02496 arXiv:2609.02745 arXiv:2609.02783 arXiv:2609.03153 arXiv:2609.03293 arXiv:2609.03430 arXiv:2609.03807 arXiv:2609.03820 arXiv:2609.03949 arXiv:2609.04131 arXiv:2609.04148 arXiv:2609.04199 arXiv:2609.04201 arXiv:2609.04263 arXiv:2609.04490 arXiv:2609.04971 arXiv:2609.05565 arXiv:2609.06008 arXiv:2609.06128 arXiv:2609.06674 arXiv:2609.07139 arXiv:2609.07529 arXiv:2609.08887 arXiv:2609.09072 arXiv:2609.09085 arXiv:2609.09156 arXiv:2609.10226 arXiv:2609.10266 arXiv:2609.10445 arXiv:2609.11085 arXiv:2609.11209 arXiv:2609.11390 arXiv:2609.11412 arXiv:2609.11561 arXiv:2609.11596 arXiv:2609.11582 arXiv:2609.11596 arXiv:2609.11699 arXiv:2609.11758 arXiv:2609.11744 arXiv:2609.12923 arXiv:2609.13141 arXiv:2609.13285 arXiv:2609.15504 arXiv:2609.15524 arXiv:2609.15810 arXiv:2609.15938 arXiv:2609.15972 arXiv:2609.16453 arXiv:2609.16818 arXiv:2609.17012 arXiv:2609.17391 arXiv:2609.17475 arXiv:2609.17524 arXiv:2609.17652 arXiv:2609.17708 arXiv:2609.17863 arXiv:2609.18063 arXiv:2609.18708 arXiv:2609.19134 arXiv:2609.19169 arXiv:2609.19499 arXiv:2609.19657 arXiv:2609.19969 arXiv:2609.20423 arXiv:2609.20511 arXiv:2609.20612 arXiv:2609.21346 arXiv:2609.22157 arXiv:2609.24797 arXiv:2609.29845 arXiv:2509.15000 arXiv:2604.05887 arXiv:2609.24991 arXiv:2609.25053 arXiv:2609.26346 arXiv:2609.26333 arXiv:2511.22880 arXiv:2501.01005 arXiv:2609.27980 arXiv:2609.25963 arXiv:2609.27158 arXiv:2609.23087 arXiv:2609.23130 arXiv:2609.27746 arXiv:2609.28870 arXiv:2609.29647 arXiv:2609.30216 arXiv:2609.31009 arXiv:2609.31093 arXiv:2609.31397 arXiv:2609.31415 arXiv:2504.07347 arXiv:2601.15232 arXiv:2606.00516 arXiv:2609.12551 arXiv:2609.32759 arXiv:2609.33485 arXiv:2609.27334 arXiv:2609.25853 arXiv:2609.34727 arXiv:2609.32259 arXiv:2609.39929 arXiv:2609.38987 arXiv:2609.36435 arXiv:2609.37725 arXiv:2609.39223 arXiv:2610.00182 arXiv:2610.00905 arXiv:2610.02772 arXiv:2610.03394 arXiv:2610.04646 arXiv:2610.06479 arXiv:2602.09323 arXiv:2609.34117 arXiv:2609.39096 arXiv:2610.07332 arXiv:2610.08430 arXiv:2610.08778 arXiv:2501.09223 arXiv:2609.23130 arXiv:2609.38169 arXiv:2609.39334 arXiv:2610.07219 arXiv:2610.10845
v3.54 morning 沿用基线 NET-new arXiv ID(6 件 · 沿用基线 453 + 6 NET-new = 459):
arXiv:2602.09323 arXiv:2609.34117 arXiv:2609.39096 arXiv:2610.07332 arXiv:2610.08430 arXiv:2610.08778
v3.55 morning 本轮新增 arXiv ID(6 件 · 沿用基线 459 + 6 NET-new = 465):
arXiv:2501.09223 arXiv:2609.23130 arXiv:2609.38169 arXiv:2609.39334 arXiv:2610.07219 arXiv:2610.10845
6.99.3 全量 CVE(共 36 件 · §IX 113 morning v3.55 升档棒沿用 · v3.54 morning 36 沿用基线)
CVE-2025-15558 CVE-2025-30165 CVE-2025-49596 CVE-2025-53109 CVE-2025-53110 CVE-2025-54135 CVE-2025-54136 CVE-2025-62164 CVE-2025-6514 CVE-2025-66448 CVE-2025-68143 CVE-2026-14890 CVE-2026-22773 CVE-2026-22778 CVE-2026-24779 CVE-2026-25960 CVE-2026-26030 CVE-2026-26384 CVE-2026-27893 CVE-2026-3059 CVE-2026-3060 CVE-2026-30623 CVE-2026-3172 CVE-2026-3288 CVE-2026-33032 CVE-2026-33626 CVE-2026-3864 CVE-2026-3865 CVE-2026-3989 CVE-2026-4342 CVE-2026-5241 CVE-2026-55765 CVE-2026-55769 CVE-2026-5760 CVE-2026-61539 CVE-2026-73558
6.99.4 全量 DOI(共 15 件 · §IX 113 morning v3.55 升档棒沿用 · v3.54 morning 15 沿用基线)
DOI:10.1109/TCAD.2026.11371745 DOI:10.1145/3749168 DOI:10.1145/3773772 DOI:10.1145/38020094 DOI:10.1145/38020284 DOI:10.1145/3803798 DOI:10.1145/3832810.3832862 DOI:10.48550/arxiv.2608.13867 DOI:10.48550/arxiv.2608.16411 DOI:10.48550/arxiv.2608.25375 DOI:10.48550/arxiv.2608.26070 DOI:10.48550/arxiv.2608.27763 DOI:10.48550/arxiv.2608.28444 DOI:10.5281/zenodo.19686729 DOI:10.5281/zenodo.21771064
6.99.5 全量 URL(共 616 件 · §IX 113 morning v3.55 升档棒 · v3.54 morning 607 沿用基线 + 9 NET-new)
https://academy.dair.ai/papers/harnessdev-can-llms-create-and-evolve-their-own-agent-harness-2609.01437 https://acecloud.ai/blog/best-vector-databases-for-multimodal-genai https://acmsocc.org/2026/accepted-papers.html https://addyo.substack.com/p/my-llm-coding-workflow-going-into https://adg.csdn.net/6a311ff410ee7a33f27df3dd.html https://advisories.gitlab.com/pypi/xinference/CVE-2026-61539 https://aiamastery.substack.com/p/production-ai-engineering-building https://aicoding.csdn.net/6a3cf743662f9a54cb844097.html https://aimultiple.com/inference-engines https://aimultiple.com/self-hosted-llm https://alexeyondata.substack.com/p/what-1000-job-descriptions-reveal https://alphasignal.ai/news/lmsys-rebuilds-sglang-s-cache-to-finally-support-hybrid-ai-models https://ampcobe.com/blog/pydantic-ai-2026-breakout https://anyscale.com/ray-summit/2026 https://app.opencve.io/cve?product=vllm&vendor=vllm-project https://ar5iv.labs.arxiv.org/html/2601.19139 https://arxiv.org/abs/2410.17043 https://arxiv.org/abs/2412.14219 https://arxiv.org/abs/2502.07115 https://arxiv.org/abs/2504.01395 https://arxiv.org/abs/2504.11320v4 https://arxiv.org/abs/2505.02922 https://arxiv.org/abs/2505.11329 https://arxiv.org/abs/2506.01333 https://arxiv.org/abs/2506.21901 https://arxiv.org/abs/2507.11507 https://arxiv.org/abs/2507.18007 https://arxiv.org/abs/2508.10991 https://arxiv.org/abs/2508.13337 https://arxiv.org/abs/2508.18572 https://arxiv.org/abs/2509.01809 https://arxiv.org/abs/2510.03215 https://arxiv.org/abs/2510.09665 https://arxiv.org/abs/2510.13910v2 https://arxiv.org/abs/2511.02230 https://arxiv.org/abs/2512.02337 https://arxiv.org/abs/2601.03236 https://arxiv.org/abs/2601.06112 https://arxiv.org/abs/2601.06288 https://arxiv.org/abs/2601.17549 https://arxiv.org/abs/2602.00328 https://arxiv.org/abs/2602.01129 https://arxiv.org/abs/2602.19594 https://arxiv.org/abs/2602.21548 https://arxiv.org/abs/2602.23374 https://arxiv.org/abs/2603.16104 https://arxiv.org/abs/2603.20397v1 https://arxiv.org/abs/2603.21354v2 https://arxiv.org/abs/2603.23710 https://arxiv.org/abs/2604.05012 https://arxiv.org/abs/2604.22906 https://arxiv.org/abs/2604.27476v1 https://arxiv.org/abs/2605.00528 https://arxiv.org/abs/2605.01280 https://arxiv.org/abs/2605.02189 https://arxiv.org/abs/2605.13734 https://arxiv.org/abs/2605.15957v1 https://arxiv.org/abs/2605.18825 https://arxiv.org/abs/2605.19537 https://arxiv.org/abs/2605.19660 https://arxiv.org/abs/2606.01927 https://arxiv.org/abs/2606.06090 https://arxiv.org/abs/2606.14589 https://arxiv.org/abs/2606.17104 https://arxiv.org/abs/2606.17107 https://arxiv.org/abs/2606.19746 https://arxiv.org/abs/2606.19803 https://arxiv.org/abs/2607.05708 https://arxiv.org/abs/2607.07386 https://arxiv.org/abs/2607.08057 https://arxiv.org/abs/2607.27090 https://arxiv.org/abs/2607.28633 https://arxiv.org/abs/2608.00902 https://arxiv.org/abs/2608.01526 https://arxiv.org/abs/2608.03036 https://arxiv.org/abs/2608.03893 https://arxiv.org/abs/2608.04771 https://arxiv.org/abs/2608.06790 https://arxiv.org/abs/2608.08097 https://arxiv.org/abs/2608.08878 https://arxiv.org/abs/2608.09444 https://arxiv.org/abs/2608.09867 https://arxiv.org/abs/2608.10288 https://arxiv.org/abs/2608.13426 https://arxiv.org/abs/2608.13867 https://arxiv.org/abs/2608.13868 https://arxiv.org/abs/2608.14333 https://arxiv.org/abs/2608.14376 https://arxiv.org/abs/2608.14680 https://arxiv.org/abs/2608.15579 https://arxiv.org/abs/2608.16157 https://arxiv.org/abs/2608.16411 https://arxiv.org/abs/2608.17050 https://arxiv.org/abs/2608.18027 https://arxiv.org/abs/2608.19269 https://arxiv.org/abs/2608.19662 https://arxiv.org/abs/2608.19758 https://arxiv.org/abs/2608.20359 https://arxiv.org/abs/2608.20953 https://arxiv.org/abs/2608.21252 https://arxiv.org/abs/2608.21281 https://arxiv.org/abs/2608.22643 https://arxiv.org/abs/2608.23041 https://arxiv.org/abs/2608.23283 https://arxiv.org/abs/2608.23392 https://arxiv.org/abs/2608.23553 https://arxiv.org/abs/2608.24040 https://arxiv.org/abs/2608.24622 https://arxiv.org/abs/2608.24636 https://arxiv.org/abs/2608.25375 https://arxiv.org/abs/2608.26021 https://arxiv.org/abs/2608.26070 https://arxiv.org/abs/2608.26530 https://arxiv.org/abs/2608.27455 https://arxiv.org/abs/2608.27763 https://arxiv.org/abs/2608.27875 https://arxiv.org/abs/2608.28444 https://arxiv.org/abs/2608.29188 https://arxiv.org/abs/2608.30795 https://arxiv.org/abs/2609.00374 https://arxiv.org/abs/2609.00768 https://arxiv.org/abs/2609.00891 https://arxiv.org/abs/2609.01316 https://arxiv.org/abs/2609.01437 https://arxiv.org/abs/2609.01481 https://arxiv.org/abs/2609.01777 https://arxiv.org/abs/2609.01925 https://arxiv.org/abs/2609.02029 https://arxiv.org/abs/2609.04263 https://arxiv.org/abs/2609.04971 https://arxiv.org/abs/2609.05565 https://arxiv.org/abs/2609.06008 https://arxiv.org/abs/2609.07139 https://arxiv.org/abs/2609.08887 https://arxiv.org/abs/2609.09072 https://arxiv.org/abs/2609.09085 https://arxiv.org/abs/2609.09156 https://arxiv.org/abs/2609.10266 https://arxiv.org/abs/2609.11412 https://arxiv.org/abs/2609.12923 https://arxiv.org/abs/2609.13285 https://arxiv.org/abs/2609.15504 https://arxiv.org/abs/2609.17391 https://arxiv.org/abs/2609.17475 https://arxiv.org/abs/2609.19169 https://arxiv.org/abs/2609.19499 https://arxiv.org/abs/2609.21346 https://arxiv.org/abs/2609.22157 https://arxiv.org/abs/2609.23087 https://arxiv.org/abs/2609.23130 https://arxiv.org/abs/2609.24991 https://arxiv.org/abs/2609.26333 https://arxiv.org/abs/2609.27746 https://arxiv.org/abs/2609.28870 https://arxiv.org/abs/2609.29647 https://arxiv.org/abs/2609.30216 https://arxiv.org/abs/2609.31009 https://arxiv.org/abs/2609.31093 https://arxiv.org/abs/2609.31397 https://arxiv.org/abs/2609.31415 https://arxiv.org/abs/2609.33485 https://arxiv.org/abs/2609.37725 https://arxiv.org/html/2502.07115v5 https://arxiv.org/html/2505.11329v2 https://arxiv.org/html/2506.09713v1 https://arxiv.org/html/2510.03215v2 https://arxiv.org/html/2510.09665v1 https://arxiv.org/html/2510.09665v2 https://arxiv.org/html/2510.13910v2 https://arxiv.org/html/2508.18572v1 https://arxiv.org/html/2601.06288v1 https://arxiv.org/html/2601.20408v2 https://arxiv.org/html/2602.00328v1 https://arxiv.org/html/2602.14516v2 https://arxiv.org/html/2602.19594 https://arxiv.org/html/2602.21548v2 https://arxiv.org/html/2603.20397v1 https://arxiv.org/html/2604.01395v1 https://arxiv.org/html/2604.22513v1 https://arxiv.org/html/2604.22906v1 https://arxiv.org/html/2605.02189v1 https://arxiv.org/html/2605.13734v1 https://arxiv.org/html/2605.29639v1 https://arxiv.org/html/2606.07362v1 https://arxiv.org/html/2606.16059v1 https://arxiv.org/html/2606.18431v1 https://arxiv.org/html/2607.02574v1 https://arxiv.org/html/2608.01526v1 https://arxiv.org/html/2608.03036v1 https://arxiv.org/html/2608.09444 https://arxiv.org/html/2608.11668v2 https://arxiv.org/html/2608.13868v1 https://arxiv.org/html/2608.16157v1 https://arxiv.org/html/2608.20953v1 https://arxiv.org/html/2608.22643v1 https://arxiv.org/html/2608.27875v1 https://arxiv.org/html/2608.29188v1 https://arxiv.org/html/2609.03807v1 https://arxiv.org/html/2609.03949v1 https://arxiv.org/html/2609.05565v1 https://arxiv.org/html/2609.17391v1 https://arxiv.org/html/2609.17475v1 https://arxiv.org/html/2609.23130v1 https://arxiv.org/html/2609.26333 https://arxiv.org/html/2609.27746v1 https://arxiv.org/html/2609.31415v1 https://arxiv.org/pdf/2503.13657 https://arxiv.org/pdf/2602.19594 https://arxiv.org/pdf/2608.01526 https://arxiv.org/pdf/2608.06790 https://arxiv.org/pdf/2608.13867 https://arxiv.org/pdf/2608.19758 https://arxiv.org/pdf/2608.20953 https://arxiv.org/pdf/2608.29188 https://arxiv.org/pdf/2609.02029 https://arxiv.org/pdf/2609.19499 https://arxiv.org/pdf/2609.19657 https://arxiv.org/pdf/2609.23130 https://atomic.chat/blog/llm-updates/sglang-vs-vllm https://aws.amazon.com/blogs/machine-learning/disaggregated-prefill-and-decode-for-llm-inference-on-sagemaker-hyperpod https://aws.amazon.com/cn/blogs/china/based-on-sglang-large-model-inference-practice/ https://baotonglu.github.io/RetroInfer_page/RetroInfer.html https://benchlm.ai/benchmarks/swe-bench-verified https://berkeleyrdi.substack.com/p/agentic-ai-weekly-berkeley-rdi-august https://blog.bytebytego.com/p/ep223-ollama-vs-vllm-vs-sglang https://blog.bytebytego.com/p/how-to-make-llms-3x-faster https://blog.bytebytego.com/p/how-to-steal-an-ai-models-private https://blog.cloudflare.com/mcp-v2 https://blog.csdn.net/2301_81666833/article/details/163924904 https://blog.csdn.net/brandy/article/details/155517690 https://blog.csdn.net/m0_69378371/article/details/158455491 https://blog.csdn.net/qq_31142761/article/details/161787922 https://blog.csdn.net/qq_73472828/article/details/160875055 https://blog.csdn.net/u011091936/article/details/150429518 https://blog.csdn.net/u011576070/article/details/146889758 https://blog.csdn.net/u013701860/article/details/146889758 https://blog.csdn.net/u013701860/article/details/148295809 https://blog.modelcontextprotocol.io/posts/2026-07-28 https://blog.premai.io/vllm-vs-sglang-vs-lmdeploy-fastest-llm-inference-engine-in-2026 https://blog.squeezebits.com/vllm-korea-meetup-highlights https://blogs.nvidia.com/blog/nvidia-to-acquire-hugging-face https://blogs.oracle.com/ai-and-datascience/llm-inference-at-scale-with-llm-d-on-oci https://buttondown.com/weekly-project-news/archive/weekly-github-report-for-llamacpp-january-16-2026 https://bytebytego.com/top-ai-github-repositories-2026 https://catalog.ngc.nvidia.com/orgs/nvidia/ai-dynamo/containers/vllm-runtime/- https://catalog.ngc.nvidia.com/orgs/nvidia/ai-dynamo/containers/vllm-runtime/1.0.0-cuda13 https://chipsandcheese.com/p/hot-chips-2026-applying-high-bandwidth https://cloudai.pt/prefill-decode-disaggregation-doubles-your-llm-throughput https://cloudnative-pg.io/releases/cloudnative-pg-1-30-0-released https://cloudnativenow.com/features/cncf-expands-efforts-to-run-ai-inference-workloads-on-kubernetes-clusters https://codingwithroby.substack.com/p/the-2026-ai-agent-stack-drawn-from https://daily.dev/blog/ai-agents-guide-for-developers-langchain-crewai https://daily.dev/posts/serving-agentic-workloads-at-scale-with-vllm-x-mooncake-h1aaslvqr https://daily.dev/posts/vllm-sessions-at-pytorch-conference-north-america-2026-pytorch-oofqixpfe https://datatracker.ietf.org/doc/draft-li-cats-kv-cache-distribution https://datatracker.ietf.org/doc/html/draft-li-cats-kv-cache-distribution-00 https://deepinfra.com/blog/vllm-vs-sglang https://deepseek.csdn.net/6a04133a54b52172bc73ae00.html https://deepseek.csdn.net/6a211c7910ee7a33f277840f.html https://deploybase.ai/articles/best-llm-inference-engine https://dev.to/x4nent/complete-guide-to-llm-d-cncf-sandbox-kubernetes-native-distributed-llm-inference-1imj https://developer.aliyun.com/article/1704054 https://developer.nvidia.com/blog/cuda-python-1-0-stable-apis-one-foundation-full-platform-access/ https://developer.nvidia.com/blog/deploying-disaggregated-llm-inference-workloads-on-kubernetes https://developers.redhat.com/articles/2025/10/03/deepseek-v32-exp-vllm-day-0-sparse-attention-long-context-inference https://devopsbeast.com/blog/vllm-vs-sglang-production-2026 https://digg.com/tech/pk1l8ong https://discuss.google.dev/t/optimizing-llm-inference-for-minimal-latency-with-vllm/289241 https://discuss.vllm.ai/ https://dl.acm.org/doi/full/10.1145/3749168 https://dl.acm.org/doi/full/10.1145/3773772 https://docs.lmcache.ai https://docs.lmcache.ai/recipes/kimi_k3.html https://docs.nvidia.com/dynamo/dev/reference/compatibility https://docs.vllm.ai/en/latest/examples/disagg/disagg_proxy_multiturn.py https://docs.vllm.ai/en/latest/features/speculative_decoding/mtp https://docs.vllm.ai/projects/vllm-omni/en/latest/user_guide/quantization/modelopt https://docs.vllm.ai/projects/vllm-omni/en/latest/user_guide/quantization/online https://doi.org/10.1145/3803798 https://emergingai.substack.com/p/local-llms-in-2026-the-simple-practical https://emergingai.substack.com/p/master-inference-engineering-the https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cloud-native-ai-inference-day https://everythinginsigcomm.group/t/kvserve-service-aware-kv-cache-compression-for-communication-efficient-disaggregated-llm-serving/458 https://explore.n1n.ai/blog/vllm-vs-sglang-vs-lmdeploy-fastest-inference-2026-2026-03-05 https://firecrawl.dev/blog/best-vector-databases https://fish.audio/blog/open-source-llm-inference-engines-2026 https://freedom.tech/project/vllm https://futureagi.substack.com/p/llm-evaluation-frameworks-metrics https://futureagi.substack.com/p/why-do-multi-agent-llm-systems-fail https://futurumgroup.com/insights/vllm-becomes-production-infrastructure-at-pytorch-conference-2026 https://gateway-api.sigs.k8s.io/ https://github.com/FFY0/DefensiveKV https://github.com/InternLM/lmdeploy/security/advisories https://github.com/MakazhanAlpamys/Soup https://github.com/RightNow-AI/TIDE https://github.com/RyanAlberts/best-of-Agent-Harnesses https://github.com/StarTrail-org/LEANN https://github.com/TanZhendong/SpecPV https://github.com/WeianMao/triattention https://github.com/ai-boost/awesome-harness-engineering https://github.com/ai-dynamo/dynamo https://github.com/ai-dynamo/dynamo/releases https://github.com/alibaba/rtp-llm https://github.com/amitshekhariitbhu/llm-inference-engineering https://github.com/bentoml/BentoML https://github.com/bs258q/kv-cache-analyzer https://github.com/dair-ai/AI-Papers-of-the-Week https://github.com/deepopen-com/deepopen https://github.com/devflowinc/uzi https://github.com/flagos-ai/awesome-LLM-driven-kernel-generation https://github.com/flashinfer-ai/flashinfer https://github.com/framsouza/inference-at-scale-on-kubernetes https://github.com/francisown/Qwen3-4B-FP8-Inference https://github.com/ggml-org/llama.cpp/discussions/6730 https://github.com/gpustack/gpustack https://github.com/headroomlabs-ai/headroom https://github.com/hogeheer499-commits/strix-halo-guide https://github.com/hpdps-group/KVServe https://github.com/ishan1410/PolyKV https://github.com/jjiantong/Awesome-KV-Cache-Optimization https://github.com/jmaczan/tiny-vllm https://github.com/karpathy/nanochat https://github.com/kubernetes-sigs/lws https://github.com/kubernetes/ingress-nginx https://github.com/kvcache-ai/Mooncake https://github.com/langgenius/dify https://github.com/llm-d/llm-d-kv-cache https://github.com/lmcache/lmcache https://github.com/malisper/pgrust https://github.com/matrixhub-ai/matrixhub https://github.com/microsoft/retrievalattention https://github.com/moonshotai/MoonEP https://github.com/mudler/LocalAI https://github.com/mudler/vllm.cpp https://github.com/okf-memory/okf-agent-memory https://github.com/qhfan/FlashPrefillv2 https://github.com/sgl-project/sglang/issues/13165 https://github.com/sgl-project/sglang/issues/20415 https://github.com/sgl-project/sglang/issues/21994 https://github.com/sgl-project/sglang/issues/23842 https://github.com/sgl-project/sglang/issues/3471 https://github.com/sgl-project/sglang/releases https://github.com/sgl-project/sglang/releases/tag/v0.5.20 https://github.com/sgl-project/sglang/security/advisories https://github.com/sihyeong/Awesome-LLM-Inference-Engine https://github.com/sjarmak/engineering-reliable-coding-agents https://github.com/thu-nics/C2C https://github.com/thushan/olla https://github.com/topics/inference https://github.com/treeai-lab/awesome-kv-cache-management https://github.com/umwyf/CRITICL https://github.com/vllm-project/vllm-project.github.io/blob/main/_posts/2026-05-26-eagle-3-1.md https://github.com/vllm-project/vllm/issues/30448 https://github.com/vllm-project/vllm/issues/31109 https://github.com/vllm-project/vllm/issues/38339 https://github.com/vllm-project/vllm/issues/48168 https://github.com/vllm-project/vllm/releases/tag/v0.28.0 https://github.com/vllm-project/vllm/releases/tag/v0.30.0 https://github.com/vllm-project/vllm/security/advisories/GHSA-7m6h-x95x-82q5 https://github.com/warpfront/hipfire https://github.com/xorbitsai/inference/releases/tag/v3.2.0 https://github.com/yshk-mxim/agent-memory https://gradientflow.substack.com/p/rag-reimagined-5-breakthroughs-you https://help.aliyun.com/zh/functioncompute/performance-comparison-of-deploying-qwen-models-using-sglang-and-vllm https://highlimitdesigns.com/blog/llm-infrastructure-breakthroughs-vllm-v1-sglang-epd-disaggregation https://highlimitdesigns.com/blog/prefill-decode-disaggregation-llm-serving-2026 https://hotinfra.org/2026/papers/hotinfra26-final59.pdf https://huggingface.co/blog https://huggingface.co/blog/MultiverseComputingCAI/quantization-aware-healing https://huggingface.co/blog/microsoft/foundry-managed-compute https://huggingface.co/blog/native-speed-vllm-transformers-backend https://huggingface.co/blog/webgpu-kernels https://huggingface.co/papers/2503.13657 https://huggingface.co/papers/2505.11329 https://huggingface.co/papers/2510.09665 https://huggingface.co/papers/2510.13910v2 https://huggingface.co/papers/2602.19594 https://huggingface.co/papers/2605.13734 https://huggingface.co/papers/2608.16157 https://huggingface.co/papers/2609.01437 https://huggingface.co/papers/2609.34117 https://hugobowne.substack.com/p/agentops-lessons-from-over-1400-production https://hugobowne.substack.com/p/llm-architecture-in-2026-agent-harnesses https://inference.net/blog/sglang-complete-guide https://inference.net/content/sglang-complete-guide https://inferenceengineering.tech/learn/vllm-vs-sglang-vs-tensorrt-llm https://inferenceops.substack.com/p/state-of-the-model-serving-communities-b93 https://introl.com/blog/tensorrt-llm-optimization-nvidia-inference-stack-guide https://jamwithai.substack.com/p/the-2026-roadmap-production-aiml https://kairos.io/blog/11-minute-k8s-upgrade https://kenhuangus.substack.com https://kenhuangus.substack.com/p/announcing-the-10-part-series-the https://kodemsecurity.com https://kodemsecurity.com/cve-archive/cve-2026-61539 https://kodemsecurity.com/resources/cve-2026-22778-critical-remote-code-execution-in-vllm-multimodal-inference https://kubernetes.io/blog/2026/03/20/running-agents-on-kubernetes-with-agent-sandbox https://kubernetes.io/blog/2026/09/02/kubernetes-v1-37-hpa-scale-to-zero-beta https://kubesimplify.com https://kvcache-ai.github.io/Mooncake https://kvcache.ai/blog https://labs.cloudsecurityalliance.org/agentic/agentic-mcp-security-best-practices-v1 https://labs.cloudsecurityalliance.org/research/csa-research-note-mcp-security-crisis-20260504-csa-styled https://leetllm.com/blog/llm-inference-engine-comparison-2026 https://linkedin.com/posts/googleresearch_introducing-turboquant-our-new-compression-activity-7442298962369216512-DKnF https://linkedin.com/posts/rebellions-ai_the-vllm-korea-meetup-2026-brought-together-activity-7450528454762053632-qZen https://linkedin.com/posts/servergurus-india_disaggregated-inference-why-splitting-prefill-activity-7489815467243540480-FGVo https://llm-d.ai/blog/production-grade-llm-inference-at-scale-kserve-llm-d-vllm https://llms3.com/node/dell-objectscale https://lmcache.ai https://lmsys.org/blog https://lmsys.org/blog/2026-06-15-next-generation-speculative-decoding-dflash-v2 https://lmsys.org/blog/2026-08-11-unified-radix-cache https://lmsys.org/blog/2026-08-17-advanced-cuda-graph https://lmsys.org/blog/2026-08-18-miles-v0-1 https://lmsys.org/blog/2026-08-19-deepseek-v4-pro-engine-optimization-h20 https://localaimaster.com/blog/sglang-vs-vllang-comparison https://lucaberton.com/blog/ai-model-serving-kubernetes-vllm-triton-nim-2026 https://lyceum.technology/magazine/vllm-vs-tensorrt-llm-production-benchmark https://martinfowler.com/articles/engineering-practices-llm.html https://mcp.csdn.net/6a2e4df2662f9a54cb7eeb74.html https://media.defense.gov/2026/Jun/02/2003943289/-1/-1/0/CSI_MCP_SECURITY.PDF https://medium.com/@adityaj5400/the-kv-cache-is-killing-your-llm-at-scale-heres-the-low-level-physics-nobody-talks-about-b577c4c7549e https://medium.com/@pratik-rupareliya/top-15-vector-databases-in-2026-a-production-decision-guide-from-100-enterprise-deployments-dd58a04f51a5 https://medium.com/@surbhi19/we-put-our-production-database-on-kubernetes-heres-what-dbre-taught-us https://medium.com/data-science-in-your-pocket/vllm-x-qwen3-next-hybrid-attention-multi-token-prediction-and-thinking-controls-for-a0f6b3dcc120 https://mlflow.org/articles/the-role-of-kubernetes-in-ai-serving-2026-guide https://mlops.substack.com/p/how-to-defeat-non-determinism-in https://mlsys.org/virtual/2026/papers.html https://modelers.csdn.net/69a698f47bbde9200b9c9954.html https://moondream.ai/blog/photon-2-launch https://mp.weixin.qq.com/s?__biz=MzkyMzI3NzQ0Mg%3D%3D&mid=2247494105&idx=1&sn=8d7409e0fb846a3c7803c142b5d1a8e7 https://nand-research.com/high-bandwidth-flash-bridging-the-gap-between-expensive-hbm-flash-memory https://nebius.com/events/ray-summit-2026 https://newreleases.io/project/pypi/vllm/release/0.28.0 https://nvidia.com/en-us/events/ray-summit https://nvidia.github.io/TensorRT-LLM/release-notes.html https://omidsaffari.com/blog/best-ai-inference-orchestration-platforms-2026 https://openai.com/index/hugging-face-incident-and-the-road-ahead https://opencve.io/cve?product=vllm&vendor=vllm-project https://opendatascience.com/vllm-transformers-backend-bridging-hugging-face-compatibility-and-high-performance-inference https://openeuler.csdn.net/6a20ddae10ee7a33f2776b12.html https://openeuler.csdn.net/6a5f9f8310ee7a33f2911aa6.html https://orca.security/resources/blog/cve-2026-22778-vllm-rce-vulnerability https://orca.security/resources/blog/sglang-llm-framework-rce-vulnerabilities https://paper.dou.ac/arxiv/2608.21281 https://particula.tech/blog/sglang-vs-vllm-inference-engine-comparison https://pecollective.com/tools/pgvector https://premai.io/blog/10-best-vllm-alternatives-for-llm-inference-in-production-2026 https://premai.io/blog/vllm-vs-sllm-vs-lmdeploy-fastest-llm-inference-engine-in-2026 https://pytorch.org/announcements https://pytorch.org/blog/vllm-sessions-at-pytorch-conference-north-america-2026 https://relvehq.com/events/ray-summit https://research.yandex.com/blog/the-kv-cache-as-an-agent-runtime https://rmarcus.info/dbscholar/papers/h77ccaef2a0eb0f9c https://rockybhatia.substack.com/p/how-to-learn-agentic-rag-in-2026 https://rocm.blogs.amd.com/artificial-intelligence/kimi-k3-mad/README.html https://rocm.blogs.amd.com/artificial-intelligence/turboquant-vllm-agentic/README.html https://rocm.blogs.amd.com/vllm-multimodal-dp-vision https://saeed.github.io/files/arc_niac26.pdf https://securitywall.co/blog/mcp-security-testing-guide https://shattered.io/posts/kubernetes-1-37-ga https://sidsaladi.substack.com/p/agent-frameworks-101-the-complete https://siliconangle.com/2026/07/17/ey-re-envisions-rag-around-multimodal-knowledge-graphs-improve-accuracy/ https://simonw.substack.com/p/fireside-chat-about-agentic-engineering https://sjarmak.ai/books/engineering-reliable-coding-agents/explore https://srekubecraft.io/posts/llm-d-distributed-inference https://startupcorners.com/digest/devtools-digest-2026-08-06 https://strix.ai https://substack.com/@systemdesignone/note/c-253508965 https://technspire.com/en/blog/vector-search-2026-azure-pgvector-managed https://tensorwave.com/blog/breaking-free-from-gpu-fragmentation https://theaiengineer.substack.com/p/the-2026-ai-agent-stack-2026-edition https://theaiengineer.substack.com/p/the-ai-agents-stack-2026-edition https://theaiengineer.substack.com/p/vllm-vs-ollama-vs-sglang-vs-tensorrt https://thenewstack.com https://tianpan.co/forum/t/distributed-systems-architecture-patterns-for-edge-computing-in-2026/1225 https://uvik.net/blog/langchain-vs-langgraph https://vals.ai/benchmarks/swebench https://vecdb-ws.github.io/vldb2026 https://vllm-project.github.io/blog/2026-04-14-disaggregated-serving-for-hybrid-ssm-models https://vllm-project.github.io/blog/2026-04-21-disaggregated-serving-for-hybrid-ssm-models https://vllm-project.github.io/blog/2026-04-21-state-of-fp8-kv-cache https://vllm-project.github.io/blog/2026-04-28-vllm-x-mooncake https://vllm-project.github.io/blog/2026-07-29-25k-tps-qwen35 https://vllm-project.github.io/blog/2026-08-06-decode-context-parallelism https://vllm-project.github.io/blog/2026-08-07-efficient-decode-context-parallelism https://vllm.ai/blog https://vllm.ai/blog/2026-04-07-vllm-korea-meetup-2026-wrap-up https://vllm.ai/blog/2026-04-14-vllm-korea-meetup-2026 https://vllm.ai/blog/2026-05-06-mooncake-store https://vllm.ai/blog/2026-09-08-vllm-agentx https://vllm.ai/blog/2026-09-10-tiered-kv-offloading https://vllm.ai/blog/glm-5-2-24xb300-sla-optimization https://vllm.ai/blog/minimax-m3-day-0-serving https://vllm.ai/events/vllm-conference/2026 https://www.air.security/blog-posts/plugin4shell https://www.alibabacloud.com/blog/hybrid-model-support-%7C-sglangs-support-scheme-for-hybrid-architecture-models-like-mamba-transformer_602857 https://www.alphaxiv.org/abs/2505.11329 https://www.alphaxiv.org/abs/2510.13910v2 https://www.alphaxiv.org/abs/2512.02337v1 https://www.alphaxiv.org/abs/2602.19594 https://www.alphaxiv.org/abs/2605.29639v1 https://www.alphaxiv.org/abs/2608.13867 https://www.alphaxiv.org/abs/2608.16157 https://www.alphaxiv.org/abs/2609.00891 https://www.alphaxiv.org/abs/2609.01437 https://www.aussieai.com/blog/llm-inference-optimization https://www.axios.com/2026/08/17/a2a-agentic-ai-foundation-open-ai-standards https://www.bio-itworld.com/news/2026/09/03/nvidia-acquires-hugging-face-for--12.93-billion https://www.braintrust.dev/articles/best-vector-databases-for-rag-2026 https://www.buildmvpfast.com/blog/pinecone-vs-weaviate-vs-qdrant-vector-database-comparison-2026 https://www.cncf.io/blog/2026/03/24/welcome-llm-d-to-the-cncf-evolving-kubernetes-into-sota-ai-infrastructure https://www.cncf.io/blog/2026/08/05/opencost-llm-d-integration-1-121-0 https://www.coalitionforsecureai.org/securing-the-ai-agent-revolution-a-practical-guide-to-mcp-security https://www.crusoe.ai/resources/blog/crusoe-managed-inference-optimize-performance-for-demanding-ai-inference-workloads https://www.cve.org/CVERecord?id=CVE-2026-73558 https://www.developersdigest.tech/blog/agentchaos-fault-injection-agent-robustness https://www.digitalapplied.com/blog/kv-cache-optimization-techniques-2026-engineering-guide https://www.digitalapplied.com/blog/nvidia-dynamo-1-0-open-source-inference-os-ai-factories https://www.emergentmind.com/papers/2608.16157 https://www.glukhov.org/ai-systems/comparisons/a2a-protocol-2026-adoption https://www.gmicloud.ai/en/blog/agent-memory-architecture-working-session-and-long-term-memory-in-production https://www.interconnects.ai/p/glm-53-how-chinese-labs-keep-stride https://www.ionix.io/threat-center/cve-2026-61539 https://www.k8gb.io/blog/k8gb-cncf-incubating https://www.kodemsecurity.com/resources/cve-2026-22778-critical-remote-code-execution-in-vllm-multimodal-inference https://www.kubenatives.com/blog/how-vllm-serves-models-kubernetes https://www.kubenatives.com/p/how-vllm-serves-models-kubernetes https://www.kunalganglani.com/blog/milvus-vs-qdrant https://www.langchain.com/state-of-agent-engineering https://www.leerichtext.de/blog/llm-inference-engine-comparison https://www.linkedin.com/posts/nikolayklyagin_llm-inference-optimization-techniques-redwerk-activity-7428060973493665792-KFgt https://www.linkedin.com/posts/servergurus-india_disaggregated-inference-why-splitting-prefill-activity-7489815467243540480-FGVo https://www.linkedin.com/posts/sharada-yeluri_the-paper-challenges-and-research-directions-activity-7423387134209839104-ZJgh https://www.llmrumors.com/news/inference-kernel-performance-race-wafer-runinfra https://www.lmsys.org/blog/2026-06-15-next-generation-speculative-decoding-dflash-v2 https://www.lmsys.org/blog/2026-08-11-unified-radix-cache https://www.lmsys.org/blog/2026-08-17-advanced-cuda-graph https://www.lmsys.org/blog/2026-08-18-miles-v0-1 https://www.lmsys.org/blog/2026-08-19-deepseek-v4-pro-engine-optimization-h20 https://www.marsdevs.com/guides/agentic-rag-2026-guide https://www.mindstudio.ai/blog/what-is-google-turboquant-kv-cache-compression https://www.modular.com/blog/introducing-max-24-6-a-gpu-native-generative-ai-platform https://www.modular.com/blog/max-25-2-unleash-the-power-of-your-h200s-without-cuda https://www.modular.com/blog/mojo-open-source https://www.modular.com/blog/three-trends-from-mlsys-2026 https://www.networkbachelor.com/vllm-infrastructure-guide-pagedattention-kv-cache https://www.nvidia.com/en-us/on-demand/session/gtc26-s82033 https://www.opentrain.ai/papers/freetoken-efficient-edge-native-moe-serving-with-bandwidth-adaptive-execution--arxiv-2608.16157 https://www.opentrain.ai/papers/quantization-aware-healing-a-practical-recipe-for-recovering-compressed-4-bit-ll--arxiv-2608.20953 https://www.paralleliq.ai/blog/vllm-oom-errors-root-cause-diagnosis https://www.practical-devsecops.com/mcp-security-statistics-2026-report https://www.premai.io/blog/vllm-vs-sglang-vs-lmdeploy-fastest-llm-inference-engine-in-2026 https://premai.io/blog/vllm-vs-sglang-vs-lmdeploy-fastest-llm-inference-engine-in-2026 https://www.redhat.com/en/blog/red-hat-ai-inference-brings-llm-d-any-managed-kubernetes-starting-coreweave-and-microsoft-azure https://www.redhat.com/en/blog/red-hat-ai-tops-mlperf-inference-v60-vllm-qwen3-vl-whisper-and-gpt-oss-120b https://www.salttechno.ai/datasets/vector-database-performance-benchmark-2026 https://www.sciencedirect.com/science/article/abs/pii/S0925231226016450 https://www.sector88.co/blog/how-to-fix-vllm-oom https://www.semanticscholar.org/paper/A-Survey-on-Large-Language-Model-Acceleration-based-Li-Li/6bcd708d2e49b34f34f157daa6bf1c3e062f57c5 https://www.sentinelone.com/cybersecurity-101/cybersecurity/mcp-security https://www.sentinelone.com/vulnerability-database/cve-2026-22778 https://www.servermo.com/blogs/sglang-vs-vllm-benchmark https://www.shakudo.io/blog/deploy-ai-agents-on-kubernetes https://www.sitepoint.com/vllm-production-deployment-guide-2026 https://www.spheron.network/blog/context-engineering-production-ai-agents-kv-cache-long-context https://www.spheron.network/blog/deploy-lmcache-vllm-kv-cache-sharing-gpu-cloud https://www.spheron.network/blog/google-turboquant-llm-compression-gpu-cloud https://www.spheron.network/blog/kubernetes-gpu-orchestration-2026 https://www.spheron.network/blog/llm-inference-optimization-2026 https://www.spheron.network/blog/modular-max-mojo-gpu-cloud-llm-inference https://www.spheron.network/blog/multi-token-prediction-mtp-gpu-cloud-deployment-guide https://www.spheron.network/blog/nvidia-grove-kubernetes-disaggregated-inference-guide https://www.spheron.network/blog/prefill-decode-disaggregation-gpu-cloud https://www.spheron.network/blog/sglang-production-deployment-guide https://www.spheron.network/blog/vllm-vs-sglang https://www.spheron.network/blog/vllm-vs-sglang-2026 https://www.spheron.network/blog/vllm-vs-tensorrt-llm-vs-sglang-benchmarks https://www.sysdig.com/blog/cve-2026-33626-how-attackers-exploited-lmdeploy-llm-inference-engines-in-12-hours https://www.techtimes.com/articles/325476/20260825/ray-summit-2026-rl-post-training-forces-open-source-ai-infrastructure-converge.htm https://www.tessell.com/blog/postgresql-vector-database-with-pgvector https://www.thefuture.im/blogs/quantization-aware-healing-compressed-4-bit-model https://www.truefoundry.com/blog/sglang-vs-vllm-vs-tensorrt-llm https://www.yottalabs.ai/post/vllm-vs-tensorrt-llm-which-inference-engine-should-you-use-in-2026 https://www.youngju.dev/blog/culture/2026-05-16-database-engines-postgres-mysql-clickhouse-duckdb-tidb-cockroach-cassandra-scylla-2026-deep-dive.en https://www.youtube.com/watch?v=wgDryFT9ZXI https://x.com/vllm_project/status/2092789782464315594 https://zylos.ai/research/2026-04-03-inference-acceleration-ai-agent-loops https://arxiv.org/abs/2609.32259 https://arxiv.org/abs/2609.39929v1 https://arxiv.org/abs/2609.38987 https://arxiv.org/abs/2609.36435 https://arxiv.org/abs/2501.09223 https://arxiv.org/abs/2610.10845 https://arxiv.org/abs/2609.39334 https://arxiv.org/abs/2610.07219 https://arxiv.org/abs/2609.38169 https://huggingface.co/blog/agent-glossary https://stackoverflow.blog/2026/10/08/part-5-operating-an-llm-system-observability-cost-routing-and-the-platform-underneath https://www.storagereview.com/news/mlperf-inference-v6-1-5-7x-per-accelerator-gains-a-512-gpu-run-and-vera-rubins-first-peer-reviewed-numbers https://gpuinsights.net/mlperf-inference-v6-1-per-gpu-results-2026
v3.53 morning 沿用 NET-new URL(13 件 · v3.52 morning 587 沿用基线 + 13 NET-new = 600):
https://arxiv.org/abs/2609.12551 https://arxiv.org/html/2609.12551v1 https://arxiv.org/abs/2610.03394 https://arxiv.org/html/2610.03394v1 https://arxiv.org/pdf/2610.03394 https://arxiv.org/abs/2609.32759 https://arxiv.org/abs/2610.02772 https://arxiv.org/abs/2609.39223 https://arxiv.org/pdf/2609.39223 https://www.baseten.co/blog/agentic-inference-optimization-faster-than-sota https://mlai.qa/blog/vllm-vs-sglang-vs-tensorrt-llm https://docs.nvidia.com/nim/large-language-models/latest/deployment/model-profiles-and-selection.html https://papers.cool/arxiv/2610.03394
v3.54 morning 本轮新增 URL(7 件 · 沿用基线 600 + 7 NET-new = 607):
https://arxiv.org/abs/2609.34117 https://arxiv.org/abs/2610.08430 https://arxiv.org/abs/2609.39096 https://arxiv.org/abs/2602.09323 https://arxiv.org/abs/2610.08778 https://arxiv.org/abs/2610.07332 https://huggingface.co/papers/2609.34117
v3.55 morning 本轮新增 URL(9 件 · 沿用基线 607 + 9 NET-new = 616):
https://localaimaster.com/blog/sglang-vs-vllm-comparison https://rockybhatia.substack.com/p/how-to-learn-agentic-ai-in-2026 https://www.crusoe.ai/resources/blog/crusoe-managed-inference-optimize-performance-for-demanding-ai-workloads https://arxiv.org/abs/2501.09223 https://arxiv.org/abs/2610.10845 https://arxiv.org/abs/2609.39334 https://arxiv.org/abs/2610.07219 https://arxiv.org/abs/2609.38169 https://huggingface.co/blog/agent-glossary https://stackoverflow.blog/2026/10/08/part5-operating-an-llm-system-observability-cost-routing-and-the-platform-underneath https://www.storagereview.com/news/mlperf-inference-v6-1-5-7x-per-accelerator-gains-a-512-gpu-run-and-vera-rubins-first-peer-reviewed-numbers https://gpuinsights.net/mlperf-inference-v6-1-per-gpu-results-2026
§7 沿革摘要(精简)
[余沿革摘要详见 archive/llm-infra-changelog.md]
本次变更
2026-10-10 05:00 (Wave3 E1 §IX 113 morning v3.55)
v3.55 morning(本轮):🌅 24h 滑动净窗口(2026-10-09 18:40 CST → 2026-10-10 05:00 CST · §IX 112 morning v3.54 morning 全量沿用基线 · v3.54 morning arXiv 459/CVE 36/DOI 15/URL 607 精确闭合沿用)+ 6 NET-new arXiv ID(① arXiv:2610.10845 galahad-kv 50M Token NVMe 持久化 ⭐⭐⭐⭐ paper_card 1731 ✓ 主分类 llm-infra 10-9 入池 · vLLM + H100 + Gemma 4 12B/31B + 50M token 100/100 byte-exact + KV state 按 ~16K token 分块加密 NVMe = KV Cache 持久化纪元 + Memory-as-a-Layer 范式预备级第 1 例 ② arXiv:2501.09223 Foundations of LLMs ⭐⭐⭐ paper_card 1742 ✓ 主分类 llm-infra 10-9 入池 · 教科书型 · 六章覆盖预训练/生成/提示/对齐/推理/推理 ③ arXiv:2610.07219 Cascadia 消费级 975B MoE ⭐⭐⭐⭐ jay 10-9 14:50 inference-systems-deep-dive · 消费级 Eleven AI PC + DGX Spark + NVFP4 + 完整 artifact = 「消费级硬件 + 超大 MoE 推理」预备级第 1 例 ④ arXiv:2609.39334 SpecScale test-time scaling 投机解码 ⭐⭐⭐⭐ jay 10-9 14:50 · Qwen2.5/Llama3 4 generator-verifier 组合 + GSM8K/MATH-500/OlympiadBench + 单 A100 = 「投机解码开销 vs 收益临界条件量化」预备级第 1 例 ⑤ arXiv:2609.23130 vLLM→llm-d→控制平面演化综述 ⭐⭐⭐⭐⭐ Twinkll Sisodia Boston University · 跨 fleet 协调状态/相位/加速器/SLO · Tesla/Red Hat/KServe Llama 3.1 70B 四卡 AMD MI300X ~3× 吞吐量 + 2× TTFT 改善 = 「推理控制平面演化叙事」预备级第 1 例 ⑥ arXiv:2609.38169 STEPQuant 100▲ HF Daily ⭐⭐⭐ tom 10-9 09:00 HF Daily + stephen noon 主棒位 v71 已锚 · Delta 规则循环状态量化中何时何处出错 = 「量化失败模式分析」预备级第 1 例)+ 1 NET-new paper_card 主分类 llm-infra 入池(1731 galahad-kv 10-9 入池)+ 1 NET-new paper_card 教科书承接级(1742 Foundations of LLMs 10-9 入池)+ 9 NET-new URL + 0 NET-new CVE/DOI + 6 件承接稳态精修预备级锚定承接(① Stack Overflow Blog Part 5 LLM System Operability 6 级成熟度第 5 级 · Gateway 单一瓶颈 + decision_id 贯穿 + 跨租户隔离四原则 + Agent 平台两个失败极端 ② GMI Cloud Agent Memory Architecture 三层 Working/Session/Long-term + LLM 调用成本 10-100× 存储 + session 边界批量合并 1 次 vs 每消息合并 40 次 ③ Network Bachelor vLLM 2026 Infrastructure Engineer Guide · PagedAttention 62-80% 浪费 + KV cache 内存公式 Llama-3-70B @ BF16 4K ≈ 1.3 GB/request ④ vllm.cpp mudler C++ 推理引擎 1:1 vLLM 兼容 + CUDA EXL3 Qwen3.8-27B + DFlash2 + C ABI 26 + Vulkan TQ1_0 + ROCm AMD gfx1151 零 CPU fallback ⑤ HF Agent Glossary Harness vs Scaffold 四层定义 · Model / Scaffold / Harness / Agent · scaffold = 如何让 model 调用工具 · harness = 如何让 agent 在 loop 中运转 · OpenClaw = harness 典型实现 ⑥ MLPerf Inference v6.1 9 月 30 日发布 5.7× 单加速器 + 16/30 submitters API-centric + Edge Agentic test VLM-Interactive gpt-oss DeepSeek-R1 Llama 3.1-8B text-to-video)+ 6 矛盾/待核续写(D249 ⚠⚬⚬ galahad-kv NVMe 持久化生产环境安全验证 + D250 ⚠⚬⚬ Cascadia 消费级 MoE 部署稳定性 + D251 ⚠⚬⚬ SpecScale H100/B200 生产实测 + D252 ⚠⚬⚬ STEPQuant 循环状态量化生产引擎稳定性 + D253 ⚠⚬⚬ vLLM→llm-d→控制平面演化综述引用数据独立验证 + D254 ⚠⚬ MLPerf v6.1 与 v6.0 B200 8 卡基线对比)+ 沿用 D1-D248 + C1-C352 v3.54 morning 候选预备级 + C353-C356 v3.55 morning 候选预备级(galahad-kv KV Cache 持久化纪元 + Cascadia 消费级 975B MoE + vLLM→llm-d→控制平面演化综述 + MLPerf v6.1)+ O1-O568 + O569-O576 v3.55 morning 新增 8 件(galahad-kv 安全验证 + Foundations of LLMs 承接级 + Cascadia 跨平台验证 + EdgeAgent 工程化 + SpecScale 实测 + vLLM→llm-d 综述独立验证 + STEPQuant ROI + MiMo-V2.6 工程实现)+ P0 #1-#288 v3.54 morning 沿用 + P0 #289 v3.55 = O569 + P0-10 v3.55 morning 新增 · galahad-kv NVMe 持久化 KV Cache 在生产环境的工程化验证 + T155 v3.55 morning 新增 → arXiv/CVE/DOI/URL 459→465/36/15/607→616(精确闭合 · 6 NET-new arXiv + 9 NET-new URL)· spark 10-9 18:40 llm-infra e1prep 65.4KB(主承载 · 11 条主增量 = 1 主分类 + 1 教科书承接级 + 5 强邻接承接级 + 4 承接稳态精修预备级)+ jay 10-9 09:36 inference briefing + jay 10-9 16:20 csdn-sglang-multimodal-deep-dive 14.4KB + jay 10-9 12:21 csdn-lln-inference-rag 11.3KB + jay 10-9 13:35 github-trending 16.4KB + jay 10-9 14:52 engineering-filter-silent-failures-observability 11.1KB + jay 10-9 10:50 llmops-operability-vllm-memory-harness 18KB + jay 10-9 17:35 evening-briefing-hf-security-agents-glossary-stack2026 25KB + tom 10-9 22:20 inference-e1prep 15.8KB + flyp 10-9 18:30 risk-e1prep 51KB + stephen 10-9 noon 协调棒位 + paper_card 1731 galahad-kv 主分类 llm-infra 10-9 入池 + paper_card 1742 Foundations of LLMs 主分类 llm-infra 10-9 入池 + web_search 2026-10-09 SGLang vs vLLM 跨源双确认 + web_search 2026-10-10 MLPerf Inference v6.1 9 月 30 日发布核实 = 「v3.55 morning 在 v3.54 morning 已成型的 2026 Q4 初 llm-infra 工程化完成级 + AI Agent 优化推理系统 + MoE 阶段层路由新范式 + Agentic RL 推理支撑 + 推理引擎代际切换双栖 + 推理框架代际切换双栖 + AI for Systems 五栖预备级体系稳态基础上」+「6 NET-new arXiv(1 主分类承接级 galahad-kv 50M Token NVMe 持久化 ⭐⭐⭐⭐ + 1 教科书承接级 Foundations of LLMs + 4 强邻接承接级 Cascadia 消费级 975B MoE + SpecScale test-time scaling 投机解码 + vLLM→llm-d→控制平面演化综述 + STEPQuant 量化失败模式)」+「9 NET-new URL」+「6 承接稳态精修预备级锚定承接(LLM System Operability Part 5 + Agent Memory 三层架构 + vLLM 2026 Infrastructure Engineer Guide + vllm.cpp C++ + HF Agent Glossary + MLPerf v6.1)」+「6 矛盾续写(D249-D254)」形成 「2026 Q4 初 llm-infra 工程化完成级 + AI Agent 优化推理系统 + KV Cache 持久化纪元 + Memory-as-a-Layer 范式 + 消费级超大 MoE 推理 + 投机解码系统级条件量化 + 推理控制平面演化叙事 + LLM Operability 6 级成熟度 + Agent Memory 三层架构 + 推理引擎代际切换双栖 + 推理框架代际切换双栖 + AI for Systems 五栖 + 评测方法学 frontier lab × 第三方 × 官方赛场 × AI Agent 自动四栖预备级 + Harness vs Scaffold 术语统一」**体系预备级 → §IX 113 evening 棒位预备候选。