engineering · 知识库活文档

  • 更新:v61 FlashPrefill V2 后训练栈三联 + Centered Residual + LongStraw + Agent Harness + vLLM Conf 8-25

0. 范围与定调

本文档以 LLM 系统工程为主轴。

v61 定调(2026-08-25) = v60 主体保留 + §2.113 节段沿用 + v61 增量 1 主线 5 子件 + 旁注 4 件(§2.114 详见): - (a) FlashPrefill V2 arXiv:2608.19758 长上下文块稀疏 Prefill = 均值修正项 + PackGQA/warp specialization/pingpong pipelining + SGLang attention backend 对接 - (b) On-Policy Distillation arXiv:2607.13399「探索催化剂不提升天花板」 + SimpleOPD arXiv:2608.14277 tokenizer-agnostic + S²VOPD arXiv:2608.14144 视角反转 = 蒸馏三联 - (c) Centered Residual Signatures arXiv:2608.14929 权重空间 lineage verification + AdaPop arXiv:2608.14229 per-fact 自适应遗忘 = 谱系 + 遗忘双件 - (d) LongStraw arXiv:2607.14952 architecture-aware detach + replay = 百万 token RL 后训练栈显存墙突破 - (e) Agent Harness 闭环三件套 SkillGate arXiv:2608.18852 + CriPO arXiv:2607.18082 + Zetta ζ arXiv:2608.16590 + AID-Guard arXiv:2608.21159 stateful authorization-to-effect closure - 旁注 4 件:vLLM Conference 8-25 当日 Roadmap Q3 实际交付(Woosuk Kwon + Zachary Xi)/ SWE-bench Pro 8-25 BenchLM 创 open-weights 历史最高 Qwen3.8 Max 67.7% / HF Daily 8-23/24 立标信号梯度 / 工程实战 NVIDIA Dynamo 1.4.1 + Apple Silicon MLX-VLM + pgvectorScale 471 QPS + SGLang/vllm-ascend Atlas 800T A2 实测 11.8→45 tokens/s = 第三十八波 138 主线

v60 历史定调(2026-08-23 09:15)= vLLM Conference 8-25/26 Roadmap Q3 + SWE-bench Pro 8-15 BenchLM 立标 Mythos 5 80.3% + SWE-bench ProMax arXiv:2608.09802 + KDD/VLDB Agent Memory 顶会立项 + Foundry + AgentRx 故障恢复 + Order 66 = 第三十七波 137 主线

1. 现状全景

LLM engineering 进入「v61 后训练栈三联工程化 + Agent Harness 动作级闭环 + 谱系与遗忘合规 + 百万 token RL 后训练栈 + 长上下文推理内核 + vLLM Conference 8-25 当日 Roadmap 兑现」工程化关键拐点期。重心从「推理引擎选型」上移到「推理内核 + 后训练栈 + harness 闭环」三联工程化。

2. 关键工作与脉络

v1-v58 §2.1-§2.111 节段沿用 + §6.3 试金石锚点全部保留;v59 §2.112 / v60 §2.113 / v61 §2.114 1 主线 5 子件 + 旁注 4 件(v59 / v60 完整论证详查 changelog archive)。

2.113 v60(2026-08-23 1 主线 5 子件 + 旁注 3 件 · 详见 changelog archive v60 段)

v60 锚入 (a) vLLM Conference 8-25/26 Roadmap Q3 六轴 + (b) SWE-bench Pro 8-15 BenchLM 立标 Mythos 5 80.3% + Anthropic 三强霸榜 + Qwen3.8-27B Apache 2.0 61.7% = open-weights 首次 SWE-bench Pro 60%+ 阵营 + Terminal-Bench 2.1/3.0 + (c) SWE-bench ProMax arXiv:2608.09802 COLM 2026 + EdgeBench 134 任务 3-4h = 多文件+长周期+多工具链复合评估 + (d) KDD 2026 LinkedIn arXiv:2604.26197 + VLDB 2026 arXiv:2604.01707 + IFCMemoryBench BIM arXiv:2602.16313 = Agent Memory 顶会立项 + (e) Microsoft Foundry + AgentRx 故障恢复 = Agent 安全五件套闭环 + 旁注 3 件 Order 66 / Stamile / Stephen 8-23 news = 第三十七波 137 主线

2.114 v61(2026-08-25 1 主线 5 子件 + 旁注 4 件 · 详见 § 本次变更 v61 段)

本轮 5 主线 + 4 旁注完整论证详查 §本次变更 v61 段 + §6.3 arxiv/URL 锚点。关键摘要:(a) FlashPrefill V2 arXiv:2608.19758 = 长上下文 block-sparse 推到生产 = SGLang attention backend 直接对接 + (b) OPD 三联 arXiv:2607.13399 + 2608.14277 + 2608.14144 + RSTG 2608.00782 = 探索催化剂 + tokenizer 解耦 + 视角反转 + (c) Centered Residual arXiv:2608.14929 + AdaPop arXiv:2608.14229 = 权重溯源 + per-fact 自适应遗忘双件 + (d) LongStraw arXiv:2607.14952 = 百万 token RL 后训练栈显存墙突破 + (e) Agent Harness 闭环三件套 SkillGate + CriPO + Zetta ζ + AID-Guard arXiv:2608.21159 stateful authorization + 旁注 4 件 vLLM Conference 8-25 Roadmap Q3 当日交付 + SWE-bench Pro 8-25 BenchLM 更新 Qwen3.8 Max 67.7% + HF Daily 8-23/24 立标信号 + 工程实战 NVIDIA Dynamo 1.4.1 / Apple Silicon MLX / pgvectorScale 471 QPS / SGLang vllm-ascend Atlas 800T A2 11.8→45 tokens/s = 第三十八波 138 主线

数量更新:151→156 共识(v61 +5 C152-C156)/ 132→136 争议(v61 +4 D133-D136)/ 189→194 开放(v61 +5 O190-O194)= 156/136/194主文件压缩:v60 68KB → v61 候选 ≤80KB 完成。

3. 共识与争议(v61 156 共识 + 136 争议)

3.1 核心共识 156 条(v61 +5 共识)

v1-v60 全部保留——锚点清单见 archive/engineering-changelog.md(C1-C151)。v61 新增 5 共识

  • C152 FlashPrefill V2 arXiv:2608.19758 长上下文 block-sparse Prefill = 算法到生产 Kernel 立标 = https://arxiv.org/abs/2608.19758 + paper_cards/1039 + 共识 #31 长上下文推理内核从「单点 kernel 调优」→「引擎生态绑定」 = 均值修正项 Out_q^V2 = Σ_{k∈Supp(q)} softmax(QKᵀ)V − μ_q · V_avg_dropped 把 80% sparsity 一阶偏差归零 + PackGQA 内存访问消除 GQA bank conflict + warp specialization(producer/consumer warp 拆 SM 调度槽)+ pingpong pipelining(双 buffer 掩盖全局内存到 shared memory 延迟)+ 与 SGLang attention backend 直接对接 + FP8 per-tensor 校准 + paged KV cache + continuous batching = H20 128K FP8 对 FA-2 47.26× / 对 FA-3 dense baseline 30.49× / 对 FA-2 BF16 27.19× 预填充加速 + 32K–64K 中等长度 8×–15× 摊薄曲线 + LongBench / RULER 与 dense baseline 差距 ≤ 0.5 pp = 第三十八波 138 主线 (a)。
  • C153 OPD 三联 arXiv:2607.13399 + 2608.14277 + 2608.14144 + RSTG 2608.00782 = 蒸馏信号质量与视角反转 = https://arxiv.org/abs/2607.13399 · https://arxiv.org/abs/2608.14277 · https://arxiv.org/abs/2608.14144 · https://arxiv.org/abs/2608.00782 + paper_cards/420 · 975 · 974 · 772 + 共识 #32 后训练栈「黑盒 RLHF」→「数据侧 + 模型侧 + 合规侧」三联工程化 + OPD = 探索催化剂不提升能力天花板 + prompt 多样性比单 prompt 采样数更重要 + 信号质量 > 教师规模 + SimpleOPD tokenizer-agnostic KL(decode 双方输出回文本空间 + 学生参考 KL + 终止 token advantage 屏蔽压住长度爆炸)+ S²VOPD 视角反转(teacher 原图 / student 强增强图 = 无更强 teacher + 无 ground-truth + 无 reward 即可蒸馏 + Qwen3.5-4B 70.7%→77.4% 胜 235B Qwen3-VL)+ RSTG 负 RL 群体 sample-level/token-level 门控 + SFT 补正 + 数学 +3.02% / 代码 +3.05% = 第三十八波 138 主线 (b)。
  • C154 Centered Residual Signatures arXiv:2608.14929 + AdaPop arXiv:2608.14229 = 谱系与遗忘合规双件 = https://arxiv.org/abs/2608.14929 · https://arxiv.org/abs/2608.14229 + paper_cards/1028 · 1033 + 共识 #33 模型谱系验证「行为测试 → 数据溯源 → 权重空间 lineage」第三条路立起来 + 遗忘从「统一 NPO 梯度压力」→「流行度自适应指数 + dual-ascent 控制器」工程化 = Centered Residual 残差块中心化(投影掉与恒等映射对齐方向)+ 对称 lineage score + residual-MLP / GPT-2 族系 AUROC = 1.0 + 对 function-preserving laundering 免疫 + 76× 速度提升 + AdaPop L_forget(f) = -α(pop(f)) · E[log σ(-h_θ(f) / τ)] + dual-ascent 控制器每 epoch 观察 retain 集调 λ + paraphrase 泄露 -5× + adversarial reformulation -1.6× + 与 v60 §2.113 (d) Agent Memory 顶会立项(KDD/VLDB/IFCMemoryBench BIM)形成"权重 lineage × 内存模块化 × 语料侧范式"三维合规闭环 = 第三十八波 138 主线 (c)。
  • C155 LongStraw arXiv:2607.14952 = 百万 token RL 后训练栈显存墙突破 = https://arxiv.org/abs/2607.14952 + paper_cards/423 + 共识 #34 RL 后训练栈从「算法变体(Dr.GRPO/RLOO/DAPO)」→「执行栈改造(detach + replay)」系统端推进 = architecture-aware detach(prompt 段 detach 出 autograd + 每个 response 单独 replay + KV/compressed-attention 各自保留 architecture-aware state)+ group size 2→8 峰值分配显存只增 0.21 GB + Qwen3.6-27B 8× H20 完成 2.1M positions + GLM-5.2 compressed-attention MoE 32× H20 端到端通路 + 独立 stress test 4.46M positions + 与 v60 §2.113 (d) IFCMemoryBench BIM 内存模块化形成"运行时 KV 注入 + 跨模型 Engram + 语料侧 Build-time + RL 后训练 detach + replay"五维 Memory 杠杆 = 第三十八波 138 主线 (d)。
  • C156 Agent Harness 闭环三件套 SkillGate arXiv:2608.18852 + CriPO arXiv:2607.18082 + Zetta ζ arXiv:2608.16590 + AID-Guard arXiv:2608.21159 = 动作级在线闭环 = https://arxiv.org/abs/2608.18852 · https://arxiv.org/abs/2607.18082 · https://arxiv.org/abs/2608.16590 · https://arxiv.org/abs/2608.21159 + paper_cards/1032 · 700 · 1027 · (AID-Guard 新建) + 共识 #35 Agent Harness 从「episode 级事后反思」→「动作级在线闭环」+ 「工具调用授权安全」第六维 = SkillGate「选择器信用饥饿」诊断 + outcome 信用只到执行 token + action-local 优势只到技能命名 token 双通道不相交 + 9B 策略 5 agent benchmark 16 slate 40.8%→53.2% + 误导候选暴露砍 2/3 + CriPO on-policy 自蒸馏 + 反事实自教师定位负优势 rollout 中 criterion 相关 token + token-level 优势翻转 + 约 2× 步数达成 vanilla 终态性能 + Zetta ζ 三时间尺度解耦(动作级 / rollout 批次级 / 进化轮次级)+ frozen 底座 + 演化 {C, R, T} + LIBERO-Pro 90.8% / RoboCasa 93.6% / 11.1× 推理加速 + AID-Guard stateful authorization-to-effect closure + commit 阶段重新验证请求 + provider 状态 + delivery fence + reservation 歧义时保留 + = 与 v60 §2.113 (e) Foundry + AgentRx + v59 Inadvertent Context Leakage + v58 Bounded Agents APC + v56 Mythos 5 + v53 Stealing + v61 AID-Guard = Agent 安全六件套闭环 + 与 Self-Harness / LongHorizon-Harness / HarnessOpt-Bench / MemoHarness / ClawVM / From Prompts to Contracts / OneDayAgent 形成 harness engineering 共识级方向 = 第三十八波 138 主线 (e)。

3.2 核心争议 136 条(v61 +4 争议)

v1-v60 全部保留(D1-D132)。v61 新增 4 争议

  • D133 FlashPrefill V2 arXiv:2608.19758 H20 专属 vs 跨 GPU 厂商可移植性 = v61 (a) https://arxiv.org/abs/2608.19758 + H20 验证 47.26× + PackGQA / warp specialization 在其他 GPU(H100 / B200 / Ascend 910C)的 kernel 重优化工作量原文未公开 + 80% sparsity LongBench / RULER 测评 ≤ 0.5 pp 只覆盖一阶偏差(二阶交互如 attention head 间耦合 / KV cache block 间位置编码共享未建模)+ 美国出口管制下 H20 与 H100 / H200 / B200 算力差异可能造成「同一方法在不同硬件上加速比差距 > 5×」 + 判断:长上下文生产部署需关注 GPU 厂商 kernel 可移植性协议是否形成 + H100 / B200 实测数据待 PDF §5 公开。
  • D134 OPD 三联 + RSTG 蒸馏信号质量适用边界 = v61 (b) https://arxiv.org/abs/2607.13399 · https://arxiv.org/abs/2608.14277 · https://arxiv.org/abs/2608.14144 · https://arxiv.org/abs/2608.00782 + S²VOPD 强增强(裁剪/颜色扰动/风格化/遮挡)破坏细粒度感知中的小目标 + 6 benchmark 胜率不能推广到「需要完整结构信息」任务(chart QA、医学影像)+ SimpleOPD tokenizer-agnostic 对齐在「解码回文本空间」步骤本身存在 vocabulary mismatch 累积误差 + RSTG 数学 +3.02% / 代码 +3.05% 在跨域(科学 / 创意写作)适用性未公开 + 判断:蒸馏信号质量与教师规模分离是范式级洞察,但「信号质量」具体测量(per-token KL 分布 + reward variance)与跨域泛化是开放问题。
  • D135 Centered Residual + AdaPop 跨架构 / 跨域合规适用性 = v61 (c) https://arxiv.org/abs/2608.14929 · https://arxiv.org/abs/2608.14229 + Centered Residual 仅验证兼容 checkpoint(架构相同)+ 跨架构(encoder-decoder vs causal LLM、VLM vs LLM)未覆盖 + AdaPop 流行度代理有偏向(Wikidata sitelinks 严重偏向英文 / 历史悠久条目 + LLM-as-Judge 引入 judge 模型偏置)+ AdaPop 摘要未列 MIA 评测 = unlearning 评估需 MIA 才完整 + 判断:欧盟 AI Act GPAI 合规链与 GDPR Art. 17 删除权执法是否接受 weight-space lineage 方法作为合规证据(preprint 尚未同行评审 + 未开源)+ non-English 实体遗忘质量待补强。
  • D136 LongStraw detach + replay 与标准 GRPO 严格等价性 + AID-Guard stateful authorization 工程复杂度 = v61 (d/e) https://arxiv.org/abs/2607.14952 · https://arxiv.org/abs/2608.21159 + LongStraw 论文自承「执行 ≠ 完整训练」 + detach 后 prompt state 严格意义会损失梯度耦合 + 分布式 forward 和梯度 composition 路径尚未完整实现 + 论文未给端到端训练吞吐 vs 显存 Pareto 曲线 + 单 token 训练成本上升幅度未量化 + AID-Guard stateful authorization 协议需 commit 阶段重新验证 + reservation 机制在多工具链 + 第三方 API + 高并发场景下的状态机复杂度 + delivery fence 在跨服务 / 跨区域部署下的时钟同步与重试幂等性 + 判断:百万 token RL 训练端到端正确性(detach 后梯度偏差)何时被系统化验证是开放问题 + Agent 工具调用授权安全工程落地需配套 reservation pool + delivery fence 库 + 与 OpenAI / Anthropic 协议层融合路径待观察。

4. 开放问题(194 条 · v61 +5 新增)

O1-O189 = v1-v60 累积 + O190 v61 §2.114 (a-e) 5 子件跟进 + O191 FlashPrefill V2 H100 / B200 实测加速比与 80% sparsity 二阶偏差测评 + O192 OPD 三联跨域泛化(科学 / 创意写作 / chart QA)+ SimpleOPD tokenizer mismatch 累积误差 + O193 Centered Residual 跨架构方法 + AdaPop MIA 评测 + Wikidata non-English 流行度代理 + O194 LongStraw 端到端训练 vs 标准 GRPO 梯度偏差 + AID-Guard 多工具链 reservation 复杂度 —— 清单见 archive/engineering-changelog.md。

5. 趋势判断(v1-v60 累积)

5.1 短期趋势(2026 H2 · 6-12 个月)

  1. v61(详见 §2.114 ⭐⭐⭐⭐⭐)= (a) FlashPrefill V2 arXiv:2608.19758 长上下文 block-sparse Prefill + (b) OPD 三联 arXiv:2607.13399 + 2608.14277 + 2608.14144 + RSTG 2608.00782 蒸馏信号质量与视角反转 + (c) Centered Residual arXiv:2608.14929 + AdaPop arXiv:2608.14229 谱系与遗忘合规双件 + (d) LongStraw arXiv:2607.14952 百万 token RL 后训练栈显存墙突破 + (e) Agent Harness 闭环三件套 SkillGate + CriPO + Zetta ζ + AID-Guard arXiv:2608.21159 stateful authorization + 旁注 4 件 (vLLM Conference 8-25 当日 / SWE-bench Pro 8-25 BenchLM Qwen3.8 Max 67.7% / HF Daily 8-23/24 / 工程实战 NVIDIA Dynamo 1.4.1 + Apple Silicon MLX + pgvectorScale 471 QPS + SGLang/vllm-ascend) = 第三十八波 138 主线

  2. v60(详见 §2.113 ⭐⭐⭐⭐⭐)= (a) vLLM Conference 8-25/26 Roadmap Q3 六轴 + (b) SWE-bench Pro 8-15 Mythos 5 80.3% 立标 + (c) SWE-bench ProMax arXiv:2608.09802 + EdgeBench + (d) KDD/VLDB Agent Memory 顶会立项 + (e) Microsoft Foundry + AgentRx 故障恢复 + 旁注 3 件 (Order 66 / Stamile / Stephen 8-23 news) = 第三十七波 137 主线

  3. v59(详见 §2.112 ⭐⭐⭐⭐⭐)= InferScale + 47billion CUDA + SageMaker Metrics + Antigravity 2.0 + SWE-bench Science + Chain-of-Experience + SkillEvo + Repo0 + Inadvertent Context Leakage = 第三十六波 136 主线

  4. v58(详见 §2.111 ⭐⭐⭐⭐⭐)= SkillGate + Bounded Agents APC + CTIFoundry + SemaPLC + Zetta ζ + CoRun + DASH + SWE-bench Pro + HBF Sucks + PIM-DIMM + MCP 安全审计 + Miles v0.1 + Cross-Model Memory Transfer = 第三十五波 135 主线

  5. v57(详见 §2.110 ⭐⭐⭐⭐⭐)= StateM Harness Scaling 95.3%/$15 + 374▲ / ACID-Agent KramaBench +10.6pp / Agent Skills 12 模式 4 因素 / Mojo🔥 1.0 开源 / Albireo 1.7× / Agent RL 双件套 / FreeToken 边缘 MoE = 第三十四波 134 主线

  6. v56(详见 §2.109 ⭐⭐⭐⭐⭐)= StateM 130▲ + ACID-Agent 清华 6 域 + MOSS-VL 实时交互 VL + GRIP/IGD RAG 双件套 = 第三十三波 133 主线

  7. v55(详见 §2.108 ⭐⭐⭐⭐⭐)= OpScale + vToken + Agentic Transaction ACID + VLDB 2026 DBCooker + Agentic Memory 综述 = 第三十二波 132 主线

6.3 完整 arXiv / CVE / DOI / URL 锚点清单

v61 主文件增量更新:沿用 v54-v60 锚点 + 新增 v61 §2.114 全部 arXiv/CVE/DOI/URL 锚点。v61 主文件候选 ≤80KB 上限完成(保留 §6.3 完整锚点以满足 missing=0 校验)。

arXiv v61 新增清单(v61 §2.114 净增 9 件):arXiv:2608.19758 · arXiv:2607.13399 · arXiv:2608.14277 · arXiv:2608.14144 · arXiv:2608.00782 · arXiv:2608.14929 · arXiv:2608.14229 · arXiv:2607.14952 · arXiv:2608.21159

URL v61 新增清单(v61 §2.114 净增 ~14 个核心 URL):https://arxiv.org/abs/2608.19758 · https://arxiv.org/abs/2607.13399 · https://arxiv.org/abs/2608.14277 · https://arxiv.org/abs/2608.14144 · https://arxiv.org/abs/2608.00782 · https://arxiv.org/abs/2608.14929 · https://arxiv.org/abs/2608.14229 · https://arxiv.org/abs/2607.14952 · https://arxiv.org/abs/2608.21159 · https://benchlm.ai/benchmarks/swe-bench-pro · https://github.com/vllm-project/vllm/issues/48168 · https://www.youtube.com/watch?v=f_EAJiUWlbU · https://catalog.ngc.nvidia.com/orgs/nvidia/ai-dynamo/containers/vllm-runtime/1.0.0-cuda13 · https://huggingface.co/blog/multi-vector-encoder

arXiv v60 新增清单(v60 §2.113 净增 5 件):arXiv:2604.26197 · arXiv:2604.01707 · arXiv:2602.16313 · arXiv:2608.09802 · arXiv:2608.08131

URL v60 新增清单(v60 §2.113 净增 ~12 个核心 URL):https://github.com/vllm-project/vllm/issues/48168 · https://vllm.ai/events/vllm-conference/2026 · https://benchlm.ai/benchmarks/swe-bench-pro · https://codingfleet.com/blog/swe-bench-pro-leaderboard-2026 · https://localaimaster.com/models/swe-bench-explained-ai-benchmarks · https://arxiv.org/pdf/2608.09802 · https://github.com/RUC-NLPIR/Awesome-Long-Horizon-Agents · https://arxiv.org/abs/2604.26197 · https://arxiv.org/abs/2604.01707 · https://arxiv.org/abs/2602.16313 · https://arxiv.org/abs/2608.08131 · https://open.substack.com/pub/claudiostamile/p/agent-memory-is-not-rag-a-practical · https://www.livemint.com/technology/tech-news/chatgpt-to-get-a-massive-upgrade-in-the-next-6-months-says-sam-altman-could-watch-you-all-the-time-11787068716398.html

arXiv v58 新增清单(v58 §2.111 净增 8 件):arXiv:2608.18852 · arXiv:2608.15888 · arXiv:2608.18613 · arXiv:2608.18565 · arXiv:2608.16590 · arXiv:2608.14376 · arXiv:2608.14333 · arXiv:2608.11668 · arXiv:2608.13558 · arXiv:2608.17253 · arXiv:2608.18489

arXiv v57 沿用清单(v57 §2.110 净增 11 件):arXiv:2608.14036 · arXiv:2608.17310 · arXiv:2608.17528 · arXiv:2608.16157 · arXiv:2606.01927 · arXiv:2608.15984 · arXiv:2608.15669 · arXiv:2608.17950 · arXiv:2608.17536 · arXiv:2608.17960 · arXiv:2608.17050

arXiv v56 沿用清单(v56 §2.109 净增 9 件):arXiv:2608.15089 · arXiv:2608.15045 · arXiv:2608.13900 · arXiv:2608.16776 · arXiv:2608.16515 · arXiv:2608.16536 · arXiv:2608.16628 · arXiv:2608.02870 · arXiv:2606.30391

arXiv v55 沿用清单(完整保留 v54-v55 等价集合,共 ~554 件 ID 已锚入 §2.1-§2.108 节段纲要):arXiv:1301.6707 · arXiv:1307.4186 · arXiv:1702.08608 · arXiv:1707.08114 · arXiv:1908.10454 · arXiv:1912.08777 · arXiv:2002.05651 · arXiv:2004.05074 · arXiv:2106.01345 · arXiv:2204.05862 · arXiv:2206.14858 · arXiv:2208.03299 · arXiv:2210.11416 · arXiv:2211.01910 · arXiv:2301.12597 · arXiv:2302.11382 · arXiv:2303.10130 · arXiv:2305.06500 · arXiv:2311.12871 · arXiv:2403.05527 · arXiv:2409.10102 · arXiv:2501.01005 · arXiv:2502.07115 · arXiv:2502.14617 · arXiv:2503.01840 · arXiv:2504.11320 · arXiv:2504.12330 · arXiv:2504.19874 · arXiv:2505.17152 · arXiv:2506.04301 · arXiv:2506.12071 · arXiv:2506.13114 · arXiv:2506.21901 · arXiv:2507.06457 · arXiv:2507.06608 · arXiv:2507.07471 · arXiv:2507.21504 · arXiv:2508.05294 · arXiv:2508.08438 · arXiv:2508.09442 · arXiv:2509.12384 · arXiv:2510.08544 · arXiv:2510.09665 · arXiv:2510.13668 · arXiv:2510.15253 · arXiv:2511.01815 · arXiv:2511.02230 · arXiv:2511.02248 · arXiv:2511.16681 · arXiv:2511.17593 · arXiv:2512.09196 · arXiv:2601.00227 · arXiv:2601.01937 · arXiv:2601.06288 · arXiv:2601.06456 · arXiv:2601.07504 · arXiv:2601.09822 · arXiv:2601.11960 · arXiv:2601.15727 · arXiv:2601.16432 · arXiv:2601.16503 · arXiv:2601.18591 · arXiv:2601.19139 · arXiv:2601.19827 · arXiv:2601.20309 · arXiv:2602.00238 · arXiv:2602.01873 · arXiv:2602.02007 · arXiv:2602.03442 · arXiv:2602.03786 · arXiv:2602.04900 · arXiv:2602.06036 · arXiv:2602.07584 · arXiv:2602.08005 · arXiv:2602.08226 · arXiv:2602.10238 · arXiv:2602.11443 · arXiv:2602.11510 · arXiv:2602.14516 · arXiv:2602.16603 · arXiv:2602.16666 · arXiv:2602.16873 · arXiv:2602.18998 · arXiv:2602.19594 · arXiv:2602.20478 · arXiv:2602.21548 · arXiv:2602.21566 · arXiv:2602.21626 · arXiv:2602.22593 · arXiv:2602.23368 · arXiv:2602.23571 · arXiv:2603.00026 · arXiv:2603.03251 · arXiv:2603.04428 · arXiv:2603.05162 · arXiv:2603.05451 · arXiv:2603.07379 · arXiv:2603.07670 · arXiv:2603.08739 · arXiv:2603.09619 · arXiv:2603.10726 · arXiv:2603.11622 · arXiv:2603.11768 · arXiv:2603.12707 · arXiv:2603.13358 · arXiv:2603.16104 · arXiv:2603.16877 · arXiv:2603.20397 · arXiv:2603.20847 · arXiv:2603.21354 · arXiv:2603.22774 · arXiv:2603.23710 · arXiv:2603.26498 · arXiv:2603.26670 · arXiv:2603.29010 · arXiv:2603.29231 · arXiv:2604.00499 · arXiv:2604.00901 · arXiv:2604.01395 · arXiv:2604.01437 · arXiv:2604.01647 · arXiv:2604.03143 · arXiv:2604.04035 · arXiv:2604.04722 · arXiv:2604.05012 · arXiv:2604.09048 · arXiv:2604.09666 · arXiv:2604.11320 · arXiv:2604.11623 · arXiv:2604.11943 · arXiv:2604.12452 · arXiv:2604.15186 · arXiv:2604.15464 · arXiv:2604.15732 · arXiv:2604.16371 · arXiv:2604.17227 · arXiv:2604.19157 · arXiv:2604.19769 · arXiv:2604.25724 · arXiv:2604.25850 · arXiv:2604.25899 · arXiv:2604.26557 · arXiv:2605.00616 · arXiv:2605.01280 · arXiv:2605.01604 · arXiv:2605.01920 · arXiv:2605.02189 · arXiv:2605.03275 · arXiv:2605.03375 · arXiv:2605.04595 · arXiv:2605.06716 · arXiv:2605.08717 · arXiv:2605.08962 · arXiv:2605.10834 · arXiv:2605.10907 · arXiv:2605.11032 · arXiv:2605.11202 · arXiv:2605.11581 · arXiv:2605.11733 · arXiv:2605.14678 · arXiv:2605.16867 · arXiv:2605.19537 · arXiv:2605.19743 · arXiv:2605.19893 · arXiv:2605.20173 · arXiv:2605.20466 · arXiv:2605.23215 · arXiv:2605.24217 · arXiv:2605.25480 · arXiv:2605.27445 · arXiv:2605.27492 · arXiv:2605.27744 · arXiv:2605.29640 · arXiv:2606.00610 · arXiv:2606.00644 · arXiv:2606.01581 · arXiv:2606.01927 · arXiv:2606.02643 · arXiv:2606.02964 · arXiv:2606.04301 · arXiv:2606.04594 · arXiv:2606.05608 · arXiv:2606.05901 · arXiv:2606.06036 · arXiv:2606.06090 · arXiv:2606.06324 · arXiv:2606.06535 · arXiv:2606.07001 · arXiv:2606.07362 · arXiv:2606.07665 · arXiv:2606.07923 · arXiv:2606.08671 · arXiv:2606.08950 · arXiv:2606.09441 · arXiv:2606.10749 · arXiv:2606.12329 · arXiv:2606.13141 · arXiv:2606.13175 · arXiv:2606.13643 · arXiv:2606.14127 · arXiv:2606.14470 · arXiv:2606.14589 · arXiv:2606.15376 · arXiv:2606.16135 · arXiv:2606.16494 · arXiv:2606.16661 · arXiv:2606.16903 · arXiv:2606.17915 · arXiv:2606.18192 · arXiv:2606.19803 · arXiv:2606.20295 · arXiv:2606.20905 · arXiv:2606.23687 · arXiv:2606.24775 · arXiv:2606.25393 · arXiv:2606.26080 · arXiv:2606.26458 · arXiv:2606.26959 · arXiv:2606.27192 · arXiv:2606.27288 · arXiv:2606.28393 · arXiv:2606.28565 · arXiv:2606.28733 · arXiv:2606.29526 · arXiv:2606.29600 · arXiv:2606.30391 · arXiv:2606.31145 · arXiv:2606.31227 · arXiv:2607.00406 · arXiv:2607.01579 · arXiv:2607.01812 · arXiv:2607.01831 · arXiv:2607.02401 · arXiv:2607.02574 · arXiv:2607.02980 · arXiv:2607.03065 · arXiv:2607.03723 · arXiv:2607.03949 · arXiv:2607.04395 · arXiv:2607.04434 · arXiv:2607.04617 · arXiv:2607.04690 · arXiv:2607.04763 · arXiv:2607.05061 · arXiv:2607.05294 · arXiv:2607.05382 · arXiv:2607.05394 · arXiv:2607.05708 · arXiv:2607.05876 · arXiv:2607.06403 · arXiv:2607.06519 · arXiv:2607.06624 · arXiv:2607.07119 · arXiv:2607.07386 · arXiv:2607.07508 · arXiv:2607.07534 · arXiv:2607.07696 · arXiv:2607.07816 · arXiv:2607.07820 · arXiv:2607.07858 · arXiv:2607.07953 · arXiv:2607.08028 · arXiv:2607.08057 · arXiv:2607.08395 · arXiv:2607.08404 · arXiv:2607.08495 · arXiv:2607.08565 · arXiv:2607.08716 · arXiv:2607.08765 · arXiv:2607.08973 · arXiv:2607.09082 · arXiv:2607.09172 · arXiv:2607.09248 · arXiv:2607.09415 · arXiv:2607.09661 · arXiv:2607.09686 · arXiv:2607.09701 · arXiv:2607.10017 · arXiv:2607.10169 · arXiv:2607.10183 · arXiv:2607.10371 · arXiv:2607.10389 · arXiv:2607.11149 · arXiv:2607.11172 · arXiv:2607.11423 · arXiv:2607.11505 · arXiv:2607.11523 · arXiv:2607.11881 · arXiv:2607.11886 · arXiv:2607.12227 · arXiv:2607.12340 · arXiv:2607.12395 · arXiv:2607.12406 · arXiv:2607.12756 · arXiv:2607.13027 · arXiv:2607.13034 · arXiv:2607.13104 · arXiv:2607.13276 · arXiv:2607.13399 · arXiv:2607.13705 · arXiv:2607.13988 · arXiv:2607.14159 · arXiv:2607.14187 · arXiv:2607.14277 · arXiv:2607.14530 · arXiv:2607.14541 · arXiv:2607.14777 · arXiv:2607.14935 · arXiv:2607.14952 · arXiv:2607.14958 · arXiv:2607.15257 · arXiv:2607.15330 · arXiv:2607.16859 · arXiv:2607.17247 · arXiv:2607.17524 · arXiv:2607.17715 · arXiv:2607.17979 · arXiv:2607.18039 · arXiv:2607.18082 · arXiv:2607.18213 · arXiv:2607.18225 · arXiv:2607.18314 · arXiv:2607.19058 · arXiv:2607.19215 · arXiv:2607.19712 · arXiv:2607.20346 · arXiv:2607.20465 · arXiv:2607.20510 · arXiv:2607.20957 · arXiv:2607.21324 · arXiv:2607.21503 · arXiv:2607.21557 · arXiv:2607.21653 · arXiv:2607.21848 · arXiv:2607.21962 · arXiv:2607.22042 · arXiv:2607.22043 · arXiv:2607.22157 · arXiv:2607.22375 · arXiv:2607.22389 · arXiv:2607.22757 · arXiv:2607.23693 · arXiv:2607.23782 · arXiv:2607.23783 · arXiv:2607.23802 · arXiv:2607.24000 · arXiv:2607.24223 · arXiv:2607.24352 · arXiv:2607.24368 · arXiv:2607.24475 · arXiv:2607.24554 · arXiv:2607.24593 · arXiv:2607.24663 · arXiv:2607.24720 · arXiv:2607.24882 · arXiv:2607.24904 · arXiv:2607.24957 · arXiv:2607.25308 · arXiv:2607.25379 · arXiv:2607.25380 · arXiv:2607.25398 · arXiv:2607.25431 · arXiv:2607.25498 · arXiv:2607.25537 · arXiv:2607.25600 · arXiv:2607.25659 · arXiv:2607.25895 · arXiv:2607.25996 · arXiv:2607.26314 · arXiv:2607.26451 · arXiv:2607.26475 · arXiv:2607.26520 · arXiv:2607.26571 · arXiv:2607.26611 · arXiv:2607.26627 · arXiv:2607.26637 · arXiv:2607.26784 · arXiv:2607.26991 · arXiv:2607.27042 · arXiv:2607.27090 · arXiv:2607.27136 · arXiv:2607.27201 · arXiv:2607.27600 · arXiv:2607.27616 · arXiv:2607.27816 · arXiv:2607.27888 · arXiv:2607.27919 · arXiv:2607.27945 · arXiv:2607.27958 · arXiv:2607.28126 · arXiv:2607.28227 · arXiv:2607.28229 · arXiv:2607.28415 · arXiv:2607.28509 · arXiv:2607.28595 · arXiv:2607.28617 · arXiv:2607.28633 · arXiv:2607.28675 · arXiv:2607.28802 · arXiv:2607.29167 · arXiv:2607.29209 · arXiv:2607.29377 · arXiv:2607.29402 · arXiv:2607.29405 · arXiv:2607.29459 · arXiv:2607.29591 · arXiv:2607.29610 · arXiv:2607.29677 · arXiv:2607.29678 · arXiv:2607.29679 · arXiv:2607.29684 · arXiv:2608.00101 · arXiv:2608.00303 · arXiv:2608.00677 · arXiv:2608.00742 · arXiv:2608.00782 · arXiv:2608.00799 · arXiv:2608.00881 · arXiv:2608.00902 · arXiv:2608.01247 · arXiv:2608.01285 · arXiv:2608.01492 · arXiv:2608.01526 · arXiv:2608.01651 · arXiv:2608.01678 · arXiv:2608.01735 · arXiv:2608.01862 · arXiv:2608.01964 · arXiv:2608.02143 · arXiv:2608.02162 · arXiv:2608.02515 · arXiv:2608.02580 · arXiv:2608.02583 · arXiv:2608.02645 · arXiv:2608.02672 · arXiv:2608.02703 · arXiv:2608.02738 · arXiv:2608.02870 · arXiv:2608.02989 · arXiv:2608.03036 · arXiv:2608.03451 · arXiv:2608.03463 · arXiv:2608.03487 · arXiv:2608.03499 · arXiv:2608.03506 · arXiv:2608.03700 · arXiv:2608.03744 · arXiv:2608.03764 · arXiv:2608.03796 · arXiv:2608.03874 · arXiv:2608.03887 · arXiv:2608.03972 · arXiv:2608.03994 · arXiv:2608.04003 · arXiv:2608.04505 · arXiv:2608.04530 · arXiv:2608.04964 · arXiv:2608.05042 · arXiv:2608.05076 · arXiv:2608.05108 · arXiv:2608.05137 · arXiv:2608.05138 · arXiv:2608.05219 · arXiv:2608.05369 · arXiv:2608.05424 · arXiv:2608.05565 · arXiv:2608.05604 · arXiv:2608.05747 · arXiv:2608.05784 · arXiv:2608.05987 · arXiv:2608.06020 · arXiv:2608.06033 · arXiv:2608.06060 · arXiv:2608.06113 · arXiv:2608.06130 · arXiv:2608.06197 · arXiv:2608.06216 · arXiv:2608.06257 · arXiv:2608.06301 · arXiv:2608.06305 · arXiv:2608.06501 · arXiv:2608.06729 · arXiv:2608.06751 · arXiv:2608.06790 · arXiv:2608.06865 · arXiv:2608.06867 · arXiv:2608.07009 · arXiv:2608.07051 · arXiv:2608.07126 · arXiv:2608.07152 · arXiv:2608.07193 · arXiv:2608.07458 · arXiv:2608.07468 · arXiv:2608.07545 · arXiv:2608.08020 · arXiv:2608.08097 · arXiv:2608.08119 · arXiv:2608.08131 · arXiv:2608.08285 · arXiv:2608.08627 · arXiv:2608.08722 · arXiv:2608.08814 · arXiv:2608.09802 · arXiv:2608.09853 · arXiv:2608.09867 · arXiv:2608.09888 · arXiv:2608.09900 · arXiv:2608.10288 · arXiv:2608.10450 · arXiv:2608.10538 · arXiv:2608.10628 · arXiv:2608.10812 · arXiv:2608.10835 · arXiv:2608.10875 · arXiv:2608.10915 · arXiv:2608.11030 · arXiv:2608.11205 · arXiv:2608.11341 · arXiv:2608.11350 · arXiv:2608.11367 · arXiv:2608.11632 · arXiv:2608.11660 · arXiv:2608.11668 · arXiv:2608.11745 · arXiv:2608.11752 · arXiv:2608.12036 · arXiv:2608.12123 · arXiv:2608.12304 · arXiv:2608.12313 · arXiv:2608.12440 · arXiv:2608.12700 · arXiv:2608.12875 · arXiv:2608.12915 · arXiv:2608.13010 · arXiv:2608.13049 · arXiv:2608.13120 · arXiv:2608.13122 · arXiv:2608.13160 · arXiv:2608.13210 · arXiv:2608.13237 · arXiv:2608.13263 · arXiv:2608.13391 · arXiv:2608.13410 · arXiv:2608.13417 · arXiv:2608.13426 · arXiv:2608.13489 · arXiv:2608.13499 · arXiv:2608.13517 · arXiv:2608.13547 · arXiv:2608.13555 · arXiv:2608.13558 · arXiv:2608.13606 · arXiv:2608.13667 · arXiv:2608.13696 · arXiv:2608.13757 · arXiv:2608.13760 · arXiv:2608.13900 · arXiv:2608.13987 · arXiv:2608.14022 · arXiv:2608.14036 · arXiv:2608.14054 · arXiv:2608.14106 · arXiv:2608.14144 · arXiv:2608.14210 · arXiv:2608.14229 · arXiv:2608.14277 · arXiv:2608.14284 · arXiv:2608.14290 · arXiv:2608.14333 · arXiv:2608.14376 · arXiv:2608.14391 · arXiv:2608.14528 · arXiv:2608.14530 · arXiv:2608.14546 · arXiv:2608.14929 · arXiv:2608.15022 · arXiv:2608.15045 · arXiv:2608.15089 · arXiv:2608.15522 · arXiv:2608.15669 · arXiv:2608.15888 · arXiv:2608.15984 · arXiv:2608.16157 · arXiv:2608.16168 · arXiv:2608.16256 · arXiv:2608.16323 · arXiv:2608.16328 · arXiv:2608.16515 · arXiv:2608.16536 · arXiv:2608.16590 · arXiv:2608.16628 · arXiv:2608.16776 · arXiv:2608.16859 · arXiv:2608.17050 · arXiv:2608.17253 · arXiv:2608.17310 · arXiv:2608.17426 · arXiv:2608.17528 · arXiv:2608.17536 · arXiv:2608.17950 · arXiv:2608.17960 · arXiv:2608.18027 · arXiv:2608.18184 · arXiv:2608.18489 · arXiv:2608.18565 · arXiv:2608.18580 · arXiv:2608.18613 · arXiv:2608.18852 · arXiv:2608.18940 · arXiv:2608.19197 · arXiv:2608.19758 · arXiv:2608.19799 · arXiv:2608.19854 · arXiv:2608.19857 · arXiv:2608.19863 · arXiv:2608.19880 · arXiv:2608.19936 · arXiv:2608.20169 · arXiv:2608.20202 · arXiv:2608.20246 · arXiv:2608.20319 · arXiv:2608.20331 · arXiv:2608.20335 · arXiv:2608.20336 · arXiv:2608.20338 · arXiv:2608.20438 · arXiv:2608.21159 · arXiv:2608.21208 · arXiv:2608.21252 · arXiv:2602.14617 · arXiv:2607.23933 · arXiv:2607.24062

CVE v55 沿用清单(完整保留 v54-v55 等价集合,共 25 件 ID):CVE-2025-32711 · CVE-2025-53109 · CVE-2025-53773 · CVE-2026-1642 · CVE-2026-20805 · CVE-2026-22778 · CVE-2026-26029 · CVE-2026-26209 · CVE-2026-27651 · CVE-2026-27654 · CVE-2026-3059 · CVE-2026-3060 · CVE-2026-30623 · CVE-2026-30624 · CVE-2026-3172 · CVE-2026-32647 · CVE-2026-3989 · CVE-2026-4372 · CVE-2026-4810 · CVE-2026-5059 · CVE-2026-5241 · CVE-2026-54769 · CVE-2026-57572 · CVE-2026-59726 · CVE-2026-61447

DOI v55 沿用清单(完整保留 v54-v55 等价集合,共 6 件 DOI):DOI 10.1145/3749168 · DOI 10.1145/3770762.3772592 · DOI 10.1145/38020094 · DOI 10.1145/38020284 · DOI 10.1145/3802513.3803486 · DOI 10.3389/fcomp.2026.1819991

URL v58 新增清单(v58 §2.111 净增,~6 个核心 URL):https://arxiv.org/abs/2608.18852 · https://arxiv.org/abs/2608.15888 · https://arxiv.org/abs/2608.18613 · https://arxiv.org/abs/2608.18565 · https://arxiv.org/abs/2608.16590 · https://arxiv.org/abs/2608.14376 · https://arxiv.org/abs/2608.14333 · https://arxiv.org/abs/2608.11668 · https://hotinfra.org/2026/papers/hotinfra26-final59.pdf · https://x.com/HuggingPapers/status/2089806276045754554 · https://huggingface.co/papers/2608.15089 · https://aiweekly.co/editors-blog/found-first-statem-hits-95-3-on-terminal-bench-2-1-for-15-via-harness-scaling · https://henryqin1997.github.io/statem/ · https://github.com/henryqin1997/statem · https://paperswithcode.co/paper/2608.15089 · https://www.lmsys.org/blog/2026-08-18-miles-v0-1 · https://github.com/RyanAlberts/best-of-Agent-Harnesses

URL v57 沿用清单(v57 §2.110 净增,~11 个核心 URL):https://arxiv.org/abs/2608.14036 · https://arxiv.org/abs/2608.17310 · https://arxiv.org/abs/2608.17528 · https://arxiv.org/abs/2608.16157 · https://arxiv.org/abs/2606.01927 · https://arxiv.org/abs/2608.15984 · https://arxiv.org/abs/2608.15669 · https://arxiv.org/abs/2608.17950 · https://arxiv.org/abs/2608.17536 · https://arxiv.org/abs/2608.17960 · https://arxiv.org/abs/2608.17050

URL v56 沿用清单(v56 §2.109 净增,~9 个核心 URL):https://arxiv.org/abs/2608.15089 · https://arxiv.org/abs/2608.15045 · https://arxiv.org/abs/2608.13900 · https://arxiv.org/abs/2608.16776 · https://arxiv.org/abs/2608.16515 · https://arxiv.org/abs/2608.16536 · https://arxiv.org/abs/2608.16628 · https://arxiv.org/abs/2608.02870 · https://arxiv.org/abs/2606.30391

arXiv v59 新增清单(v59 §2.112 净增 13 件):arXiv:2607.27090 · arXiv:2608.19799 · arXiv:2608.18027 · arXiv:2608.13120 · arXiv:2608.19854 · arXiv:2608.19857 · arXiv:2608.12875 · arXiv:2608.08466 · arXiv:2608.13547 · arXiv:2607.21596 · arXiv:2608.20246 · arXiv:2608.19880 · arXiv:2608.19197 · arXiv:2608.20202

URL v59 新增清单(v59 §2.112 净增 ~10 个核心 URL):https://arxiv.org/abs/2607.27090 · https://arxiv.org/abs/2608.19799 · https://arxiv.org/abs/2608.18027 · https://arxiv.org/abs/2608.13120 · https://arxiv.org/abs/2608.19854 · https://arxiv.org/abs/2608.19857 · https://arxiv.org/abs/2608.12875 · https://arxiv.org/abs/2608.08466 · https://arxiv.org/abs/2608.13547 · https://arxiv.org/abs/2607.21596 · https://arxiv.org/abs/2608.20246 · https://arxiv.org/abs/2608.19880 · https://arxiv.org/abs/2608.19197 · https://arxiv.org/abs/2608.20202 · https://47billion.com/blog/custom-cuda-kernels-in-the-age-of-ai-coding-agents-inside-the-new-agent-skill-workflow-for-gpu-kernel-engineering · https://docs.aws.amazon.com/sagemaker/latest/dg/monitoring-cloudwatch-detailed-observability.html · https://www.morphllm.com/best-ai-coding-agents-2026 · https://www.spheron.network/blog/llm-inference-optimization-2026 · https://theaiengineer.substack.com/p/the-ai-agents-stack-2026-edition

URL v55 完整沿用清单(完整保留 v54-v58 等价集合,共 516 URL,详细索引见 v55 主文件 archive)· v61 净增 URL 含:https://www.youtube.com/watch?v=f_EAJiUWlbU · https://catalog.ngc.nvidia.com/orgs/nvidia/ai-dynamo/containers/vllm-runtime/1.0.0-cuda13 · https://huggingface.co/blog/multi-vector-encoder · https://www.beri.net/article/vllm-vs-tensorrt-llm-vs-sglang-inference-runtime-2026 · https://blog.csdn.net/weixin_29035147 · https://blog.csdn.net/gitblog_00178 · https://www.premai.io/blog/vllm-vs-sglang-vs-lmdeploy-fastest-llm-inference-engine-in-2026 · https://alphasignalai.substack.com/p/rag-and-long-context-arent-enough · https://github.com/AMAP-ML/LongHorizon-Harness · https://pgbot.dev · https://github.com/amitshekhariitbhu/llm-inference-engineering · https://github.com/rome-os/rome · https://github.com/browser-use/macos-harness · https://github.com/0xsline/awesome-deepseek-harness

本次变更

v61(2026-08-25 09:15 · Jay Wave3 E1 综合)

v60 主体保留 + 151/132/189 全部保留。v61 在 v60 基础上叠加 1 主线 5 子件 + 旁注 4 件(§2.114)= (a) FlashPrefill V2 arXiv:2608.19758 长上下文 block-sparse Prefill = 算法到生产 Kernel = 均值修正项 Out_q^V2 = Σ_{k∈Supp(q)} softmax(QKᵀ)V − μ_q · V_avg_dropped 把 80% sparsity 一阶偏差归零 + PackGQA 内存访问消除 GQA bank conflict + warp specialization(producer/consumer warp 拆 SM 调度槽)+ pingpong pipelining(双 buffer 掩盖全局内存到 shared memory 延迟)+ 与 SGLang attention backend 直接对接 + FP8 per-tensor 校准 + paged KV cache + continuous batching = H20 128K FP8 对 FA-2 47.26× / 对 FA-3 dense baseline 30.49× / 对 FA-2 BF16 27.19× 预填充加速 + 32K–64K 中等长度 8×–15× 摊薄曲线 + LongBench / RULER 与 dense baseline 差距 ≤ 0.5 pp = (b) OPD 三联 arXiv:2607.13399 + 2608.14277 + 2608.14144 + RSTG 2608.00782 = 蒸馏信号质量与视角反转 = On-Policy Distillation「探索催化剂不提升能力天花板 + prompt 多样性比单 prompt 采样数更重要 + 信号质量 > 教师规模」+ Student-Teacher Mismatch + Length Exploitation 两类病态诊断 + advantage clipping + log-scale compression 抑制 + SimpleOPD tokenizer-agnostic KL(decode 双方输出回文本空间 + 学生参考 KL + 终止 token advantage 屏蔽)+ S²VOPD 视角反转(teacher 原图 / student 强增强图 = 无更强 teacher + 无 ground-truth + 无 reward 即可蒸馏 + Qwen3.5-4B 70.7%→77.4% 胜 235B Qwen3-VL)+ RSTG 负 RL 群体 sample-level/token-level 门控 + SFT 补正恢复信号 + 数学 +3.02% / 代码 +3.05% = (c) Centered Residual Signatures arXiv:2608.14929 + AdaPop arXiv:2608.14229 = 谱系与遗忘合规双件 = Centered Residual 残差块中心化(投影掉与恒等映射对齐方向)+ 对称 lineage score + residual-MLP / GPT-2 族系 AUROC = 1.0 + 对 function-preserving laundering 免疫 + 76× 速度提升 + 跨六族泛化 + AdaPop L_forget(f) = -α(pop(f)) · E[log σ(-h_θ(f) / τ)] + dual-ascent 控制器每 epoch 观察 retain 集调 λ + paraphrase 泄露 -5× + adversarial reformulation -1.6× + 与 v60 §2.113 (d) KDD/VLDB/IFCMemoryBench BIM 形成"权重 lineage × 内存模块化 × 语料侧范式"三维合规闭环 = (d) LongStraw arXiv:2607.14952 = 百万 token RL 后训练栈显存墙突破 = architecture-aware detach(prompt 段 detach 出 autograd + 每个 response 单独 replay + KV/compressed-attention 各自保留 architecture-aware state)+ group size 2→8 峰值分配显存只增 0.21 GB + Qwen3.6-27B 8× H20 完成 2.1M positions + GLM-5.2 compressed-attention MoE 32× H20 端到端通路 + 独立 stress test 4.46M positions + 与 v60 §2.113 (d) IFCMemoryBench BIM 形成"运行时 KV 注入 + 跨模型 Engram + 语料侧 Build-time + RL 后训练 detach + replay"五维 Memory 杠杆 = (e) Agent Harness 闭环三件套 SkillGate arXiv:2608.18852 + CriPO arXiv:2607.18082 + Zetta ζ arXiv:2608.16590 + AID-Guard arXiv:2608.21159 = SkillGate「选择器信用饥饿」诊断 + outcome 信用只到执行 token + action-local 优势只到技能命名 token 双通道不相交 + 9B 策略 5 agent benchmark 16 slate 40.8%→53.2% + 误导候选暴露砍 2/3 + CriPO on-policy 自蒸馏 + 反事实自教师定位负优势 rollout 中 criterion 相关 token + token-level 优势翻转 + 约 2× 步数达成 vanilla 终态性能 + Zetta ζ 三时间尺度解耦(动作级 / rollout 批次级 / 进化轮次级)+ frozen 底座 + 演化 {C, R, T} + LIBERO-Pro 90.8% / RoboCasa 93.6% / 11.1× 推理加速 + AID-Guard stateful authorization-to-effect closure + commit 阶段重新验证请求 + provider 状态 + delivery fence + reservation 歧义时保留 = 与 v60 §2.113 (e) Foundry + AgentRx + v59 Inadvertent Context Leakage + v58 Bounded Agents APC + v56 Mythos 5 + v53 Stealing + v61 AID-Guard = Agent 安全六件套闭环 + 旁注 4 件 = (1) vLLM Conference 8-25 当日 Roadmap Q3 实际交付(YouTube 8-18 Woosuk Kwon + Zachary Xi "State of vLLM 2026" preview + Issue #48168)= Flat Model + MRV2 核心重写 + Rust Frontend 持续改进 + llm-d 整合 + Q3 production agentic workload + 1000+ TPS SpecDec 目标 + (2) SWE-bench Pro 8-25 BenchLM 更新 = Claude Mythos 5 80.3% + Fable 5 80% + Opus 5 79.2% + Qwen3.8 Max 67.7%(v60 时为 61.7%,现已升至 67.7% 创 open-weights SWE-bench Pro 历史最高)+ Sakana Fugu-Ultra 73.7% + Qwen3.6-27B 53.5% + Muse Glimmer 30B 51.2% = open-weights 阵营显著抬升 + (3) HF Daily 8-23/24 立标信号梯度 = EnvHarness 246→254▲ + Zetta ζ 141→142▲ + SemaPLC 115▲ + FACET 112→114▲ + SWE-bench Science 58→61▲ + SPADE 47→48▲ + SkillEvo 29→30▲ + MemTrapBench 31▲ + 4DAnyone 66→69▲ + Co-RL 91→92▲ + OmniScientist 88▲ + 化学逆合成 32▲ + ForgeWM 22▲ + SemComp-Bench 155▲ + WithEveryone 39▲ + (4) 工程实战信号 = NVIDIA Dynamo 1.4.1(8-22 更新 KV Router + NIXL + Planner + OpenAI 兼容 HTTP API + disaggregated serving)+ Apple Silicon MLX/MLX-VLM Qwen3.8 本地/边缘推理栈正式确立 + pgvectorScale 5000万向量 471 QPS / 99% recall vs Qdrant 41 QPS(差距 ~11×)+ Multi-Vector Late Interaction 现状(Qdrant v1.10+ / Weaviate v1.29+ / Vespa / LanceDB v0.15+ / VectorChord / Milvus v2.6.4)+ SGLang + vllm-ascend Atlas 800T A2 Llama 2-70B 实测 11.8→45 tokens/s(3.8×)+ NPU 利用率 30%→90% + DeepSeek-R1 A100 80G vLLM 128.5 vs SGLang 112.3 tokens/s + AWQ/FP8 vLLM 显存优化 50+ 组实测 + LongHorizon-Harness arXiv:2608.01964 Alibaba DreamX MEA Loop + StateM 311 ⭐ = 第三十八波 138 主线数量更新:151→156 共识(v61 +5 C152 FlashPrefill V2 / C153 OPD 三联 / C154 Centered Residual + AdaPop / C155 LongStraw / C156 Agent Harness 闭环三件套 + AID-Guard)+ 132→136 争议(v61 +4 D133 FlashPrefill V2 H20 专属 / D134 OPD 三联 + RSTG 蒸馏信号质量适用边界 / D135 Centered Residual + AdaPop 跨架构 / D136 LongStraw 严格等价性 + AID-Guard 工程复杂度)+ 189→194 开放 = 156/136/194主文件压缩:v60 68KB → v61 候选 ≤80KB 完成。fresh sources(本轮新增材料,精选): - inbox jay 8-25 inference-vector-k8s-agent(vLLM vs SGLang vs TensorRT-LLM 2026 实战选型 + NVIDIA Dynamo 1.4.1 KV Router + NIXL + Planner + OpenAI 兼容 HTTP API + Qwen3.8 Apple MLX/MLX-VLM 本地/边缘推理栈正式确立 + Multi-Vector Late Interaction 现状 + pgvectorScale 471 QPS vs Qdrant 41 QPS · 10×) - inbox jay 8-25 csdn-inference-quantization-highvalue(5 件 CSDN 高价值:⭐⭐⭐⭐⭐ 昇腾 NPU SGLang + vllm-ascend Atlas 800T A2 Llama 2-70B 实测 11.8→45 tokens/s · NPU 利用率 30%→90% + ⭐⭐⭐⭐ DeepSeek-R1 vLLM 128.5 vs SGLang 112.3 tokens/s + ⭐⭐⭐⭐ SGLang vs vLLM RadixAttention vs PagedAttention 5× advantage + Ascend NPU 实战 + vLLM PagedAttention + Continuous Batching + AWQ/FP8 显存优化 50+ 组实测) - inbox jay 8-25 engineering-articles-secondary-filter(LongHorizon-Harness arXiv:2608.01964 Alibaba DreamX MEA Loop · MIT · 311 ⭐ + c-CRAB + Self-Harness) - inbox jay 8-25 ai-engineering-github-trending(pgbot 601 ⭐ + llm-inference-engineering 218 ⭐ + rome-os 274 ⭐ + browser-use/macos-harness 752 ⭐ + awesome-deepseek-harness 882 ⭐ + Agent S TMLR 2026 OSWorld SOTA 72.60%) - inbox tom 8-25 agent-rag-longcontext-radar(EnSI-RAG arXiv:2608.21252 Entity-Structure-Indexed + δ-mem arXiv 8×8 关联记忆矩阵 + PV-SST arXiv:2608.20438 + Human-Centric Intelligence Survey arXiv:2608.18184 + FlavourBench + Hydra-0 + SparsePR + PhysCaP) - inbox tom 8-25 agents-lite(AID-Guard arXiv:2608.21159 stateful authorization-to-effect closure + PV-SST 词法收敛 + Specification Portability arXiv:2608.21208) - inbox tom 8-24 hf-daily-2026-08-24(EnvHarness 254▲ + SemComp-Bench 155▲ + Zetta ζ 142▲ + SemaPLC 115▲ + FACET 114▲ + Co-RL 92▲ + OmniScientist 88▲ + 4DAnyone 69▲ + SWE-bench Science 61▲ + SPADE 48▲ + MemTrapBench 31▲ + SkillEvo 30▲ + 化学逆合成 32▲ + ForgeWM 22▲ + WithEveryone 39▲) - inbox flyp 8-25 reliability-science-long-horizon-agents-critical-read(Reliability Science Framework arXiv:2603.29231 RDC/VAF/GDS/MOP + HORIZON arXiv:2604.11978 COLM 2026 leaderboard + "memory scaffold 普遍伤害"反向论据 + ReliabilityBench(rel 2026b)) - inbox stephen 8-25 morning棒 0517 coordination(vLLM Conference 8-25 关键事件观察 · Woosuk Kwon + Zachary Xi 主讲 "State of vLLM 2026" · 当日观察 = 8-25 evening 棒补全 + Flyp R52 risk 备料补强 + Tom R70 inference / rag 部分实质触发) - paper_cards 1027-1057 + paper_cards/420(On-Policy Distillation 5 cites)+ 423(LongStraw 2 cites)+ 460(Xiaomi-Robotics-1)+ 608 + 700(CriPO 1 cite)+ 772(RSTG 1 cite)+ 974(S²VOPD)+ 975(SimpleOPD)+ 1028(Centered Residual)+ 1032(SkillGate)+ 1033(AdaPop)+ 1039(FlashPrefill V2) - 2 web_search:vLLM Conference 8-25 Roadmap Q3(YouTube f_EAJiUWlbU 8-18 Woosuk Kwon "State of vLLM 2026" preview)+ SWE-bench Pro 8-25 BenchLM 数据更新(Mythos 5 80.3% + Fable 5 80% + Opus 5 79.2% + Qwen3.8 Max 67.7% + Qwen3.6-27B 53.5% + Muse Glimmer 30B 51.2%) - organized/promo/surveys/2026-08-25-engineering.md(spark 8-25 engineering 综述,5 主线 v2 三段式反方)

v60(2026-08-23 09:15 · Jay Wave3 E1 综合)

v59 主体保留 + 147/129/184 全部保留。v60 在 v59 基础上叠加 1 主线 5 子件 + 旁注 3 件(§2.113)= (a) vLLM Conference 2026-08-25/26 Roadmap Q3 六轴 + (b) SWE-bench Pro 8-15 BenchLM 立标 Mythos 5 80.3% + Anthropic 三强霸榜 + Qwen3.8-27B Apache 2.0 61.7% = open-weights 首次 SWE-bench Pro 60%+ 阵营 + Terminal-Bench 2.1/3.0 + (c) SWE-bench ProMax arXiv:2608.09802 COLM 2026 + EdgeBench 134 任务 3-4h = 多文件+长周期+多工具链复合评估 + (d) KDD 2026 LinkedIn arXiv:2604.26197 + VLDB 2026 arXiv:2604.01707 + IFCMemoryBench BIM arXiv:2602.16313 = Agent Memory 顶会立项 + (e) Microsoft Foundry + AgentRx 故障恢复 = Agent 安全五件套闭环 + 旁注 3 件 Order 66 arXiv:2608.08131 Skill 生态恶意供应链审计 / Stamile 三维度设计图 / Stephen 8-23 news 10 件(Claude Code Auto Mode 8-14 默认 + Sam Altman 6 个月内 ChatGPT 后代「掌握完整人生」)= 第三十七波 137 主线数量更新:147→151 共识(v60 +4 C148-C151)/ 129→132 争议(v60 +3 D130-D132)/ 184→189 开放(v60 +5 O185-O189)= 151/132/189主文件压缩:v59 79.0KB → v60 78KB ≤80KB 完成。fresh sources:tom 8-23 hf-daily 15 篇 + radar 8 条 / jay 8-23 csdn-llm-inference-rag + 8-22 ai-engineering-trending-aug22 + 8-22 five-category-supplement / stephen 8-23 news-x-vip-radar 10 件 / paper_cards 1040-1057(21 张含 4DAnyone/Repo0/IAR/SWE-bench Science/Chain-of-Experience/SkillEvo/Inadvertent Context Leakage/HSI 等)/ 2 web_search (vLLM Conference 8-25/26 Issue #48168 + SWE-bench Pro BenchLM 8-15)。

v59(2026-08-22 09:15 · Jay Wave3 E1 综合)

v58 主体保留 + §2.1-§2.111 节段纲要 + 143/126/179 全部保留。v59 在 v58 基础上叠加 1 主线 5 子件 + 旁注 4 件(§2.112)= (a) InferScale arXiv:2607.27090 GPU 原生 KV Injection 个性化 LLM Serving = prompt injection vs KV injection 双轨机制 = KV store per conversation 1.8–4.8 GB + Jasper proximity graph <25 MB + SGLang decode latency 实测 + 与 v58 §2.111 (h) Cross-Model Memory Transfer Engram + CTIFoundry 形成"运行时 KV 注入 + 跨模型 Engram + 语料侧 Build-time"三维 Memory 杠杆 = (b) 47billion Custom CUDA Kernels + agent-as-kernel-engineer 工作流 = 推理加速三层(引擎级插件 → 生产调优 → 内核级定制) = dispatch/build/debugging 三阶段 + RoPE/GEGLU fused attention + (c) AWS SageMaker vLLM/SGLang Native Metrics = 生产推理可观测性标准路径 + (d) Antigravity 2.0 / Morph SWE-bench Pro 8-2 立标 + SWE-bench Science arXiv:2608.19799 + Chain-of-Experience arXiv:2608.18027 三轨并进 + (e) SkillEvo arXiv:2608.13120 + Repo0 arXiv:2608.19854 + Inadvertent Context Leakage arXiv:2608.19857 三件套 + 旁注 4 件 = HF Daily 8-22 5 件立标 + Tom 8-22 0840 radar 高价值 3 条 + Stephen 8-22 0910 news 7 件 + 1 web_search(Spheron + Morph)= 第三十六波 136 主线

URL v55 完整沿用清单(v54-v58 等价集合,v61 沿用 486 件 URL 锚点)

https://2026.hpca-conf.org/details/hpca-2026-industry-track/2/Characterizing-Cloud-Native-LLM-Inference-at-ByteDance-and-Exposing-Optimization-Chal https://2026.sigmod.org/sigmod_papers.shtml https://aclanthology.org/2026.acl-demo.1 https://aclanthology.org/2026.acl-demo.18 https://aclanthology.org/2026.acl-demo.19 https://aclanthology.org/2026.acl-demo.2 https://aclanthology.org/2026.acl-demo.21 https://aclanthology.org/2026.acl-demo.23 https://aclanthology.org/2026.acl-demo.3 https://adg.csdn.net/6a6992bf10ee7a33f293d657.html https://aembit.io/blog/the-ultimate-guide-to-mcp-security-vulnerabilities https://aiagentssimplified.substack.com/p/2026s-q1-ai-updates https://aimultiple.com/inference-engines https://aishwaryasrinivasan.substack.com/p/all-you-need-to-know-about-loop-engineering https://aixfunda.substack.com/p/top-llm-rag-and-agent-updates-of https://aixfunda.substack.com/p/top-llm-rag-and-agent-updates-of-0d2 https://alexeyondata.substack.com/p/what-1000-job-descriptions-reveal https://ali-liu.com/blog/2026-rag-is-not-dead-its-just-boring-now https://ali-liu.com/blog/2026-the-year-of-the-os-for-llms https://ali-liu.com/blog/3-lessons-from-running-rag-in-production https://ali-liu.com/blog/how-modern-rag-systems-are-actually-built https://ali-liu.com/blog/the-data-moat-is-dead-how-2026s-llm-apps-really-make-money https://ali-liu.com/blog/the-difference-between-context-engineering-and-rag https://ali-liu.com/blog/the-real-cost-of-self-hosted-llms https://ali-liu.com/blog/the-rise-of-self-hosted-llms https://ali-liu.com/blog/why-evaluation-is-the-foundation-of-production-ai https://ali-liu.com/blog/why-llm-rag-broke-the-data-moat https://ali-liu.com/blog/why-most-llm-evaluations-fail-the-rag-mismatch https://alicelabs.ai/en/insights/best-ai-agent-frameworks-2026 https://artificialanalysis.ai/evaluations/terminalbench-hard https://arxiv.org/abs/2403.05527 https://arxiv.org/abs/2504.11320 https://arxiv.org/abs/2510.09665 https://arxiv.org/abs/2511.02248 https://arxiv.org/abs/2511.16681v2 https://arxiv.org/abs/2601.07504 https://arxiv.org/abs/2602.14617 https://arxiv.org/abs/2603.07379 https://arxiv.org/abs/2603.20397 https://arxiv.org/abs/2603.29231 https://arxiv.org/abs/2605.19743 https://arxiv.org/abs/2605.20173 https://arxiv.org/abs/2605.27744v2 https://arxiv.org/abs/2606.06324 https://arxiv.org/abs/2606.14589 https://arxiv.org/abs/2606.20295 https://arxiv.org/abs/2607.08028 https://arxiv.org/abs/2607.14159 https://arxiv.org/abs/2607.17715 https://arxiv.org/abs/2607.20510 https://arxiv.org/abs/2607.21503 https://arxiv.org/abs/2607.22389 https://arxiv.org/abs/2607.23782 https://arxiv.org/abs/2607.23783 https://arxiv.org/abs/2607.23802 https://arxiv.org/abs/2607.24000 https://arxiv.org/abs/2607.24882 https://arxiv.org/abs/2607.25379 https://arxiv.org/abs/2607.25431 https://arxiv.org/abs/2607.25498 https://arxiv.org/abs/2607.25600 https://arxiv.org/abs/2607.25996 https://arxiv.org/abs/2607.26475 https://arxiv.org/abs/2607.26571 https://arxiv.org/abs/2607.26627 https://arxiv.org/abs/2607.26637 https://arxiv.org/abs/2607.27201 https://arxiv.org/abs/2607.27600 https://arxiv.org/abs/2607.27616 https://arxiv.org/abs/2607.27816 https://arxiv.org/abs/2607.27919 https://arxiv.org/abs/2607.27958 https://arxiv.org/abs/2607.28227 https://arxiv.org/abs/2607.28229 https://arxiv.org/abs/2607.28415 https://arxiv.org/abs/2607.28509 https://arxiv.org/abs/2607.28595 https://arxiv.org/abs/2607.28617 https://arxiv.org/abs/2607.28633 https://arxiv.org/abs/2607.28675 https://arxiv.org/abs/2607.29167 https://arxiv.org/abs/2607.29209 https://arxiv.org/abs/2607.29402 https://arxiv.org/abs/2607.29405 https://arxiv.org/abs/2607.29459 https://arxiv.org/abs/2607.29591 https://arxiv.org/abs/2607.29610 https://arxiv.org/abs/2607.29678 https://arxiv.org/abs/2607.29679 https://arxiv.org/abs/2608.00101 https://arxiv.org/abs/2608.00303 https://arxiv.org/abs/2608.00677 https://arxiv.org/abs/2608.00742 https://arxiv.org/abs/2608.00881 https://arxiv.org/abs/2608.00902 https://arxiv.org/abs/2608.01526 https://arxiv.org/abs/2608.01651 https://arxiv.org/abs/2608.01678 https://arxiv.org/abs/2608.01735 https://arxiv.org/abs/2608.01964 https://arxiv.org/abs/2608.02143 https://arxiv.org/abs/2608.02162 https://arxiv.org/abs/2608.02515 https://arxiv.org/abs/2608.02583 https://arxiv.org/abs/2608.02645 https://arxiv.org/abs/2608.02672 https://arxiv.org/abs/2608.02703 https://arxiv.org/abs/2608.02989 https://arxiv.org/abs/2608.03036 https://arxiv.org/abs/2608.03451 https://arxiv.org/abs/2608.03487 https://arxiv.org/abs/2608.03506 https://arxiv.org/abs/2608.03700 https://arxiv.org/abs/2608.03744 https://arxiv.org/abs/2608.03764 https://arxiv.org/abs/2608.03874 https://arxiv.org/abs/2608.03972 https://arxiv.org/abs/2608.03994 https://arxiv.org/abs/2608.04003 https://arxiv.org/abs/2608.04530 https://arxiv.org/abs/2608.04964 https://arxiv.org/abs/2608.05042 https://arxiv.org/abs/2608.05076 https://arxiv.org/abs/2608.05108 https://arxiv.org/abs/2608.05137 https://arxiv.org/abs/2608.05138 https://arxiv.org/abs/2608.05369 https://arxiv.org/abs/2608.05424 https://arxiv.org/abs/2608.05565 https://arxiv.org/abs/2608.05604 https://arxiv.org/abs/2608.05747 https://arxiv.org/abs/2608.05784 https://arxiv.org/abs/2608.05987 https://arxiv.org/abs/2608.06020 https://arxiv.org/abs/2608.06033 https://arxiv.org/abs/2608.06060 https://arxiv.org/abs/2608.06130 https://arxiv.org/abs/2608.06197 https://arxiv.org/abs/2608.06216 https://arxiv.org/abs/2608.06257 https://arxiv.org/abs/2608.06301 https://arxiv.org/abs/2608.06305 https://arxiv.org/abs/2608.06729 https://arxiv.org/abs/2608.07193 https://arxiv.org/abs/2608.08020 https://arxiv.org/abs/2608.10450 https://arxiv.org/abs/2608.11341 https://arxiv.org/abs/2608.11632 https://arxiv.org/abs/2608.11660 https://arxiv.org/abs/2608.11745 https://arxiv.org/abs/2608.11752 https://arxiv.org/abs/2608.12440 https://arxiv.org/abs/2608.12700 https://arxiv.org/abs/2608.12915 https://arxiv.org/abs/2608.13122 https://arxiv.org/abs/2608.13263 https://arxiv.org/abs/2608.13410 https://arxiv.org/abs/2608.13417 https://arxiv.org/abs/2608.13426 https://arxiv.org/abs/2608.13499 https://arxiv.org/abs/2608.13517 https://arxiv.org/abs/2608.13555 https://arxiv.org/abs/2608.13606 https://arxiv.org/abs/2608.13696 https://arxiv.org/abs/2608.13757 https://arxiv.org/abs/2608.13760 https://arxiv.org/abs/2608.13987 https://arxiv.org/abs/2608.14054 https://arxiv.org/abs/2608.14210 https://arxiv.org/abs/2608.14290 https://arxiv.org/abs/2608.14391 https://arxiv.org/abs/2608.14528 https://arxiv.org/abs/2608.14530 https://arxiv.org/html/2504.11320 https://arxiv.org/html/2509.12384v1 https://arxiv.org/html/2511.02248v1 https://arxiv.org/html/2512.18470v6 https://arxiv.org/html/2602.06036v1 https://arxiv.org/html/2602.16873v1 https://arxiv.org/html/2603.07379v1 https://arxiv.org/html/2604.01647v2 https://arxiv.org/html/2605.01280v1 https://arxiv.org/html/2605.08717v1 https://arxiv.org/html/2605.11733v1 https://arxiv.org/html/2605.19743v1 https://arxiv.org/html/2605.20173v1 https://arxiv.org/html/2605.24217v1 https://arxiv.org/html/2606.01927v1 https://arxiv.org/html/2606.06324v1 https://arxiv.org/html/2606.14589v1 https://arxiv.org/html/2606.30391v1 https://arxiv.org/html/2607.04763v1 https://arxiv.org/html/2607.06519v1 https://arxiv.org/html/2607.08028v1 https://arxiv.org/html/2607.10169v1 https://arxiv.org/html/2607.12340v1 https://arxiv.org/html/2607.12406v1 https://arxiv.org/html/2607.12756v1 https://arxiv.org/html/2607.14159v1 https://arxiv.org/html/2607.14277v1 https://arxiv.org/html/2607.16859v1 https://arxiv.org/html/2607.18314v1 https://arxiv.org/html/2607.20465v1 https://arxiv.org/html/2607.20957v1 https://arxiv.org/html/2607.21503v1 https://arxiv.org/html/2607.21557v1 https://arxiv.org/html/2607.21653v1 https://arxiv.org/html/2607.21848v1 https://arxiv.org/html/2607.21962v1 https://arxiv.org/html/2607.22042v1 https://arxiv.org/html/2607.22043v1 https://arxiv.org/html/2607.22157v1 https://arxiv.org/html/2607.22375v1 https://arxiv.org/html/2607.24223v1 https://arxiv.org/html/2607.24368v1 https://arxiv.org/html/2607.24593v1 https://arxiv.org/html/2607.24720v1 https://arxiv.org/html/2607.24904v1 https://arxiv.org/html/2607.24957v1 https://arxiv.org/html/2607.25398v1 https://arxiv.org/html/2607.25537v1 https://arxiv.org/html/2607.25895v1 https://arxiv.org/html/2608.01526v1 https://arxiv.org/html/2608.13900v1 https://arxiv.org/html/2608.16515v1 https://arxiv.org/html/2608.16536v1 https://arxiv.org/html/2608.16628v1 https://arxiv.org/html/2608.16776v1 https://arxiv.org/pdf/2511.02248 https://arxiv.org/pdf/2607.20957 https://beta.hyper.ai/en/papers/2608.13900 https://blog.csdn.net/Gaga246/article/details/155610267 https://blog.csdn.net/ProceNest/article/details/160956302 https://blog.csdn.net/gitblog_00554/article/details/151436503 https://blog.csdn.net/gitblog_00885/article/details/151437102 https://blog.csdn.net/ld326/article/details/161401770 https://blog.csdn.net/qq_51605551/article/details/163329288 https://blog.csdn.net/weixin_35774598/article/details/162381479 https://blog.csdn.net/xx_nm98/article/details/158851692 https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash https://blog.kubesimplify.com/day-5-local-llm-inference-engines-wrappers-and-what-to-pick https://blog.lmcache.ai/en/2026/06/23/vllm-lmcache-a-starter-guide-no-gpu-required https://blog.modelcontextprotocol.io/posts/2026-07-28 https://blog.modelcontextprotocol.io/posts/2026-07-28-release-candidate https://buzzgrewal.medium.com/ai-agents-dont-eed-vector-search-anymore-inside-the-agentic-search-stack-replacing-rag-in-2026-58efcabe4f6f https://buzzgrewal.medium.com/ai-agents-dont-need-vector-search-anymore-inside-the-agentic-search-stack-replacing-rag-in-2026-58efcabe4f6f https://byteiota.com/vllm-v0-25-model-runner-v2-default-pagedattention-gone https://cameronrwolfe.substack.com/p/agent-evals https://cameronrwolfe.substack.com/p/agentic-rl https://christian-schneider.net/blog/rag-security-forgotten-attack-surface https://cruxdigits.nl/blog/context-engineering-ai-agents-2026 https://dbgroup.cs.tsinghua.edu.cn/ligl/publications.html https://deepmind.google/blog https://dev-discuss.pytorch.org/t/intel-gpu-cpu-enabling-status-and-feature-plan-2026-h1-update/3320 https://dev.to/gabrielanhaia/70-of-enterprise-rag-deployments-fail-before-production-heres-what-kills-them-26ml https://dev.to/hadil/why-ai-agents-fail-in-production-and-how-engineering-teams-are-fixing-it-in-2026-job https://developer.cloud.tencent.com/article/2707601 https://developer.nvidia.com/blog/boost-inference-performance-up-to-15x-on-nvidia-blackwell-using-dflash-speculative-decoding https://developer.nvidia.com/blog/mastering-llm-techniques-inference-optimization https://dextralabs.com/blog/how-to-build-a-24-7-ai-customer-service-agent https://dkennetz.substack.com/p/llm-inference-curriculum https://dl.acm.org/doi/10.1145/3749168 https://dl.acm.org/doi/10.1145/3786583.3786904 https://dl.acm.org/doi/10.1145/38020094 https://dl.acm.org/doi/10.1145/38020284 https://docs.nvidia.com/aiperf/dev/tutorials/load-patterns-scheduling/control-hooks-by-server-v-llm-sg-lang-trt-llm https://docs.nvidia.com/deeplearning/frameworks/vllm-release-notes/index.html https://docs.nvidia.com/nim/large-language-models/2.0.2/about-nim-llm/release-notes.html https://emergentmind.com/topics/swe-bench-3b86a734-b378-4ee6-bcaf-640949ed7afb https://flashinfer.ai/whl/cu129 https://freedom.tech/project/vllm https://futureagi.com/blog/llm-incident-response-playbook-2026 https://futureagi.substack.com/p/the-complete-guide-to-llm-evaluation https://futureagi.substack.com/p/the-llm-incident-runbook-six-steps-f27 https://gateway-api-inference-extension.sigs.k8s.io https://gitcode.csdn.net/69fdd2c9cc6cf6495d58314b.html https://github.com/JustVugg/colibri https://github.com/LLMSecurity/awesome-agent-skills-security https://github.com/LMCache/LMCache https://github.com/Michaelvll/llm-ie-benchmarks https://github.com/MoonshotAI/FlashKDA https://github.com/NVIDIA/skills https://github.com/NousResearch/hermes-agent https://github.com/TsinghuaC3I/Awesome-Memory-for-Agents https://github.com/Yigtwxx/awesome-rag-production https://github.com/Zijian-Ni/awesome-ai-agents-2026 https://github.com/advisories/GHSA-4r2x-xpjr-7cvv https://github.com/ai-boost/awesome-harness-engineering https://github.com/astral-sh/uv https://github.com/badlogic/pi-mono https://github.com/deepagents-ai/agent-backend https://github.com/different-ai/openwork https://github.com/firecrawl/anydoc https://github.com/firecrawl/firecrawl https://github.com/herdrdev/herdr https://github.com/langchain-ai/deepagents https://github.com/langflow-ai/langflow https://github.com/langgenius/dify https://github.com/llm-d/llm-d https://github.com/lyogavin/airllm https://github.com/malisper/pgrust https://github.com/memvid/memvid https://github.com/moeru-ai/airi https://github.com/obra/superpowers https://github.com/ollama/ollama https://github.com/oramasearch/oramacore https://github.com/patchy631/time-to-first-token https://github.com/s7a9/C2KV https://github.com/simonw/uv-init-demos https://github.com/stas00/ml-engineering https://github.com/syhya/mlsys26-flashinfer-contest https://github.com/vllm-project/vllm/issues/34018 https://github.com/vllm-project/vllm/issues/40608 https://github.com/vllm-project/vllm/pull/31987 https://github.com/vllm-project/vllm/pull/32319 https://github.com/vllm-project/vllm/pull/32668 https://github.com/vllm-project/vllm/releases https://github.com/vllm-project/vllm/releases/tag/v0.27.0 https://github.com/xAI/grok-build https://haoailab.com/blogs/distserve-retro https://help.openai.com/en/articles/6825453-chatgpt-release-notes https://huggingface.co/blog/ResterChed/kimi-k3-model-overview-mxfp4-quantization-open-wei https://huggingface.co/blog/Svngoku/agentic-coding-trends-2026 https://huggingface.co/blog/daya-shankar/open-source-llm-models-to-run-locally https://huggingface.co/blog/daya-shankar/open-source-llms https://huggingface.co/blog/huggingface/state-of-open-models-summer-2026 https://huggingface.co/blog/huggingface/state-of-os-hf-spring-2026 https://huggingface.co/blog/icml-2026-open-reproductions https://huggingface.co/blog/icml-reproduction https://huggingface.co/blog/security-incident-july-2026 https://huggingface.co/blog/state-of-open-models-summer-2026 https://huggingface.co/docs/text-generation-inference https://huggingface.co/khoichk/kimi-k3 https://huggingface.co/papers https://huggingface.co/spaces/agent-sandbox/graphify https://hugobowne.substack.com/p/agentops-lessons-from-over-1400-production https://inferenceops.substack.com/p/state-of-the-model-serving-communities-b93 https://interconnects.ai/p/5-useful-things-youll-learn-in-my https://interconnects.ai/p/glm-53-how-chinese-labs-keep-stride https://jamwithai.substack.com/p/the-2026-roadmap-production-aiml https://k-ai.ai/en/news/rag-cross-source-contradiction-failure-mode https://karozieminski.substack.com/p/context-engineering-product-builders-guide-2026 https://labs.cloudsecurityalliance.org/agentic/agentic-mcp-security-best-practices-v1 https://lambda.ai/blog/flashattention-4-gives-the-nvidia-blackwell-platform-its-most-optimized-attention-ket-yet https://leaddev.com/ai/your-llm-inference-benchmark-is-lying-to-you https://learnaitogethernewsletter.substack.com/p/lai-137-where-ai-engineering-is-going https://leetllm.com/blog/llm-inference-engine-comparison-2026 https://lilianweng.github.io/posts/2026-07-04-harness/ https://lucaberton.com/blog/ai-model-serving-kubernetes-vllm-triton-nim-2026 https://machinelearning.apple.com/research/quantspec https://magazine.sebastianraschka.com/p/controlling-reasoning-effort-in-llms https://magazine.sebastianraschka.com/p/using-local-coding-agents https://medium.com https://medium.com/@addyosmani/my-llm-coding-workflow-going-into-2026-52fe1681325e https://medium.com/@wasowski.jarek/i-benchmarked-6-vector-databases-for-rag-none-wins-everywhere-in-2026-900971966b7d https://medium.com/data-science-collective/355k-github-stars-in-5-months-17-defense-rate-the-complete-honest-guide-to-openclaw https://metafiedlab.com/blog/how-to-build-a-production-ready-rag-pipeline-in-2026 https://micheallanham.substack.com/p/comparative-analysis-of-rag-architectures https://mlflow.org/articles/building-production-ready-ai-agents-in-2026 https://moondream.ai/blog/photon-2-launch https://newsletter.pragmaticengineer.com/p/what-is-inference-engineering https://open.substack.com/pub/semianalysis/p/can-amd-break-the-cuda-moat-amd-advancing https://openai.com/index/gpt-5-6-frontier-intelligence-efficiency https://openai.com/index/gpt-daybreak https://openai.com/index/hugging-face-model-evaluation-security-incident https://openai.com/index/huggingface-model-evaluation-security-incident https://openai.com/index/previewing-ultrafast https://openeuler.csdn.net/6a508ec510ee7a33f28c08ff.html https://orca.security/resources/blog/cve-2026-22778-vllm-rce-vulnerability https://orca.security/resources/blog/sglang-llm-framework-rce-vulnerabilities https://papers.cool/arxiv/2608.13263 https://papers.cool/arxiv/2608.13499 https://particula.tech/blog/sglang-vs-vllm-inference-engine-comparison https://petronellatech.com/blog/openclaw-ai-agent-guide-2026 https://pradeepkj.substack.com/p/production-deployment-challenges https://proceedings.iclr.cc/paper_files/paper/2026/file/444a3737adaee10d86ad2ef5f74468e6-Paper-Conference.pdf https://ranksquire.com/2026/05/27/vector-database-news-may-2026 https://redis.io/blog/rag-at-scale https://saeed.github.io/files/arc_niac26.pdf https://sebastianraschka.com/blog/2026/gpt-5-6-configurations.html https://simonwillison.net/2026/Jul/22/openai-cyberattack https://simonwillison.net/2026/Jul/27/kimi-k3 https://simonwillison.net/2026/Jul/28/uv https://snorkel.ai/leaderboard/terminal-bench-2-1 https://snyk.io/news/snyk-2026-state-of-agentic-ai-adoption-volume-ii https://soulhacked.substack.com/p/research https://stateofmlops.substack.com/p/state-of-mlops-2026jun1 https://stolen-thoughts.com https://techcrunch.com/2026/07/30/anthropic-says-its-own-ai-models-breached-three-companies-during-security-tests https://techsy.io/en/blog/vllm-vs-sglang https://thakicloud.com/tech-blog/en/dev/vllm-v0-25-0-model-runner-v2 https://theaiengineer.pub/p/the-ai-agents-stack-2026-edition https://theaiengineer.substack.com/p/agentic-rag-vs-cua-vs-a2a https://theaiengineer.substack.com/p/how-claude-code-actually-works https://theaiengineer.substack.com/p/how-doordash-built-their-rag-system https://theaiengineer.substack.com/p/the-4-single-agent-patterns https://theaiengineer.substack.com/p/vllm-vs-ollama-vs-sglang-vs-tensorrt https://theaiengineer.substack.com/p/why-ai-agents-keep-failing-in-production https://thebuild.com/blog/pgvector-082-and-the-trouble-with-parallel-hnsw https://thedatafirst.com/why-vllm-is-eating-the-inference-stack-in-2026/ https://thehackernews.com/2026/07/worlds-largest-ai-model-repository.html https://theneuralmaze.substack.com/p/building-the-ai-roadmap-for-2026 https://thenuancedperspective.substack.com/p/the-ai-agent-stack-in-2026 https://turion.ai/blog/vllm-vs-sglang-inference-comparison-2026 https://unipat.ai/benchmarks/MonthlySWEBench https://venturebeat.com/infrastructure/mcp-just-got-its-biggest-update-ever-heres-what-changes-for-ai-agents https://venturebeat.com/technology/43-of-ai-generated-code-changes-need-debugging-in-production-survey-finds https://vllm-project.github.io https://vllm-project.github.io/blog/decode-context-parallelism/ https://vllm-project.github.io/blog/eagle3-amd-quark/ https://vllm-project.github.io/blog/semantic-router-v0.1-iris/ https://vllm-project.github.io/blog/speculative-decoding/ https://vllm-project.github.io/blog/state-of-fp8-kv-cache/ https://vllm-project.github.io/blog/vllm-25k-tps-gpu-qwen3.5/ https://vllm.ai/blog/2026-03-24-mrv2 https://vllm.ai/blog/2026-07-16-keeping-vllm-production-quality https://vllm.ai/blog/2026-07-22-kimi-k3-preview https://vllm.ai/blog/2026-07-27-k3 https://vllm.ai/blog/2026-08-07-decode-context-parallelism https://warsawainews.substack.com/p/warsawai-news-6-12072026 https://witness.ai/blog/rag-security https://workos.com/blog/mcp-2026-spec-agent-authentication https://www.actian.com/blog/databases/how-to-evaluate-vector-databases-in-2026 https://www.aisi.gov.uk https://www.alphaxiv.org/abs/2601.11868 https://www.anthropic.com/news/claude-opus-5 https://www.atlan.com/blog/multi-agent-debugging-7-failure-modes-fixes-2026 https://www.aussieai.com/research/prefix-sharing https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultrafast-with-openai https://www.cidrdb.org/cidr2026/papers/p5-jin.pdf https://www.clawbot.blog/blog/openclaw-overtakes-react-in-github-stars-an-ai-agent-framework-phenomenon https://www.cncf.io/announcements/2026/08/10/cncf-reveals-kubecon-cloudnativecon-north-america-2026-schedule-adds-new-ai-inference-agentic-track https://www.cncf.io/blog/2026/01/28/introducing-kthena-llm-inference-for-the-cloud-native-era https://www.cncf.io/wp-content/uploads/2026/03/State-of-Cloud-Native-Development-Q1-2026.pdf https://www.codercops.com/blog/database-trends-postgres-sqlite-vector-2026 https://www.comet.com/site/blog/f1-radio-rag-ai-eval-example https://www.crusoe.ai/resources/blog/crusoe-memoryalloy-reinventing-kv-caching-for-cluster-scale-inference https://www.datadoghq.com/state-of-ai-engineering https://www.developersdigest.tech/blog/mcp-2026-07-28-breaking-changes https://www.digitalapplied.com/blog/kv-cache-optimization-techniques-2026-engineering-guide https://www.digitalapplied.com/blog/rag-anti-patterns-7-failure-modes-2026-engineering-guide https://www.faros.ai/blog/harness-engineering https://www.firecrawl.dev/blog/best-vector-databases https://www.gmicloud.ai/en/blog/affordable-fast-llm-inference-top-picks-2026 https://www.inferenceengineering.tech/learn/vllm-vs-sglang-vs-tensorrt-llm https://www.inngest.com/blog/principles-of-production-ai https://www.ithome.com/0/984/365.htm https://www.kb.cert.org/vuls/id/665416 https://www.kodemsecurity.com/resources/cve-2026-22778-critical-remote-code-execution-in-vllm-multimodal-inference https://www.kubenatives.com/p/production-runbook-vllm-oom-debugging https://www.langchain.com/state-of-agent-engineering https://www.lmsys.org/blog/2026-06-15-next-generation-speculative-decoding-dflash-v2 https://www.lmsys.org/blog/2026-07-30-sglang-google-tpu https://www.lmsys.org/blog/2026-08-07-hpc-ops-sglang https://www.lyzr.ai/blog/harness-engineering-for-ai-agents https://www.microsoft.com/en-us/research/publication/opscale-operator-level-provisioning-and-autoscaling-for-llm-serving https://www.openlayer.com/blog/post/ai-monitoring-vs-ai-observability-explained https://www.paralleliq.ai/blog/vllm-oom-errors-root-cause-diagnosis https://www.postgresql.org/about/news/pgvector-080-released-2952 https://www.postgresql.org/about/news/postgresql-18-released https://www.reddit.com/r/LocalLLaMA/comments/1k45plp https://www.salttechno.ai/datasets/vector-database-performance-benchmark-2026 https://www.sector88.co/blog/how-to-fix-vllm-oom https://www.semanticscholar.org/paper/HotPrefix%3A-Hotness-Aware-KV-Cache-Scheduling-for-in-Li-Gu/b89241bc76845411fc4aa68d26825432f1a8bb55 https://www.sentinelone.com/vulnerability-database/cve-2026-3989 https://www.sivaro.in/articles/ai-agent-deployment-pipeline-a-practitioners-guide-for-2026 https://www.solo.io/blog/llm-d-distributed-inference-serving-on-kubernetes https://www.spheron.network/blog/inference-engineering-guide-2026 https://www.spheron.network/blog/llm-d-kubernetes-disaggregated-inference-guide https://www.spheron.network/blog/nvme-kv-cache-offloading-llm-inference https://www.spheron.network/blog/token-level-gpu-pooling-multi-llm-marketplace-inference https://www.spheron.network/blog/vllm-production-deployment-2026 https://www.spheron.network/blog/vllm-vs-sglang-2026 https://www.spheron.network/blog/vllm-vs-tensorrt-llm-vs-sglang-benchmarks https://www.swebench.com https://www.themoonlight.io/en/review/vtoken-token-level-virtualization-for-reclaimable-kv-caches https://www.truefoundry.com/blog/vllm-benchmark https://www.usenix.org/system/files/osdi26-wu-haonan.pdf https://www.varonis.com/blog/huggingface-breach https://www.vastdata.com/blog/accelerating-inference https://www.yottalabs.ai/post/best-llm-inference-engines-in-2026-vllm-tensorrt-llm-tgi-and-sglang-compared https://www.yottalabs.ai/post/vllm-vs-sglang-which-inference-engine-should-you-use-in-2026 https://x.com/_akhaliq/status/2081921773910499494 https://x.com/svpino/status/2079581522936676819 https://ydb.tech/blog/post/2026/03/24/how-io_uring-overtook-libaio-benchmarks-on-nvme https://z-lab.ai/projects/dflash