Jay · 研究简报 · 2026-09-29 下午

主题:推理引擎格局 · KV Cache 系统 · Vector DB 生产基准 · Agent Stack 2026 检索范围:GitHub Trending · Hugging Face Blog/Daily · arXiv (Sep 2026) · Substack · 技术博客 本次产出:1 个草稿文件


一、推理引擎格局(2026 年 9 月快照)

vLLM vs SGLang vs LMDeploy

维度 vLLM SGLang LMDeploy
GitHub Stars ~92,700 ~36,400 —
H100 吞吐 ~12,500 tok/s ~16,200 tok/s ~16,200 tok/s
Prefix Cache xgrammar (默认) grammar-cache 复用更激进 —
生态成熟度 ✅ 最成熟 增长快 量化受限硬件最优
RadixAttention 优势 PagedAttention 前缀共享 >60% 时显著领先 —
Structured Output xgrammar 默认 schema 复用更激进 —

关键结论(来源:PremAI Blog、Spheron Blog、Winder AI):

  • SGLang + LMDeploy 吞吐量相当,均领先 vLLM 约 29%
  • vLLM 凭借生态成熟度和 broad hardware support 仍是生产首选
  • SGLang 的 RadixAttention 在 prefix-heavy traffic(系统 prompt、RAG doc、tool definition block 共享)场景下优势明显
  • 决策阈值:前缀重叠率 >60% → SGLang;否则两者差距 <5%
  • SGLang v1.2.1(2026 年 9 月)、vLLM v0.5.20、LMDeploy 持续更新

引用: - https://www.premai.io/blog/vllm-vs-sglang-vs-lmdeploy-fastest-llm-inference-engine-in-2026 - https://www.spheron.network/blog/vllm-vs-sglang-2026 - https://winder.ai/vllm-vs-ollama-vs-sglang-llm-inference-comparison


二、KV Cache 系统研究前沿(arXiv Sep 2026)

高价值论文

1. Inference Control Plane 综述

标题:From Inference Engine to Inference Control Plane: Connecting vLLM, llm-d, and the Evolution of Efficient Distributed LLM Serving 来源:arXiv:2609.23130v1 · Sep 19, 2026 · Twinkll Sisodia (BU) 核心观点: - vLLM 已成为 control plane 的基础层而非终点 - llm-d (LLM disaggregation) 正在将 prefill/decode 物理分离,形成新的调度平面 - OpenTelemetry tracing 现在覆盖 Gateway → Endpoint Picker → KV-cache → prefill/decode proxy → model server 全路径

可信度:高(arXiv 2026-09 新文,系统综述)

2. Fluid-Guided Online Scheduling

标题:Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints 来源:arXiv:2504.11320v4 · Aug 24, 2026 · MIT + Alibaba 核心观点: - 研究 KV cache 内生增长导致的内存溢出问题 - 提出 (Nested) WAIT 调度算法:通过 fluid approximation 识别 fluid stability region - 维度约减:诱导 threshold/boundary queues 捕获 memory-coupled state - 适用场景:prefill/decode 混合高并发场景

可信度:高(MIT + 阿里联合,PGM 2026 接收)

3. Sustainable Distributed LLM Inference

标题:Toward Sustainable Distributed LLM Inference: A Systems Synthesis and Research Agenda for an Energy-, Carbon-, and Cache-Aware llm-d Control Plane 来源:arXiv:2609.05565v1 · Sep 2026 核心观点: - KV state 既性能资产也是能源资产 - Peer-to-peer cache sharing 策略:比较复制 KV vs 重计算的跨点(延迟/能量/内存压力/未来复用概率) - Phase-specific power control:prefill/decode 阶段 compute/memory 行为不同,能源最优点也不同 - PowerSlider 研究了 time-varying power caps 下的阶段不对称性

可信度:中高(系统综述,文献覆盖 2023-2026)

4. Multi-Segment Attention

标题:Multi-Segment Attention: Enabling Efficient KV-Cache Management for Faster LLM Serving 来源:arXiv:2606.02964v1 核心观点: - block-level, position-aware eviction,可与 Autellix/Tokencake/ThunderAgent 调度层无缝集成 - 与 Continuum(TTL pin)、KVFlow(Agent Step Graph)、InferCept(min-waste eviction)正交

可信度:中高(新研究方向)

5. ObjectCache: Layerwise Object-Storage Retrieval

标题:ObjectCache: Layerwise Object-Storage Retrieval for KV Cache Reuse 来源:arXiv:2605.22850v1 核心观点: - S3-compatible object tier 的 KV chunk 获取、聚合和按 layer 顺序交付 - 解决 S3 读取延迟问题(TCP/RDMA-buffer/RDMA-direct 多路径) - 与 CacheFlow(重计算加速)解决不同瓶颈

可信度:中高

6. SmartGen: Selective KV Cache Transfer

标题:SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer 来源:arXiv:2607.28150v1 · Jul 30, 2026 核心观点: - disaggregated prefill/decode 环境下的选择性 KV transfer - RetroInfer (VLDB 2026):向量存储引擎支持长上下文 LLM inference

可信度:高(VLDB 引用)

7. SwiftCache: Heterogeneous KV Cache Sharing

标题:SwiftCache: Efficient LLM Serving for Multi-turn Conversations with Heterogeneous KV Cache Sharing 来源:arXiv:2606.16135v1 核心观点: - 单服务器多 GPU 异构模型协调 - 利用 GPU 内存和 GPU 间带宽资源 - 指出 vLLM/SGLang/LightLLM 的 layer-major layout 不支持动态 KV cache resize

可信度:中高

8. "Internet for KV Cache"

标题:An Internet for the KV Cache: Rethinking Classical Infrastructure Boundaries in the LLM Inference Age 来源:arXiv:2608.01526v1 相关引用: - CXL-SpecKV (FPGA speculative KV-Cache, FPGA 2026) - DroidSpeak (NSDI 2026):跨 fine-tuned 模型变体共享 KV cache

可信度:高(NSDI 引用)


三、Vector DB 生产基准(2026 年 9 月)

格局总结(按场景推荐)

场景 推荐方案 理由
现有 PostgreSQL(<5000 万向量) pgvector + pgvectorscale 继承 WAL,写入便宜
新项目 <1000 万向量 Qdrant Cloud(最佳免费层)或 ChromaDB(原型快速) 性能 vs 简易平衡
新项目 1000-1 亿向量 Pinecone(零运维)或 Weaviate(混合搜索)或 Milvus(自托管成本低) 按规模分层
新项目 >1 亿向量 Milvus/Zilliz Cloud 或 Pinecone serverless 分布式横向扩展
电商搜索(1 亿产品) Milvus(分布式)或 Qdrant(过滤查询) 规模 + 过滤
企业知识搜索 Weaviate 或 pgvectorscale(混合搜索) 语义 + 关键词
LLM 语义缓存 Redis(超低延迟) 毫秒级
合规敏感(数据主权) Pinecone BYOC 或 pgvectorscale 自己控制
研究/原型 Chroma(免费,无需配置) 零配置

关键性能数字

  • Redis(1 亿向量 benchmark):90% 精度时 median latency ~200ms(top 100 NN),QPS ~66,000 insertion/s;95% 精度时 median latency ~1.3s
  • Qdrant(1M 向量,1536 维,~550 QPS):p50 ~2.9ms,p99 ~8ms
  • Weaviate(同等规模):p50 ~5.8ms,p99 ~18ms
  • pgvector:适合 <5000 万向量,PostgreSQL 生态内统一管理

2026 新动态

  • VectorDBBench(Zilliz,v2.0.0,Sep 10 2026):新增 pgvector CLI 模块,1760 commits
  • Benchmark 可信度警告(Actian):供应商自评 benchmark 存在架构特定优势伪装成客观比较的问题
  • 方向:2026 年倾向于集成平台(PostgreSQL + pgvector → 企业混合引擎),而非"纯向量"孤岛

引用: - https://redis.io/blog/best-open-source-vector-databases-comparison - https://www.actian.com/blog/databases/how-to-evaluate-vector-databases-in-2026 - https://github.com/zilliztech/vectordbbench


四、Substack 高价值工程洞察

1. AI Agents Stack 2026(The AI Engineer,Eric Roby)

核心观点: - Agent ≠ Chatbot:Agent 需要跨多步执行的 state management、tool access governed by protocols、跨 session 持久 memory、autonomous reasoning loops、real-time guardrails - Memory 三层架构(2024:"选个向量数据库做 RAG" → 2026:第一性 architectural primitive): 1. In-context cache(前缀共享) 2. Session memory(短期状态) 3. Long-term memory(持久化知识) - Guardrails 独立学科:不再是输入/输出过滤器,而是授权 tool calls、执行 rate limits、验证 agent 实际行为 - Model Routing 模式:分类/分流用小模型;困难推理用 frontier 模型;嵌入/评估用专用模型。RouteLLM 研究表明 routing 可显著降低成本同时保留质量 - 生产代理很少用单一模型

可信度:高(AI Engineer newsletter,Eric Roby,圈内广泛阅读)

引用:https://theaiengineer.substack.com/p/the-ai-agents-stack-2026-edition

2. RAG 2026 完整指南

核心观点: - Cross-Encoder Re-Ranking:Hybrid Search 先 top-100 候选,再用 Cross-Encoder (Cohere Rerank 3.5 / BGE-Reranker) 精排到 top-5~10 - 语义分块(Semantic Chunking):不用固定字符数切割,而用 cosine distance 阈值检测主题转换 - Small-to-Big (Parent-Document) 策略:子 chunk(100 tokens)精确匹配 + 父文档(整页)丰富上下文 - Hybrid Search + RRF:向量 + 关键词 + 元数据信号融合,Reciprocal Rank Fusion 合并排名 - Naive RAG(纯语义相似度)在生产中已很少部署

可信度:高(Shrinivasan,AI 工程师圈广泛引用)

引用:https://aishwaryasrinivasan.substack.com/p/all-you-need-to-know-about-rag-in

3. AI Agents Stack 选择指南(The Nuanced Perspective)

核心观点: - 2026 评估转向 continuous improvement loop:LLM judges 成为 grading agent 输出的默认方式 - Arize Alyx:从 observability 数据中自动发现常见失败模式 - Managed retrieval 服务(如 Glean):RAG + 传统搜索融合,牺牲 knob-level control换取免运维 - Hybrid search = 生产默认,naive RAG 已过时

引用:https://thenuancedperspective.substack.com/p/how-to-choose-your-ai-agent-stack

4. 学术 × AI 代理(Kiran Garimella)

核心观点: - LLM 用户论文产量提升 30-90%(Science 研究,Kusumegi et al.) - 但:传统质量信号与实际质量正在脱钩——LLM 写作语言复杂度更高,但被接收率反而更低 - 对已有训练的研究者(已具备判断力)增益更大;对新手风险更高(产出量大但质量低)

可信度:高(Science peer-reviewed 研究)

引用:https://kirangarimella.substack.com/p/ai-agents-and-academia


五、Hugging Face 生态快照(Sep 2026)

模型下载排行(30 天)

家族 下载量 占比
Qwen 240.1M ~60%
Gemma 32.6M ~8%
Ornith 13.5M ~3%
Llama 12.9M ~3%
gpt-oss 11.9M ~3%
DeepSeek 4.4M ~1%

结论:Qwen 是 2026 年社区 base model,Derivatives 最多

State of Open Models: Summer 2026 要点

  • Qwen 成为社区 base model(衍生模型最多)
  • Small models仍是 practical layer:llama.cpp + ggml 在 2026 年 2 月加入 HF,项目完全开源、社区治理
  • 注意力 ≠ 采纳:高星标不等于高下载
  • 国家格局:美国 + 中国主导;韩国政府 AI sovereignty 计划(LG AI Research、SK Telecom、Naver Cloud 等)

引用:https://huggingface.co/blog/state-of-open-models-summer-2026

  • YOLO26(Ultralytics):NMS-free 统一模型族,检测/分割/姿态估计
  • SmolDocling(IBM Granite):256M 参数端到端文档转换
  • PagedAttention/vLLM 原始论文仍在 HF papers trending

引用:https://huggingface.co/papers/trending


六、分类标签

推理引擎 | vLLM | SGLang | LMDeploy | KV Cache | Vector DB
RAG | Agent Stack | Hugging Face | arXiv | Substack
Memory Systems | Quantization | Benchmark | Production

七、建议写入路径

  • 已写入:/shared/research-kb/inbox/jay/2026-09-29-jay-afternoon-briefing-inference-vecdb-stack2026.md

八、后续行动建议

优先级 行动 理由
🔴 高 精读 arXiv:2609.23130v1(Inference Control Plane 综述) 系统性梳理 vLLM → llm-d 演进,对理解 2026 推理基础设施格局至关重要
🔴 高 追踪 SmartGen (2607.28150) + RetroInfer (VLDB 2026) disaggregated inference + 长上下文向量存储,生产风向标
🟡 中 审稿 SwiftCache vs vLLM layer-major layout 限制 如果做多 GPU 异构推理调度,直接相关
🟡 中 核实 SGLang v1.2.1 + vLLM v0.5.20 最新 release notes 验证 prefix caching / grammar cache 差异是否已变化
🟢 低 更新 Vector DB 选型决策树 当前推荐(pgvector/Qdrant/Pinecone/Milvus)已覆盖主流场景

Jay · 研究知识库 · 2026-09-29 13:35 CST 数据来源:arXiv · Hugging Face · Substack · AIMultiple · Spheron · Winder AI · PremAI · Redis Blog · Actian · VectorDBBench