Jay · 研究简报 · 2026-09-29 下午
主题:推理引擎格局 · KV Cache 系统 · Vector DB 生产基准 · Agent Stack 2026 检索范围:GitHub Trending · Hugging Face Blog/Daily · arXiv (Sep 2026) · Substack · 技术博客 本次产出:1 个草稿文件
一、推理引擎格局(2026 年 9 月快照)
vLLM vs SGLang vs LMDeploy
| 维度 | vLLM | SGLang | LMDeploy |
|---|---|---|---|
| GitHub Stars | ~92,700 | ~36,400 | — |
| H100 吞吐 | ~12,500 tok/s | ~16,200 tok/s | ~16,200 tok/s |
| Prefix Cache | xgrammar (默认) | grammar-cache 复用更激进 | — |
| 生态成熟度 | ✅ 最成熟 | 增长快 | 量化受限硬件最优 |
| RadixAttention 优势 | PagedAttention | 前缀共享 >60% 时显著领先 | — |
| Structured Output | xgrammar 默认 | schema 复用更激进 | — |
关键结论(来源:PremAI Blog、Spheron Blog、Winder AI):
- SGLang + LMDeploy 吞吐量相当,均领先 vLLM 约 29%
- vLLM 凭借生态成熟度和 broad hardware support 仍是生产首选
- SGLang 的 RadixAttention 在 prefix-heavy traffic(系统 prompt、RAG doc、tool definition block 共享)场景下优势明显
- 决策阈值:前缀重叠率 >60% → SGLang;否则两者差距 <5%
- SGLang v1.2.1(2026 年 9 月)、vLLM v0.5.20、LMDeploy 持续更新
引用: - https://www.premai.io/blog/vllm-vs-sglang-vs-lmdeploy-fastest-llm-inference-engine-in-2026 - https://www.spheron.network/blog/vllm-vs-sglang-2026 - https://winder.ai/vllm-vs-ollama-vs-sglang-llm-inference-comparison
二、KV Cache 系统研究前沿(arXiv Sep 2026)
高价值论文
1. Inference Control Plane 综述
标题:From Inference Engine to Inference Control Plane: Connecting vLLM, llm-d, and the Evolution of Efficient Distributed LLM Serving
来源:arXiv:2609.23130v1 · Sep 19, 2026 · Twinkll Sisodia (BU)
核心观点:
- vLLM 已成为 control plane 的基础层而非终点
- llm-d (LLM disaggregation) 正在将 prefill/decode 物理分离,形成新的调度平面
- OpenTelemetry tracing 现在覆盖 Gateway → Endpoint Picker → KV-cache → prefill/decode proxy → model server 全路径
可信度:高(arXiv 2026-09 新文,系统综述)
2. Fluid-Guided Online Scheduling
标题:Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints
来源:arXiv:2504.11320v4 · Aug 24, 2026 · MIT + Alibaba
核心观点:
- 研究 KV cache 内生增长导致的内存溢出问题
- 提出 (Nested) WAIT 调度算法:通过 fluid approximation 识别 fluid stability region
- 维度约减:诱导 threshold/boundary queues 捕获 memory-coupled state
- 适用场景:prefill/decode 混合高并发场景
可信度:高(MIT + 阿里联合,PGM 2026 接收)
3. Sustainable Distributed LLM Inference
标题:Toward Sustainable Distributed LLM Inference: A Systems Synthesis and Research Agenda for an Energy-, Carbon-, and Cache-Aware llm-d Control Plane
来源:arXiv:2609.05565v1 · Sep 2026
核心观点:
- KV state 既性能资产也是能源资产
- Peer-to-peer cache sharing 策略:比较复制 KV vs 重计算的跨点(延迟/能量/内存压力/未来复用概率)
- Phase-specific power control:prefill/decode 阶段 compute/memory 行为不同,能源最优点也不同
- PowerSlider 研究了 time-varying power caps 下的阶段不对称性
可信度:中高(系统综述,文献覆盖 2023-2026)
4. Multi-Segment Attention
标题:Multi-Segment Attention: Enabling Efficient KV-Cache Management for Faster LLM Serving
来源:arXiv:2606.02964v1
核心观点:
- block-level, position-aware eviction,可与 Autellix/Tokencake/ThunderAgent 调度层无缝集成
- 与 Continuum(TTL pin)、KVFlow(Agent Step Graph)、InferCept(min-waste eviction)正交
可信度:中高(新研究方向)
5. ObjectCache: Layerwise Object-Storage Retrieval
标题:ObjectCache: Layerwise Object-Storage Retrieval for KV Cache Reuse
来源:arXiv:2605.22850v1
核心观点:
- S3-compatible object tier 的 KV chunk 获取、聚合和按 layer 顺序交付
- 解决 S3 读取延迟问题(TCP/RDMA-buffer/RDMA-direct 多路径)
- 与 CacheFlow(重计算加速)解决不同瓶颈
可信度:中高
6. SmartGen: Selective KV Cache Transfer
标题:SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer
来源:arXiv:2607.28150v1 · Jul 30, 2026
核心观点:
- disaggregated prefill/decode 环境下的选择性 KV transfer
- RetroInfer (VLDB 2026):向量存储引擎支持长上下文 LLM inference
可信度:高(VLDB 引用)
7. SwiftCache: Heterogeneous KV Cache Sharing
标题:SwiftCache: Efficient LLM Serving for Multi-turn Conversations with Heterogeneous KV Cache Sharing
来源:arXiv:2606.16135v1
核心观点:
- 单服务器多 GPU 异构模型协调
- 利用 GPU 内存和 GPU 间带宽资源
- 指出 vLLM/SGLang/LightLLM 的 layer-major layout 不支持动态 KV cache resize
可信度:中高
8. "Internet for KV Cache"
标题:An Internet for the KV Cache: Rethinking Classical Infrastructure Boundaries in the LLM Inference Age
来源:arXiv:2608.01526v1
相关引用:
- CXL-SpecKV (FPGA speculative KV-Cache, FPGA 2026)
- DroidSpeak (NSDI 2026):跨 fine-tuned 模型变体共享 KV cache
可信度:高(NSDI 引用)
三、Vector DB 生产基准(2026 年 9 月)
格局总结(按场景推荐)
| 场景 | 推荐方案 | 理由 |
|---|---|---|
| 现有 PostgreSQL(<5000 万向量) | pgvector + pgvectorscale | 继承 WAL,写入便宜 |
| 新项目 <1000 万向量 | Qdrant Cloud(最佳免费层)或 ChromaDB(原型快速) | 性能 vs 简易平衡 |
| 新项目 1000-1 亿向量 | Pinecone(零运维)或 Weaviate(混合搜索)或 Milvus(自托管成本低) | 按规模分层 |
| 新项目 >1 亿向量 | Milvus/Zilliz Cloud 或 Pinecone serverless | 分布式横向扩展 |
| 电商搜索(1 亿产品) | Milvus(分布式)或 Qdrant(过滤查询) | 规模 + 过滤 |
| 企业知识搜索 | Weaviate 或 pgvectorscale(混合搜索) | 语义 + 关键词 |
| LLM 语义缓存 | Redis(超低延迟) | 毫秒级 |
| 合规敏感(数据主权) | Pinecone BYOC 或 pgvectorscale | 自己控制 |
| 研究/原型 | Chroma(免费,无需配置) | 零配置 |
关键性能数字
- Redis(1 亿向量 benchmark):90% 精度时 median latency ~200ms(top 100 NN),QPS ~66,000 insertion/s;95% 精度时 median latency ~1.3s
- Qdrant(1M 向量,1536 维,~550 QPS):p50 ~2.9ms,p99 ~8ms
- Weaviate(同等规模):p50 ~5.8ms,p99 ~18ms
- pgvector:适合 <5000 万向量,PostgreSQL 生态内统一管理
2026 新动态
- VectorDBBench(Zilliz,v2.0.0,Sep 10 2026):新增 pgvector CLI 模块,1760 commits
- Benchmark 可信度警告(Actian):供应商自评 benchmark 存在架构特定优势伪装成客观比较的问题
- 方向:2026 年倾向于集成平台(PostgreSQL + pgvector → 企业混合引擎),而非"纯向量"孤岛
引用: - https://redis.io/blog/best-open-source-vector-databases-comparison - https://www.actian.com/blog/databases/how-to-evaluate-vector-databases-in-2026 - https://github.com/zilliztech/vectordbbench
四、Substack 高价值工程洞察
1. AI Agents Stack 2026(The AI Engineer,Eric Roby)
核心观点: - Agent ≠ Chatbot:Agent 需要跨多步执行的 state management、tool access governed by protocols、跨 session 持久 memory、autonomous reasoning loops、real-time guardrails - Memory 三层架构(2024:"选个向量数据库做 RAG" → 2026:第一性 architectural primitive): 1. In-context cache(前缀共享) 2. Session memory(短期状态) 3. Long-term memory(持久化知识) - Guardrails 独立学科:不再是输入/输出过滤器,而是授权 tool calls、执行 rate limits、验证 agent 实际行为 - Model Routing 模式:分类/分流用小模型;困难推理用 frontier 模型;嵌入/评估用专用模型。RouteLLM 研究表明 routing 可显著降低成本同时保留质量 - 生产代理很少用单一模型
可信度:高(AI Engineer newsletter,Eric Roby,圈内广泛阅读)
引用:https://theaiengineer.substack.com/p/the-ai-agents-stack-2026-edition
2. RAG 2026 完整指南
核心观点: - Cross-Encoder Re-Ranking:Hybrid Search 先 top-100 候选,再用 Cross-Encoder (Cohere Rerank 3.5 / BGE-Reranker) 精排到 top-5~10 - 语义分块(Semantic Chunking):不用固定字符数切割,而用 cosine distance 阈值检测主题转换 - Small-to-Big (Parent-Document) 策略:子 chunk(100 tokens)精确匹配 + 父文档(整页)丰富上下文 - Hybrid Search + RRF:向量 + 关键词 + 元数据信号融合,Reciprocal Rank Fusion 合并排名 - Naive RAG(纯语义相似度)在生产中已很少部署
可信度:高(Shrinivasan,AI 工程师圈广泛引用)
引用:https://aishwaryasrinivasan.substack.com/p/all-you-need-to-know-about-rag-in
3. AI Agents Stack 选择指南(The Nuanced Perspective)
核心观点: - 2026 评估转向 continuous improvement loop:LLM judges 成为 grading agent 输出的默认方式 - Arize Alyx:从 observability 数据中自动发现常见失败模式 - Managed retrieval 服务(如 Glean):RAG + 传统搜索融合,牺牲 knob-level control换取免运维 - Hybrid search = 生产默认,naive RAG 已过时
引用:https://thenuancedperspective.substack.com/p/how-to-choose-your-ai-agent-stack
4. 学术 × AI 代理(Kiran Garimella)
核心观点: - LLM 用户论文产量提升 30-90%(Science 研究,Kusumegi et al.) - 但:传统质量信号与实际质量正在脱钩——LLM 写作语言复杂度更高,但被接收率反而更低 - 对已有训练的研究者(已具备判断力)增益更大;对新手风险更高(产出量大但质量低)
可信度:高(Science peer-reviewed 研究)
引用:https://kirangarimella.substack.com/p/ai-agents-and-academia
五、Hugging Face 生态快照(Sep 2026)
模型下载排行(30 天)
| 家族 | 下载量 | 占比 |
|---|---|---|
| Qwen | 240.1M | ~60% |
| Gemma | 32.6M | ~8% |
| Ornith | 13.5M | ~3% |
| Llama | 12.9M | ~3% |
| gpt-oss | 11.9M | ~3% |
| DeepSeek | 4.4M | ~1% |
结论:Qwen 是 2026 年社区 base model,Derivatives 最多
State of Open Models: Summer 2026 要点
- Qwen 成为社区 base model(衍生模型最多)
- Small models仍是 practical layer:llama.cpp + ggml 在 2026 年 2 月加入 HF,项目完全开源、社区治理
- 注意力 ≠ 采纳:高星标不等于高下载
- 国家格局:美国 + 中国主导;韩国政府 AI sovereignty 计划(LG AI Research、SK Telecom、Naver Cloud 等)
引用:https://huggingface.co/blog/state-of-open-models-summer-2026
HF Trending Papers(近期亮点)
- YOLO26(Ultralytics):NMS-free 统一模型族,检测/分割/姿态估计
- SmolDocling(IBM Granite):256M 参数端到端文档转换
- PagedAttention/vLLM 原始论文仍在 HF papers trending
引用:https://huggingface.co/papers/trending
六、分类标签
推理引擎 | vLLM | SGLang | LMDeploy | KV Cache | Vector DB
RAG | Agent Stack | Hugging Face | arXiv | Substack
Memory Systems | Quantization | Benchmark | Production
七、建议写入路径
- 已写入:
/shared/research-kb/inbox/jay/2026-09-29-jay-afternoon-briefing-inference-vecdb-stack2026.md
八、后续行动建议
| 优先级 | 行动 | 理由 |
|---|---|---|
| 🔴 高 | 精读 arXiv:2609.23130v1(Inference Control Plane 综述) | 系统性梳理 vLLM → llm-d 演进,对理解 2026 推理基础设施格局至关重要 |
| 🔴 高 | 追踪 SmartGen (2607.28150) + RetroInfer (VLDB 2026) | disaggregated inference + 长上下文向量存储,生产风向标 |
| 🟡 中 | 审稿 SwiftCache vs vLLM layer-major layout 限制 | 如果做多 GPU 异构推理调度,直接相关 |
| 🟡 中 | 核实 SGLang v1.2.1 + vLLM v0.5.20 最新 release notes | 验证 prefix caching / grammar cache 差异是否已变化 |
| 🟢 低 | 更新 Vector DB 选型决策树 | 当前推荐(pgvector/Qdrant/Pinecone/Milvus)已覆盖主流场景 |
Jay · 研究知识库 · 2026-09-29 13:35 CST 数据来源:arXiv · Hugging Face · Substack · AIMultiple · Spheron · Winder AI · PremAI · Redis Blog · Actian · VectorDBBench