研究简报 · Jay · 2026-10-10 下午批次(15:05)

本次主题

推理引擎决策框架 · 向量数据库实测对比 · Multi-Agent 框架选型 · 学术论文 · 云原生趋势


检索范围

  • Tavily: 推理引擎、向量数据库 benchmark、Multi-Agent 框架、arXiv 论文、云原生基础设施
  • 覆盖来源: Spheron、aiml.qa、ayautomate、emasterlabs、arXiv、isovalent、CNCF blog

📦 database

🔴 KEEP — pgvector vs Qdrant 2026 实测对比(50M 向量量级)

来源: aiml.qa (June 26, updated Sep 6, 2026) URL: https://aiml.qa/blog/qdrant-vs-pgvector 质量评分: 9/10

关键数据(同硬件 AWS r6id.4xlarge 16vCPU,ANN-Benchmarks 工具,50M 向量 768dims):

配置 QPS p95 延迟 p99 延迟
pgvectorscale (HNSW m=16) 471 60ms 75ms
Qdrant (self-hosted) 41 37ms 39ms

重要结论: - pgvectorscale 吞吐量是 Qdrant 的 11.4 倍(471 vs 41 QPS),但 Qdrant 尾延迟(p95/p99)更优 - pgvector 适合:已有 Postgres 基础设施、需要向量与业务表 JOIN、≤5000万向量场景 - Qdrant 适合:百亿级向量 + 需亚 50ms P99 + 快速 HNSW 构建(重建索引无需维护窗口) - 两者可组合使用:pgvector 跑小规模 + Qdrant 承接热数据

可信度: 高(实测数据,ANN-Benchmarks fork 工具链开放) 建议: 入知识库向量数据库选型参考表;精读 benchmark 方法论章节


🟡 BORDERLINE — Vector DB Benchmark 2026 综合排行(Salt Technologies)

来源: https://www.salttechno.ai/datasets/vector-database-performance-benchmark-2026 质量评分: 6.5/10

关键数据:Qdrant p50 延迟 4ms / p99 25ms(专用向量DB 最优);Redis p50 5ms 但 RAM 受限;Pinecone p50 8ms(全托管) 结论: 数据有参考价值但方法论声明不透明("$3,000 AI Readiness Audit"),适合建立宏观认知而非工程决策


⚙️ backend / inference-engineering

🔴 KEEP — 推理引擎三选一决策框架(Spheron, 2026)

来源: https://www.spheron.network/blog/llm-inference-optimization-2026 质量评分: 9.5/10 分类: inference-engineering benchmark vllm sglang tensorrt-llm

H100 80GB + Llama 3.3 70B FP8 实测(50 并发,unique prompts):

引擎 吞吐量 TTFT p50 冷启动 适用场景
vLLM 1,850 tok/s 380ms ~62s 模型种类多、快速部署、多厂商 GPU
TensorRT-LLM 2,100 tok/s 340ms ~28min 极致吞吐、固定模型、NVIDIA 独占
SGLang 1,920 tok/s 360ms ~58s prefix 共享、长对话、多轮 Agent

重要发现: - SGLang 优势阈值:60% 前缀共享率(SGLang RadixAttention vs vLLM PagedAttention) - Prefix-heavy 场景(50 并发,80% 共享):TTFT p50 从 310ms → 195ms(-37%) - HuggingFace TGI 已于 2026-03-21 归档(read-only),新项目切 vLLM / SGLang

决策树:

固定模型 + NVIDIA + 追求极限吞吐 → TensorRT-LLM
多模型切换 + 快速迭代 + prefix 少 → vLLM(默认)
长对话 + 多轮 Agent + prefix 共享多 → SGLang

可信度: 高(实测数据,版本明确 vLLM v0.18 / SGLang v0.5.9 / cu130) 建议: 精读;纳入推理引擎知识库;关注 SGLang Diffusion(2026-01,支持视频生成推理后端)


🔴 KEEP — LMDeploy vs SGLang vs vLLM 量化对比(Premiere.ai, 2026)

来源: https://www.premai.io/blog/vllm-vs-sglang-vs-lmdeploy-fastest-llm-inference-engine-in-2026 质量评分: 8/10

H100 单卡量化结论: - SGLang / LMDeploy: ~16,200 tok/s(并列第一) - vLLM: ~12,500 tok/s(差距 29%) - 月省 GPU 成本估算:百万请求/天时,29% 吞吐差距 ≈ $15,000/月

建议: SGLang vs LMDeploy 对比可纳入量化选型参考


🤖 multi-agent-frameworks

🔴 KEEP — Multi-Agent 框架 2026 选型矩阵

来源: https://www.ayautomate.com/blog/best-multi-agent-frameworks (June 2026) 质量评分: 8.5/10 分类: agent-engineering langgraph crewai autogen orchestration

2026-10 格局快照:

框架 核心范式 生产成熟度 适用场景 企业案例
LangGraph 有状态图 + 持久化 checkpoint ⭐⭐⭐⭐⭐ 复杂多 Agent、长时 workflow、HITL Anthropic、LinkedIn、Uber
CrewAI Role-based crews ⭐⭐⭐⭐ 快速原型、业务流程自动化 65% Fortune 500 使用开源版
AutoGen 对话式多 Agent ⭐⭐⭐ 研究、Azure 集成 Microsoft 内部
OpenAI Agents SDK 轻量 handoffs ⭐⭐ OpenAI-first 栈 —
Google ADK 工作流图 ⭐⭐⭐ Gemini 生态 —
AWS Strands Agent + Tools ⭐⭐ Bedrock 生态 —

关键判断: - 2026 年所有主要框架都在向 图结构状态管理 收敛 - LangGraph 已成企业生产事实标准(有 durable execution) - CrewAI Enterprise 落地推动 Fortune 500 采纳(内容 pipeline、销售研究、客服分诊) - 补充发现(Splunk):LangGraph 有最陡学习曲线但提供最细粒度执行控制

可信度: 高(综合横向对比,有真实企业案例) 建议: 精读框架选型章节;更新 Multi-Agent 知识库


🔬 reproduction / arxiv

🔴 KEEP — arXiv 近期 AI/Agent 论文(arXiv cs.AI recent)

URL: https://arxiv.org/list/cs.AI/recent (2026-10 本周) 分类: arxiv agent-reasoning tool-call skillForge

高价值条目:

  1. arXiv:2610.09683 — System Switch: When Should a Fast Decision Model Stop and Think? - 方向:Adaptive computation / early exit - 信号:Accepted at INLG 2026

  2. arXiv:2610.09769 — SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles - 方向:Agent skill acquisition / lifelong learning - 信号:Accepted at NeurIPS 2026 Main Poster

  3. arXiv:2610.09832 — How Do Agentic LLMs Decide to Call Tools? A Tool-Call Vector Shaped by Suppression - 方向:Agent tool-use reasoning mechanism - 信号:Accepted at NeurIPS 2026 Main Poster

  4. arXiv:2602.06176 — Large Language Model Reasoning Failures - 会议:TMLR 2026 with Survey Certification - 方向:LLM reasoning failure modes / benchmark - Repo: 有配套代码仓库

可信度: 高(顶会接收 + arXiv 同行评审) 建议: NeurIPS 2026 这两篇值得加入本周精读队列


🟡 BORDERLINE — eMasterLabs 2026 Top 10 LLM 论文综述

URL: https://emasterlabs.com/llm-research-papers 质量评分: 6.5/10

值得记录的观点: - Paper #01: Sparse Attention Horizons: Scaling Context to 1M Tokens(长上下文效率) - Paper #02: Constitutional Distillation(对齐 + 可复现,RLHF 替代方案) - Paper #03: MoE-Fusion: Cross-Layer Expert Sharing(推理效率 +15%,路由开销 -30%) - 总体趋势:2026 年四大主题 = Efficiency / Alignment / Long-context / Multimodal

可信度: 中(综述文,有二次加工;需核对原文) 建议: 定位为研究线索而非一手来源


☁️ cloud-native

🟡 BORDERLINE — KubeCon + CloudNativeCon NA 2026 预览

来源: https://www.cncf.io/blog/2026/10/01/kubecon-cloudnativecon-north-america-2026-build-your-infrastructure-engineer-journey 分类: cloud-native kubernetes cilium ebpf platform-engineering

关键议题: - GPU / AI 基础设施支持成平台工程新标配 - CiliumCon(10/27):eBPF 无 sidecar 服务网格,Hubble 可观测性,Tetragon 安全 - Platform Engineering Day:内部开发者平台(IDP)规模化挑战 - eBPF 趋势:无插桩遥测 + 服务网格流量拦截 + 安全策略内核级执行

可信度: 高(CNCF 官方博客) 建议: 关注 Cilium + eBPF 在 AI 推理服务网格中的落地案例


🟡 BORDERLINE — eBPF 成为云原生基础设施"隐形层"

来源: https://cloudnativenow.com/features/ebpf-the-silent-power-behind-cloud-natives-next-phase (Sep 25, 2026) 核心判断:

"Just as Kubernetes became the control plane of containers, eBPF is becoming the substrate for networking, observability, and security."

2026 eBPF 应用层: - 服务网格:替代 sidecar(Cilium) - 可观测性:零插桩 telemetry - 安全:内核级威胁检测 + 策略执行 - 平台工程:嵌入 golden path,对开发者透明

可信度: 中(行业媒体综述,非一手数据) 建议: 纳入技术雷达 Cloud-Native 层


📋 分类标签

向量数据库 推理工程 inference-engineering benchmark pgvector qdrant vllm sglang langgraph crewai multi-agent arxiv neurips2026 cloud-native ebpf cilium


本批次汇总

分类 KEEP BORDERLINE DROP 合计
database 1 1 0 2
backend/inference 2 1 0 3
multi-agent 1 0 0 1
reproduction/arxiv 1 1 0 2
cloud-native 0 2 0 2
合计 5 5 0 10

建议写入路径

精读队列(本周): 1. arXiv:2610.09769 — SkillForge(NeurIPS 2026,Agent skill 生命周期) 2. arXiv:2610.09832 — Tool-Call Vector(NeurIPS 2026,Agent 工具调用机制) 3. Spheron 推理引擎决策框架(benchmark 数据 + 决策树)

知识库更新建议: - ai-engineering/inference/2026-engine-decision-framework-vllm-sglang-trtllm.md - ai-engineering/vector-databases/2026-pgvector-qdrant-benchmark-50m.md - ai-engineering/agents/2026-multi-agent-framework-matrix.md


Jay · 2026-10-10 15:05 · Asia/Shanghai 未执行 Git 写入 | 草稿路径: /shared/research-kb/inbox/jay/2026-10-10-1505-jay-five-category-briefing.md