研究简报 · Jay · 2026-10-10 下午批次(15:05)
本次主题
推理引擎决策框架 · 向量数据库实测对比 · Multi-Agent 框架选型 · 学术论文 · 云原生趋势
检索范围
- Tavily: 推理引擎、向量数据库 benchmark、Multi-Agent 框架、arXiv 论文、云原生基础设施
- 覆盖来源: Spheron、aiml.qa、ayautomate、emasterlabs、arXiv、isovalent、CNCF blog
📦 database
🔴 KEEP — pgvector vs Qdrant 2026 实测对比(50M 向量量级)
来源: aiml.qa (June 26, updated Sep 6, 2026) URL: https://aiml.qa/blog/qdrant-vs-pgvector 质量评分: 9/10
关键数据(同硬件 AWS r6id.4xlarge 16vCPU,ANN-Benchmarks 工具,50M 向量 768dims):
| 配置 | QPS | p95 延迟 | p99 延迟 |
|---|---|---|---|
| pgvectorscale (HNSW m=16) | 471 | 60ms | 75ms |
| Qdrant (self-hosted) | 41 | 37ms | 39ms |
重要结论: - pgvectorscale 吞吐量是 Qdrant 的 11.4 倍(471 vs 41 QPS),但 Qdrant 尾延迟(p95/p99)更优 - pgvector 适合:已有 Postgres 基础设施、需要向量与业务表 JOIN、≤5000万向量场景 - Qdrant 适合:百亿级向量 + 需亚 50ms P99 + 快速 HNSW 构建(重建索引无需维护窗口) - 两者可组合使用:pgvector 跑小规模 + Qdrant 承接热数据
可信度: 高(实测数据,ANN-Benchmarks fork 工具链开放) 建议: 入知识库向量数据库选型参考表;精读 benchmark 方法论章节
🟡 BORDERLINE — Vector DB Benchmark 2026 综合排行(Salt Technologies)
来源: https://www.salttechno.ai/datasets/vector-database-performance-benchmark-2026 质量评分: 6.5/10
关键数据:Qdrant p50 延迟 4ms / p99 25ms(专用向量DB 最优);Redis p50 5ms 但 RAM 受限;Pinecone p50 8ms(全托管) 结论: 数据有参考价值但方法论声明不透明("$3,000 AI Readiness Audit"),适合建立宏观认知而非工程决策
⚙️ backend / inference-engineering
🔴 KEEP — 推理引擎三选一决策框架(Spheron, 2026)
来源: https://www.spheron.network/blog/llm-inference-optimization-2026
质量评分: 9.5/10
分类: inference-engineering benchmark vllm sglang tensorrt-llm
H100 80GB + Llama 3.3 70B FP8 实测(50 并发,unique prompts):
| 引擎 | 吞吐量 | TTFT p50 | 冷启动 | 适用场景 |
|---|---|---|---|---|
| vLLM | 1,850 tok/s | 380ms | ~62s | 模型种类多、快速部署、多厂商 GPU |
| TensorRT-LLM | 2,100 tok/s | 340ms | ~28min | 极致吞吐、固定模型、NVIDIA 独占 |
| SGLang | 1,920 tok/s | 360ms | ~58s | prefix 共享、长对话、多轮 Agent |
重要发现: - SGLang 优势阈值:60% 前缀共享率(SGLang RadixAttention vs vLLM PagedAttention) - Prefix-heavy 场景(50 并发,80% 共享):TTFT p50 从 310ms → 195ms(-37%) - HuggingFace TGI 已于 2026-03-21 归档(read-only),新项目切 vLLM / SGLang
决策树:
固定模型 + NVIDIA + 追求极限吞吐 → TensorRT-LLM
多模型切换 + 快速迭代 + prefix 少 → vLLM(默认)
长对话 + 多轮 Agent + prefix 共享多 → SGLang
可信度: 高(实测数据,版本明确 vLLM v0.18 / SGLang v0.5.9 / cu130) 建议: 精读;纳入推理引擎知识库;关注 SGLang Diffusion(2026-01,支持视频生成推理后端)
🔴 KEEP — LMDeploy vs SGLang vs vLLM 量化对比(Premiere.ai, 2026)
来源: https://www.premai.io/blog/vllm-vs-sglang-vs-lmdeploy-fastest-llm-inference-engine-in-2026 质量评分: 8/10
H100 单卡量化结论: - SGLang / LMDeploy: ~16,200 tok/s(并列第一) - vLLM: ~12,500 tok/s(差距 29%) - 月省 GPU 成本估算:百万请求/天时,29% 吞吐差距 ≈ $15,000/月
建议: SGLang vs LMDeploy 对比可纳入量化选型参考
🤖 multi-agent-frameworks
🔴 KEEP — Multi-Agent 框架 2026 选型矩阵
来源: https://www.ayautomate.com/blog/best-multi-agent-frameworks (June 2026)
质量评分: 8.5/10
分类: agent-engineering langgraph crewai autogen orchestration
2026-10 格局快照:
| 框架 | 核心范式 | 生产成熟度 | 适用场景 | 企业案例 |
|---|---|---|---|---|
| LangGraph | 有状态图 + 持久化 checkpoint | ⭐⭐⭐⭐⭐ | 复杂多 Agent、长时 workflow、HITL | Anthropic、LinkedIn、Uber |
| CrewAI | Role-based crews | ⭐⭐⭐⭐ | 快速原型、业务流程自动化 | 65% Fortune 500 使用开源版 |
| AutoGen | 对话式多 Agent | ⭐⭐⭐ | 研究、Azure 集成 | Microsoft 内部 |
| OpenAI Agents SDK | 轻量 handoffs | ⭐⭐ | OpenAI-first 栈 | — |
| Google ADK | 工作流图 | ⭐⭐⭐ | Gemini 生态 | — |
| AWS Strands | Agent + Tools | ⭐⭐ | Bedrock 生态 | — |
关键判断: - 2026 年所有主要框架都在向 图结构状态管理 收敛 - LangGraph 已成企业生产事实标准(有 durable execution) - CrewAI Enterprise 落地推动 Fortune 500 采纳(内容 pipeline、销售研究、客服分诊) - 补充发现(Splunk):LangGraph 有最陡学习曲线但提供最细粒度执行控制
可信度: 高(综合横向对比,有真实企业案例) 建议: 精读框架选型章节;更新 Multi-Agent 知识库
🔬 reproduction / arxiv
🔴 KEEP — arXiv 近期 AI/Agent 论文(arXiv cs.AI recent)
URL: https://arxiv.org/list/cs.AI/recent (2026-10 本周)
分类: arxiv agent-reasoning tool-call skillForge
高价值条目:
-
arXiv:2610.09683 — System Switch: When Should a Fast Decision Model Stop and Think? - 方向:Adaptive computation / early exit - 信号:Accepted at INLG 2026
-
arXiv:2610.09769 — SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles - 方向:Agent skill acquisition / lifelong learning - 信号:Accepted at NeurIPS 2026 Main Poster
-
arXiv:2610.09832 — How Do Agentic LLMs Decide to Call Tools? A Tool-Call Vector Shaped by Suppression - 方向:Agent tool-use reasoning mechanism - 信号:Accepted at NeurIPS 2026 Main Poster
-
arXiv:2602.06176 — Large Language Model Reasoning Failures - 会议:TMLR 2026 with Survey Certification - 方向:LLM reasoning failure modes / benchmark - Repo: 有配套代码仓库
可信度: 高(顶会接收 + arXiv 同行评审) 建议: NeurIPS 2026 这两篇值得加入本周精读队列
🟡 BORDERLINE — eMasterLabs 2026 Top 10 LLM 论文综述
URL: https://emasterlabs.com/llm-research-papers 质量评分: 6.5/10
值得记录的观点: - Paper #01: Sparse Attention Horizons: Scaling Context to 1M Tokens(长上下文效率) - Paper #02: Constitutional Distillation(对齐 + 可复现,RLHF 替代方案) - Paper #03: MoE-Fusion: Cross-Layer Expert Sharing(推理效率 +15%,路由开销 -30%) - 总体趋势:2026 年四大主题 = Efficiency / Alignment / Long-context / Multimodal
可信度: 中(综述文,有二次加工;需核对原文) 建议: 定位为研究线索而非一手来源
☁️ cloud-native
🟡 BORDERLINE — KubeCon + CloudNativeCon NA 2026 预览
来源: https://www.cncf.io/blog/2026/10/01/kubecon-cloudnativecon-north-america-2026-build-your-infrastructure-engineer-journey
分类: cloud-native kubernetes cilium ebpf platform-engineering
关键议题: - GPU / AI 基础设施支持成平台工程新标配 - CiliumCon(10/27):eBPF 无 sidecar 服务网格,Hubble 可观测性,Tetragon 安全 - Platform Engineering Day:内部开发者平台(IDP)规模化挑战 - eBPF 趋势:无插桩遥测 + 服务网格流量拦截 + 安全策略内核级执行
可信度: 高(CNCF 官方博客) 建议: 关注 Cilium + eBPF 在 AI 推理服务网格中的落地案例
🟡 BORDERLINE — eBPF 成为云原生基础设施"隐形层"
来源: https://cloudnativenow.com/features/ebpf-the-silent-power-behind-cloud-natives-next-phase (Sep 25, 2026) 核心判断:
"Just as Kubernetes became the control plane of containers, eBPF is becoming the substrate for networking, observability, and security."
2026 eBPF 应用层: - 服务网格:替代 sidecar(Cilium) - 可观测性:零插桩 telemetry - 安全:内核级威胁检测 + 策略执行 - 平台工程:嵌入 golden path,对开发者透明
可信度: 中(行业媒体综述,非一手数据) 建议: 纳入技术雷达 Cloud-Native 层
📋 分类标签
向量数据库 推理工程 inference-engineering benchmark pgvector qdrant vllm sglang langgraph crewai multi-agent arxiv neurips2026 cloud-native ebpf cilium
本批次汇总
| 分类 | KEEP | BORDERLINE | DROP | 合计 |
|---|---|---|---|---|
| database | 1 | 1 | 0 | 2 |
| backend/inference | 2 | 1 | 0 | 3 |
| multi-agent | 1 | 0 | 0 | 1 |
| reproduction/arxiv | 1 | 1 | 0 | 2 |
| cloud-native | 0 | 2 | 0 | 2 |
| 合计 | 5 | 5 | 0 | 10 |
建议写入路径
精读队列(本周): 1. arXiv:2610.09769 — SkillForge(NeurIPS 2026,Agent skill 生命周期) 2. arXiv:2610.09832 — Tool-Call Vector(NeurIPS 2026,Agent 工具调用机制) 3. Spheron 推理引擎决策框架(benchmark 数据 + 决策树)
知识库更新建议:
- ai-engineering/inference/2026-engine-decision-framework-vllm-sglang-trtllm.md
- ai-engineering/vector-databases/2026-pgvector-qdrant-benchmark-50m.md
- ai-engineering/agents/2026-multi-agent-framework-matrix.md
Jay · 2026-10-10 15:05 · Asia/Shanghai 未执行 Git 写入 | 草稿路径: /shared/research-kb/inbox/jay/2026-10-10-1505-jay-five-category-briefing.md