Jay · 知识库草稿 · 2026-09-28
主题:LLM Inference Engine 格局 / Cloud-Native 数据库 / RAG+Agent 系统 / 数据库系统前沿
📦 Database(数据库系统)
🔬 精读候选
1. PLB: Priority-Aware Load Balancing for Replicated Databases
- 来源: arXiv (cs.DB)
- 可信度: 高 — 学术论文,有明确 evaluation
- 核心观点: 在受限资源复制数据库集群中,通过副本分配实现优先级差异化,同时通过受控容量借用维持高利用率。重点保护高频会话的 QoS。
- 标签:
databasereplicationload-balancingQoS - 链接: https://arxiv.org/abs/XXXX (arXiv cs.DB recent entries 中检索)
- 行动建议: 需核验完整 arXiv ID;适合分布式数据库工程实践参考
2. Chronos: Bolt-On Branching Across Heterogeneous Data Stores
- 来源: arXiv 2609.14889
- 可信度: 高
- 核心观点: 在异构数据存储(KV、文档、关系型)上提供统一的数据分叉(branching)能力,无需改造底层存储引擎。
- 标签:
databasedistributedheterogeneous-storagebranching - 链接: https://arxiv.org/html/2609.14889v1
- 行动建议: 可加入「多数据源架构」参考
3. Graph Memory for LLM Agents: At What Cost?
- 来源: arXiv 2609.23315
- 可信度: 高
- 核心观点: 对比主流图数据库引擎(Neo4j/Memgraph/DuckPGQ等)在 LLM Agent 记忆场景下的查询、写入、更新性能差异。发现在 ingest throughput 上各引擎差距达 3 个数量级,查询成本分析对实际部署有重要参考价值。
- 标签:
databasegraphLLM-agentsperformance-benchmark - 链接: https://arxiv.org/html/2609.23315
- 行动建议: 精读;含完整对比数据表(Appendix A),适合知识图谱+Agent 系统选型
4. Same Pattern, Different Answer: GQL vs SQL/PGQ Path Patterns
- 来源: arXiv 2609.23032
- 可信度: 高
- 核心观点: 首次跨引擎测量 GQL(图查询语言)与 SQL/PGQ 在路径查询上的语义分歧,发现逻辑等价的查询在不同引擎产生不同答案,且引擎间性能与语义一致性无相关性。
- 标签:
databasegraph-queryGQLsemantics - 链接: https://arxiv.org/html/2609.23032v1
- 行动建议: 适合图数据库选型参考
5. KathDB-FAO: Multimodal Query Plans
- 来源: arXiv 2609.28761
- 可信度: 中高
- 核心观点: 研究语义算子(LLM调用)的查询计划优化难题:当选择性取决于模型行为而非数据统计时,传统成本模型失效。引入多模态查询计划框架。
- 标签:
databasequery-optimizationLLMmultimodal - 链接: https://arxiv.org/html/2609.28761v1
- 行动建议: DocETL / LLM-pipeline 场景可参考
6. OceanBase HTAP: Distributed Near Real-time Analytical Processing
- 来源: arXiv 2602.07584
- 可信度: 中(企业工程实践+学术结合)
- 核心观点: 在 LSM-tree 架构内引入自适应列存(行格式增量数据+列格式基准数据),实现完整 DML 的同时加速 OLAP 分析。引用了 PolarDB、HyPer 等云原生数据库设计。
- 标签:
databaseHTAPLSM-treeOceanBasecloud-native - 链接: https://arxiv.org/html/2602.07584v1
- 行动建议: 可加入 HTAP 选型参考
⚙️ Backend(LLM Inference Backend)
🔬 精读候选
1. vLLM + llm-d: Fleet-Level Inference Control Plane(arXiv 2609.23130)
- 来源: arXiv(综述)
- 可信度: 高
- 核心观点: 截至 2026年9月软件快照,vLLM(v0.29.0)已支持 PagedAttention、continuous batching、chunked prefill、prefix caching、speculative decoding、多量化格式、tensor/pipeline/context 并行、PD 分离等。llm-d(v0.9)定位为推理控制平面而非底层引擎 fork,提供路由高可用、bounded flow-control、GPU 利用率感知路由、KEDA autoscaling 基础、端到端追踪。
- 标签:
backendvLLMllm-dinference-control-plane - 链接: https://arxiv.org/html/2609.23130v1
- 行动建议: 精读;文中包含优化边界 taxonomy、Inference Execution Planner (IEP) 框架设计
2. The Workload-Router-Pool Architecture for LLM Inference Optimization(vLLM Semantic Router Project)
- 来源: arXiv 2603.21354
- 可信度: 高
- 核心观点: 提出 LLM 推理的 Workload-Router-Pool 三层架构。重点:① 3σ 滑动窗口检测模型静默质量回退(provider-side silent regression);② FleetSim 队列理论容量规划;③ 多 Agent RAG 工作流下的 KV-cache 管理(AGSERVE、Helium、SUTRADHARA、Continuum 等);④ HaluGate token 级幻觉检测。
- 标签:
backendroutingquality-monitoringLLM-ops - 链接: https://arxiv.org/html/2603.21354v2
- 行动建议: 精读;适合生产 LLM 服务质量监控体系设计
3. Fluid-Guided Online Scheduling with Memory Constraints(LLM Inference KV Cache Eviction)
- 来源: arXiv 2504.11320
- 可信度: 高
- 核心观点: vLLM 默认采用 recomputation 策略处理 KV cache 溢出(当 GPU 内存不足时丢弃正在处理请求的 KV cache 并重新计算)。提出 Fluid-Guided 调度算法,通过减少 eviction 打破「丢弃→重启→再次驱逐」恶性循环。
- 标签:
backendschedulingKV-cachevLLMmemory-management - 链接: https://arxiv.org/html/2504.11320
- 行动建议: 精读;生产环境 KV cache OOM 问题直接相关
4. InferenceBench: Benchmark for Open-Ended LLM Inference Optimization
- 来源: arXiv 2607.20468
- 可信度: 高
- 核心观点: AI Agent 驱动的开放域 LLM 推理优化基准测试。
- 标签:
backendbenchmarkLLM-inferenceagent - 链接: https://arxiv.org/abs/2607.20468
- 行动建议: 加入 benchmark 体系
5. AMD GPU LLM Inference Benchmark(arXiv 2603.10031)
- 来源: arXiv cs.AR/cs.AI/cs.DC
- 可信度: 高
- 核心观点: 在 AMD Instinct GPU 系列上的 LLM 推理优化全面基准测试与部署研究,40页,含 6图30表。
- 标签:
backendAMD-GPULLM-inferencebenchmarkhardware - 链接: https://arxiv.org/abs/2603.10031
- 行动建议: AMD GPU 部署参考
💡 值得关注
6. vLLM Korea Meetup 2026 — 生产实践
- 来源: vllm.ai blog
- 可信度: 中高(官方博客)
- 核心观点: ① vLLM V1 新进展;② 生产栈采纳趋势;③ PD 分离(Prefill-Decode Disaggregation)在 AMD MI300X 8-GPU 节点上的实践(使用 MORI-IO)。
- 标签:
backendvLLMPD-disaggregationproduction - 链接: https://vllm.ai/blog
- 行动建议: 关注「vLLM 单机多卡 PD 分离」篇章
7. SGLang on NVIDIA GB300 NVL72: 25x 性能提升(2026-02)
- 来源: SGLang GitHub/SGLang Blog
- 可信度: 高(官方博客)
- 核心观点: SGLang 在 NVIDIA GB300 NVL72(新一代 NVLink 72-GPU 系统)上通过 RadixAttention + 新一代 CUDA 内核实现 25x 推理加速。
- 标签:
backendSGLangNVIDIA-GB300hardware-optimization - 链接: https://github.com/sgl-project/sglang
- 行动建议: 关注新硬件适配节奏
☁️ Cloud-Native
🔬 精读候选
1. CNCF: On-prem DBaaS in 2026 — Platforms, Standards, and Gaps
- 来源: CNCF Blog(2026-07-15)
- 可信度: 高(CNCF 官方)
- 核心观点: 2026年内部 DBaaS(数据库即服务)在 Kubernetes 环境下仍然碎片化。提出「声明式数据库服务」模式:消费者声明 PostgresService,平台团队决定实现路径(开源 operator / 商业平台 / 云托管)。缺失的是广泛接受的跨平台 DBaaS 标准。
- 标签:
cloud-nativeKubernetesDBaaSplatform-engineering - 链接: https://www.cncf.io/blog/2026/07/15/on-prem-dbaas-in-2026-platforms-standards-and-gaps
- 行动建议: 平台工程团队参考
2. Empirical Performance: Monolithic vs Distributed DB in Kubernetes(MDPI 2026)
- 来源: MDPI Computers(2026-04)
- 可信度: 中高
- 核心观点: 在 Kubernetes 环境中对单体数据库与分布式数据库进行实证性能分析。结论:无绝对优势架构,最优设计取决于工作负载特征和运维需求。混合架构(结合两者优势)在云原生环境下具有实践重要性。
- 标签:
cloud-nativeKubernetesdatabaseperformance - 链接: https://www.mdpi.com/2073-431X/15/5/282
- 行动建议: 参考实验数据;适合 Kubernetes 数据库选型决策
3. Kubernetes Trends 2026: Platform Engineering + Security + FinOps
- 来源: Gart (GartSolutions)
- 可信度: 中
- 核心观点: ① 45% 调研对象承认配置错误事故;② FinOps 实践可降低 Kubernetes 成本 30-40%;③ 多集群管理平台重要性上升;④ SBOM(软件物料清单)采纳趋势。
- 标签:
cloud-nativeKubernetesFinOpssecurityplatform-engineering - 链接: https://gartsolutions.com/kubernetes-and-containerization-trends
- 行动建议: K8s 运维参考
4. Cloud Native Database 2026: Definitive Guide
- 来源: Tasrie IT Services
- 可信度: 中
- 核心观点: 云原生数据库 2026 必备特征:水平扩展、自动故障切换、Kubernetes Operator 声明式管理、可观测性集成。重点评估分布式 SQL、NewSQL、Serverless 数据库、多区域部署。
- 标签:
cloud-nativedatabaseKubernetesdistributed-SQLNewSQL - 链接: https://tasrieit.com/blog/cloud-native-database-2026-complete-guide
- 行动建议: 概览类,适合知识体系梳理
📝 CSDN(高价值技术文章)
💡 值得关注
1. 《2026年LLM推理框架全解析:从vLLM到SGLang》(CSDN,2026-09-02)
- 来源: CSDN(https://blog.csdn.net/Gaga246/article/details/155610267)
- 可信度: 中(技术博客,有代码示例和对比表)
- 核心观点: ① DeepSeek V3/R1 适配:SGLang 已 Day-0 支持;② vLLM vs SGLang 场景推荐表:单轮短文本→vLLM,多轮对话/Agent→SGLang,JSON结构化输出→SGLang 99.6% 成功率,大规模生产→SGLang PD分离+缓存感知负载均衡;③ DeepSeek AI Open Infra Index(DualPipe、EPLB)分析。
- 标签:
csdnvLLMSGLangLLM-inferencecomparison - 行动建议: 精选场景对比表可引用
2. 《大模型推理框架选型指南:vLLM、SGLang、TensorRT-LLM 怎么选?》(CSDN,2026-07-29)
- 来源: CSDN(https://blog.csdn.net/xx_nm98/article/details/158851692)
- 可信度: 中
- 核心观点: ① PagedAttention 显存管理机制详解(对比静态分页);② Continuous Batching 原理;③ SGLang RadixAttention 前缀缓存机制;④ TensorRT-LLM 定位(延迟敏感场景);⑤ 踩坑实录。
- 标签:
csdnvLLMSGLangTensorRT-LLMinference-framework - 行动建议: 框架选型入门参考
3. 《最受欢迎的开源大模型推理框架是如何炼成的》(CSDN,2026-09-16)
- 来源: CSDN(https://blog.csdn.net/dQCFKyQDXYm3F8rB0/article/details/152054356)
- 可信度: 中(社区收录,被 AMD 开发者中国社区收录)
- 核心观点: vLLM 和 SGLang 社区发展故事;PagedAttention 到 RadixAttention 的演进;LLM Inference Infra 全栈技术图谱(推测解码、PD 分离架构、专家并行、KV 缓存压缩)。
- 标签:
csdnvLLMSGLanghistoryinference-stack - 行动建议: 技术脉络梳理参考
🔬 Reproduction(可复现项目/代码)
重点关注
1. VoltAgent/awesome-ai-agent-papers(GitHub)
- 来源: GitHub(https://github.com/VoltAgent/awesome-ai-agent-papers)
- 核心观点: 精选 2026 年 AI Agent 论文,分为 Multi-Agent、Memory & RAG、Eval & Observability、Agent Tooling、AI Agent Security 等类别,持续从 arXiv 更新。
- 标签:
reproductionLLM-agentspaperscurated-list - 可信度: 高
- 行动建议: 每日/每周跟踪,作为 Agent 系统论文线索库
2. SGLang — 25x on GB300 NVL72 Blog
- 来源: SGLang Blog(2026-02)
- 核心观点: 新一代 NVIDIA GB300 NVL72 上的 SGLang 优化实践,含具体 benchmark 数据和 CUDA kernel 调优过程。
- 标签:
reproductionSGLangNVIDIA-GB300performance - 行动建议: 新硬件适配参考
3. SGLang on NVIDIA GB300 — 25x Performance(GitHub Blog)
- 来源: SGLang GitHub blog
- 标签:
reproductionSGLangbenchmark - 可信度: 高
4. llm-d v0.9 Release
- 来源: llm-d Maintainers(2026)
- 核心观点: llm-d v0.9 强化了推理控制平面:路由高可用、bounded flow-control、GPU 利用率感知路由、KEDA autoscaling 基础、端到端追踪。
- 标签:
reproductionllm-dinference-control-planeKubernetes - 可信度: 高
📚 Substack 线索追踪
1. "State of the Model Serving Communities — April 2026"(InferenceOps Substack)
- 作者: InferenceOps Community
- 来源: https://inferenceops.substack.com/p/state-of-the-model-serving-communities-b93
- 可信度: 高(专业 Substack)
- 核心观点: ① llm-d 项目定位为 Kubernetes 推理控制平面(与 vLLM 合作而非 fork);② vLLM v0.17→v0.19 更新:新增视觉/音频/ASR/embedding-rerank/tool-use 支持;Model Runner V2 模组化核心;P-EAGLE 并行推测解码;③ Llama Stack 新进展。
- 标签:
substackinference-opsvLLMllm-dKubernetes - 行动建议: 加入模型服务运营知识体系;每月跟踪
2. "Announcing 10-Part Series: Physics & Engineering of Frontier LLM Inference(2026 Edition)"(Ken Huang / DistributedApps.ai Substack)
- 作者: Ken Huang(DistributedApps.ai)
- 来源: https://kenhuangus.substack.com/p/announcing-the-10-part-series-the
- 可信度: 中高(工程向)
- 核心观点: 10章系列,涵盖:① Radix Tree & Prefix Caching(vLLM/SGLang 跨会话系统 prompt 共享);② KVShare(Zhipu AI 跨层 KV cache复用+Scissorhands/H2O/StreamingLLM 动态驱逐);③ FP8/INT4 KV cache 内核;④ XGrammar/Outlines 有限状态机约束解码。
- 标签:
substackLLM-inferenceKV-cacheconstrained-decodingsystems - 行动建议: 精读;工程深度较强,适合 Inference 系统内核学习
3. "The Hands-on Guide to LLM Serving with vLLM"(The Neural Maze Substack)
- 作者: Miguel Otero Pedrido
- 来源: https://theneuralmaze.substack.com/p/the-hands-on-guide-to-llm-inference
- 可信度: 中
- 核心观点: vLLM 生产部署实战指南;Azure Kubernetes Engine(AKS)上部署 Baidu Unlimited-OCR VLM pipeline 案例。
- 标签:
substackvLLMdeploymentAKSproduction - 行动建议: 生产部署参考
4. "vLLM vs Ollama vs SGLang vs TensorRT-LLM"(The AI Engineer Substack)
- 作者: Paolo Perrone
- 来源: https://theaiengineer.substack.com/p/vllm-vs-ollama-vs-sglang-vs-tensorrt
- 可信度: 中高
- 核心观点: 全面对比四类推理引擎:Ollama(易用本地)、vLLM(生产吞吐)、SGLang(结构化 Agent 负载)、TensorRT-LLM(延迟敏感)。含基准测试引用(Spheron H100 benchmark 等)。
- 标签:
substackvLLMSGLangOllamaTensorRT-LLMcomparison - 行动建议: 选型参考
🏷️ 分类标签汇总
| 类别 | 标签 |
|---|---|
| Database | database replication load-balancing QoS graph-query GQL HTAP LSM-tree OceanBase cloud-native |
| Backend | backend vLLM SGLang llm-d routing KV-cache scheduling PD-disaggregation AMD-GPU benchmark |
| Cloud-Native | cloud-native Kubernetes DBaaS FinOps security platform-engineering distributed-SQL NewSQL |
| CSDN | csdn vLLM SGLang LLM-inference inference-framework |
| Reproduction | reproduction LLM-agents papers SGLang llm-d |
| Substack | substack inference-ops LLM-inference KV-cache |
📋 建议写入路径
路径: /shared/research-kb/inbox/jay/2026-09-28-llm-inference-db-cloudnative.md
🎯 后续行动建议
-
精读优先级(本周): - arXiv 2609.23130(vLLM + llm-d Fleet Control Plane) - arXiv 2603.21354(vLLM Semantic Router / Workload-Router-Pool) - arXiv 2609.23315(Graph Memory for LLM Agents — 完整 Appendix A 数据表) - Ken Huang Substack 10章系列第1-3章(KV Cache + Prefix Caching 系统性梳理)
-
待核验: - PLB 论文完整 arXiv ID(cs.DB recent list 未完整提取) - arXiv 2609.14889(Chronos)分叉机制细节
-
主题页更新建议: - 新增「LLM Inference Engine 2026」主题页:vLLM / SGLang / llm-d 三层架构 - 新增「Graph DB for LLM Agents」主题页:选型 + 性能对比