Jay 工程实践筛选 · 2026-10-10 上午批次
筛选标准:真实环境 / 命令 / 错误 / 源码 / 性能数据 / 可复现步骤
✅ KEEP(通过筛选)
1. Spheron · vLLM vs TensorRT-LLM vs SGLang: H100 Benchmarks (2026)
URL: https://www.spheron.network/blog/vllm-vs-tensorrt-llm-vs-sglang-benchmarks
质量评分: 9/10
分类标签: inference-engineering benchmark docker h100 fp8
质量信号: - 实测数据:H100 80GB + Llama 3.3 70B FP8,同硬件三引擎对比 - vLLM: 1,850 tok/s (50 并发), TTFT p50 120ms, 冷启动 ~62s - TensorRT-LLM: 2,100 tok/s, TTFT p50 105ms, 冷启动 ~28min - SGLang: 1,920 tok/s, TTFT p50 112ms, 冷启动 ~58s - 可复现 docker run 命令(所有三个引擎) - 测试方法明确:200 prompts / 60s warmup / 3min 测量窗口 / 4档并发 - 决策树:按 workload 选引擎的量化标准
保留理由:稀缺的三引擎同框 H100 实测数据,带版本号和可复现命令
2. SitePoint · vLLM Production Deployment: Complete 2026 Guide
URL: https://www.sitepoint.com/vllm-production-deployment-guide-2026
质量评分: 9/10
分类标签: inference-engineering docker production security tensor-parallel
质量信号:
- Docker 部署命令:--gpus, --shm-size, --ipc=host, --env-file 安全实践(chmod 600, 禁止明文传参)
- 多卡 TP4 tensor-parallel 部署完整命令
- PagedAttention 原理 + 实际收益数字:7B FP16 H100 80GB 从 30 并发提升到 100+
- benchmark_serving.py 命令及输出解读(Mean TTFT / P99 TTFT / E2E latency)
- --enable-prefix-caching 前缀缓存部署
- 源码:vLLM 官方 benchmark 脚本
保留理由:生产级 vLLM 部署命令集 + 安全实践 + benchmark 实操,稀缺度高
3. Spheron · vLLM vs SGLang 2026: RadixAttention vs PagedAttention
URL: https://www.spheron.network/blog/vllm-vs-sglang-2026
质量评分: 8.5/10
分类标签: inference-engineering sglang vllm benchmark prefix-caching
质量信号:
- 关键量化结论:60% 前缀共享率 = SGLang 优势阈值;80% 共享 + 50 并发下 TTFT p50 从 310ms 降至 195ms(-37%)
- 版本明确:vLLM v0.18.0, SGLang v0.5.9, cu130
- 完整 SGLang FP8 部署命令:--mem-fraction-static 0.92
- /metrics 端点可查 RadixAttention 缓存命中率
- 量化标准:prefix overlap ratio 决策框架
保留理由:SGLang vs vLLM 决策有量化门槛,缓存命中率可观测,命令可复现
4. SitePoint · Agent Memory: The Complete Guide (2026)
URL: https://www.sitepoint.com/ai-agent-memory-guide
质量评分: 7.5/10
分类标签: agent-engineering memory python production failure-modes
质量信号:
- Python 源码:safe_agent() token 预算 + 递归摘要,broken_agent() 对比
- 5 种生产内存故障模式:Context Overflow/Silent Truncation, Stale Memory Poisoning 等
- 可靠性检查清单:memory retrieval fail 测试,端到端 session 边界测试,并发读写延迟 benchmark
- 故障验尸(postmortem)结构:每个 failure pattern 有根因 + 后果 + 对策
- 引用实现:Ollama 本地 LLM
保留理由:真实 Python 源码 + token budget 逻辑 + 生产故障模式,Agent 工程稀缺内容
5. Arize · 9 Best AI Agent Debugging Tools for Production (2026)
URL: https://arize.com/resources/best-ai-agent-debugging-tools
质量评分: 7/10
分类标签: observability agent-engineering production debugging mcp
质量信号: - 生产调试三元分类法:Tool Failure / Trajectory Failure / Outcome Mismatch - Arize AX + Signal + Swarm 自动化发现-调查-修复闭环 - Ollie agent:读源码 → 提 edit → rerun → 加入测试集 → 自动执行(需审批) - MCP 集成:coding agent 可直接查询生产 traces - Opik:持久化 issue queue + local code repair - 具体失败案例:200 status 返回的 tool call 静默失败
保留理由:生产调试工作流 + MCP + agent 自动化修复,工程落地价值高
6. Orkes/Agentspan · Why We Built Agentspan
URL: https://orkes.io/blog/why-we-built-agentspan-the-production-agent-problem-nobody-wants-to-talk-about
质量评分: 6.5/10
分类标签: agent-engineering observability workflow docker debugging
质量信号: - Conductor Python SDK 安装 + 启动命令(pip install, docker run) - server-side runtime 模型:run 在 server 端持久化,worker 独立进程 - execution graph 可视化:每个 LLM call / tool call / decision 均有记录 - 动态 workflow(self-modifying execution path)vs 预编译 pipeline 的本质区别 - 调试从猜测试验到诊断的工程转变
保留理由:生产 Agent 调试模型 + execution graph 可观测性框架,概念有工程深度
⚠️ BORDERLINE(边界,需人工判断)
7. Medium/Data Science Collective · Top LLM Observability Platforms 2026
URL: https://medium.com/data-science-collective/top-llm-observability-platforms-in-2026-2c1c37619766 评分: 5.5/10 DROP 理由: 产品横向对比为主,无原创命令或可复现步骤;Opik trace 分析、Galileo Luna 2 降本 97% 等洞察有参考价值,但不适合工程师直接复用
8. MLflow · What Is Agent Observability? A 2026 Developer Guide
URL: https://mlflow.org/articles/what-is-agent-observability-a-2026-developer-guide 评分: 5.5/10 DROP 理由: 控制面框架文章,hierarchical spans vs logs 概念有价值;MLflow SDK 集成细节偏营销,适合建立认知但无直接可操作内容
9. Deploybase · Best LLM Inference Engines 2026
URL: https://deploybase.ai/articles/best-llm-inference-engine 评分: 5/10 DROP 理由: 数字(A100 Llama 70B: vLLM 3,500 tok/s, SGLang 2,800 tok/s)与 Spheron 实测数据不一致,无测试方法说明,权威性存疑
10. Yotta Labs · What Is SGLang Architecture Performance
URL: https://www.yottalabs.ai/post/what-is-sglang-architecture-performance-and-when-to-use-it 评分: 5/10 DROP 理由: LMSYS 5x 加速数据为厂商自报,无独立验证;RadixAttention vs PagedAttention 概念有价值但无实测命令
❌ DROP
| 来源 | 标题 | 丢弃理由 |
|---|---|---|
| YouTube (Karpathy RSS) | 复现 GPT-2 / 构建 Tokenizer / LLM 入门 | 教学向,非生产工程 |
| YouTube (Fireship RSS) | 63亿开权重模型被击败 / DHH 失控 | 科技八卦,无工程技术内容 |
| YouTube | LLM Engineering Roadmap 2026 (~$99 course) | 商业课程营销,章节片段无完整工程内容 |
| YouTube | Agent Observability Ultimate Guide 2026 | 视频,无可引用代码或命令 |
📋 本次筛选统计
| 类别 | 数量 |
|---|---|
| KEEP | 6 |
| BORDERLINE | 4 |
| DROP | 4 |
| 合计 | 14 |
💡 建议写入路径
精读/主题页更新建议:
- inference-engineering 主题页 → 纳入 Spheron 两篇 + SitePoint vLLM 部署指南(benchmark 数据 + docker 命令)
- agent-engineering 主题页 → 纳入 SitePoint Agent Memory + Arize 调试工具链
- observability 主题页 → 纳入 Arize AX 工作流 + Orkes Agentspan 调试模型
建议草稿路径(可选):
- /shared/research-kb/inbox/jay/2026-10-10-inference-benchmark-vllm-sglang-h100-commands.md(来自 KEEP #1, #2, #3 合并)
- /shared/research-kb/inbox/jay/2026-10-10-agent-memory-failure-modes-python-eval-debugging.md(来自 KEEP #4, #5, #6 合并)
筛选时间: 2026-10-10 10:50 CST | 筛选工具: Tavily search + extract | 未执行 Git 写入