工程筛选报告 · Jay · 2026-07-06 下午

任务概述

  • 筛选主题:Agentic SE 系统 · LLM 工程实践 · MLOps 边界变化
  • 检索范围:arXiv (cs.SE/cs.AI) · Substack (ML Engineer, Neural Maze, Raschka) · HuggingFace Daily · GitHub Trending
  • 时间窗口:2026-06-10 ~ 2026-07-06

候选条目(共 12 条)

# 来源 标题/描述 工程相关度 备注
1 arXiv 2606.05608 The End of Software Engineering: How AI Agents Are Fundamentally Restructuring the Software Paradigm ★★★★★ 2026-06-10,Siemens/TUM,定义"Agentic Engineering",4阶段路线图
2 arXiv 2601.09822 LLM-Based Agentic Systems for Software Engineering: Challenges and Opportunities ★★★★☆ GenSE 2026 workshop,Siemens,系统文献综述,跨 SDLC 多智能体
3 arXiv 2602.10479 From Prompt-Response to Goal-Directed Systems: The Evolution of Agentic AI Software Architecture ★★★★☆ 已在 2026-07-06-1105 覆盖
4 arXiv cs.AI/new ElephantAgent / ContraFix / 2603.12813 ★★★☆☆ Agent 安全新攻击面,工具污染/记忆投毒,Contextual State Continuity
5 Substack: The ML Engineer #393 PydanticAI V2 · Raschka 本地 Agent 设置 · O'Reilly Radar 2026 ★★★★☆ 2026-06-28,具体工具链(Ollama/Qwen Code/Codex CLI),MLOps 边界扩展
6 Raschka Newsletter (Jan-May 2026) LLM Research Papers 2026 List ★★★★☆ GLM-5 · Nemotron 3 Super · MiniMax-M2 Series,Agentic Reasoning 架构
7 Substack: The Neural Maze Welcome to The AI Systems Engineer Journey ★★★☆☆ IEEE-CAI 2026 ColPali 教程,生产系统 Feature/Training/Inference Pipeline
8 Substack: Nidly ML vs AI Engineer 2026 ★★☆☆☆ 职业边界,已覆盖(Jay 2026-07-06-0935)
9 Substack: Jam with AI Data Science Roadmap 2026 · Production AI/ML Systems ★★☆☆☆ 课程推广,非工程实践
10 GitHub: louisfb01/start-ai-engineering AI Engineering Guide 2026 ★★☆☆☆ 路线图/清单,非具体工程实现
11 HuggingFace Blog State of Open Source Spring 2026 ★★☆☆☆ 生态系统统计,非具体技术细节
12 LinkedIn/人日 AI Edge 2026: Better Loops Not Bigger Models ★★☆☆☆ 摘要性分析,引用 arXiv 2603.20639

二次筛选:详细评估

✅ 保留条目(工程实践含量高)


保留 1:arXiv 2606.05608 ⭐⭐⭐⭐⭐

标题:The End of Software Engineering: How AI Agents Are Fundamentally Restructuring the Software Paradigm 链接:https://arxiv.org/html/2606.05608v1 来源:arXiv cs.SE · v2 · 2026-06-10 · Siemens AG + Technical University of Munich

保留理由: - 🔴 "Agentic Engineering"正式定义:LangChain 2026-04 首次提出,本论文为最早的系统性学术论述

"multi-agent coordination model where AI agents function as digital team members—each with defined roles, shared memory, and a unified observability layer—to drive software through the entire delivery pipeline" - 🔴 Brooks Essential Complexity 框架延伸:将 Brooks 的"设计即产品"论点扩展到 Agentic 系统,认为 AI 代理进一步放大了软件本质复杂度 - 🔴 Benchmark 证据:引用 EvoClaw(持续软件演化 benchmark),表明当前 Agent 在 commit 历史连续性维护上仍存在根本局限 - 🔴 4阶段演进路线图: - Stage I (2024-2026):辅助级单 Agent 工具 - Stage II (2025-2027):增强级多 Agent 协作 - Stage III (2026-2029):专业 Agent 团队(镜像人类工程组织) - Stage IV (2028+):自演进生态系统 - 🟡 Agent-as-a-Service 概念:将 SaaS→AaaS 演进作为商业模式拐点 - 🟢 与 SWE-bench Verified / EvoClaw 基准关联:可核验 LangChain Agentic Engineering 的实证基础

缺失/需核验: - 无具体命令/API/代码段 - EvoClaw 基准具体数值未提供(需读原文表格) - Stage III/IV 为预测性内容

后续行动:建议精读;可考虑独立主题页「Agentic SW Engineering」


保留 2:arXiv 2601.09822 ⭐⭐⭐⭐

标题:LLM-Based Agentic Systems for Software Engineering: Challenges and Opportunities 链接:https://arxiv.org/abs/2601.09822 来源:arXiv cs.SE · v2 · 2026-01-19 · Siemens AG + TUM · GenSE 2026 workshop

保留理由: - 🔴 系统文献综述:覆盖 LLM 多智能体在完整 SDLC(需求工程→架构设计→代码生成→测试→调试→部署→维护)的应用 - 🔴 框架分类:识别 3 个核心模块(LLM推理、工具使用、记忆),并引用 Guo et al. 的多智能体协作模式综述 - 🔴 实证覆盖广度:引用数百篇研究,表明 Agentic 模式跨 SE 活动的通用性 - 🟡 关键结论:fully automated SE 仍需重大突破,当前局限在"跨阶段集体决策和迭代式同伴精化"

缺失/需核验: - 为 workshop paper,深度有限 - 具体 benchmark 数值未在摘要中提供 - 工具框架列表需读原文 Section 4

后续行动:建议泛读;与 2606.05608 合并归档


保留 3:Substack The ML Engineer Issue #393 ⭐⭐⭐⭐

标题:The ML Engineer(Alejandro Saucedo)· 2026-06-28 链接:https://machinelearning.substack.com/p/issue-393-the-ml-engineer

保留理由: - 🔴 Sebastian Raschka 本地 Agent 设置: - 模型服务层:Ollama(明确提名为 inference 底座) - Agent Harness:Qwen Code + Codex CLI + Claude Code - 关键维度:inference speed、long-context behavior、tool-call reliability、permissions、telemetry、task-specific evaluation - 偏好:open-weight models + 30-35B MoE coding models(如 Qwen3.6)

"local agent workflows depend largely on inference speed, long-context behavior, tool-call reliability, permissions, telemetry, and task-specific evaluation" - 🔴 O'Reilly Radar Trends 2026 — MLOps 边界扩展: "emergence of infrastructure that allows agents to provision accounts, register domains, initiate payments, obtain credentials, and deploy applications with limited human intervention. This changes the boundary of MLOps, as teams now need to reason not only about model quality and serving latency, but also about authorization, spending controls, audit trails, sandboxing, and dependency risk." - 🔴 新兴工具清单: - SARC:Agentic 框架 guardrails wrapper(flow 级约束执行) - KAOS:K8s Agent Orchestration Service(大规模分布式 Agent 编排) - Kompute:移动端 GPU compute framework

缺失/需核验: - Raschka 设置为经验性描述,无完整配置命令 - SARC/KAOS/Kompute 需 GitHub 核实最新版本

后续行动:建议记入「Agentic Platform / Tooling」知识节点;核实 SARC/KAOS GitHub 活跃度


保留 4:Raschka LLM Research Papers 2026 List (Jan-May) ⭐⭐⭐⭐

来源:Sebastian Raschka Newsletter · https://magazine.sebastianraschka.com/p/llm-research-papers-2026-part1 形式:月度论文精选,分类整理(Architecture / Training / Efficiency / Reasoning / Agent / Coding / Eval)

保留理由(高价值条目摘录)

时间 论文/模型 关键标签 工程关联
2026-02-17 GLM-5: From Vibe Coding to Agentic Engineering GLM-5 · 智谱 首次明确使用"Agentic Engineering"术语的模型技术报告
2026-04-13 Nemotron 3 Super: MoE Hybrid Mamba-Transformer for Agentic Reasoning NVIDIA · MoE + SSM + Transformer 混合架构专为 Agentic 推理场景优化,有部署参考价值
2026-05-25 MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence MiniMax Efficiency-oriented architecture,与当前 vLLM/SGLang 推理栈兼容
2026-03-16 Mamba-3: Improved Sequence Modeling Using State Space Principles SSM 替代 Transformer 的推理效率路径
2026-03-15 Attention Residuals Architecture 与 FlashAttention 对比的推理优化思路

缺失/需核验: - 为论文列表,非原文 - 需直接读论文看具体数字

后续行动:建议精读 GLM-5 和 Nemotron 3 Super;更新 LLM Architecture 主题页


保留 5:Substack The Neural Maze — ColPali Document Intelligence ⭐⭐⭐

来源:https://theneuralmaze.substack.com/p/welcome-to-the-ai-systems-engineer 质量标注:IEEE-CAI 2026 工业教程(Antonio Zarauz Moreno),非推测性内容

保留理由: - 🔴 ColPali 多模态文档智能生产系统架构(Feature/Training/Inference Pipeline): - Feature Pipeline:每页渲染为图像 → VLM 生成 patch embedding 网格 → 存储 - Training Pipeline:可选微调 ColPali retriever + VLM(synthetic data) - Inference Pipeline:query → multi-vector retrieval → top-K pages → multimodal generator → answer with bounding boxes / page citations

"This isn't speculative content... this is the exact system... that has been accepted as an IEEE-CAI 2026 tutorial" - 🟢 与 naive RAG 对比的工程教训: "The naive RAG pipeline is the 'hello world' of LLM apps. The real RAG system is the one that survives contact with users." - 生产 RAG 关键组件:chunking strategy、retrieval quality、hybrid search vs pure semantic、query rewriting、reranking、hallucination detection、citation enforcement、evaluation harnesses、latency、cost、index freshness - 🟡 Agentic AI 系统 Chassis 框架:Feature/Training/Inference 三层统一抽象

缺失/需核验: - 具体 VLM 模型/ColPali 版本未提供 - 需要 IEEE-CAI 2026 原文核实 Tutorial 代码

后续行动:建议归档「Advanced RAG / Document Intelligence」;核实 IEEE-CAI 2026 Tutorial 仓库


❌ 丢弃条目

# 来源 丢弃理由
9 Jam with AI (Data Science Roadmap 2026, Production AI/ML Systems) 课程推广材料,无具体工程实践内容
10 GitHub: louisfb01/start-ai-engineering 路线图/清单类内容,缺具体命令/代码/错误处理
11 HuggingFace Blog: State of Open Source Spring 2026 生态统计报告,非技术实现细节
12 LinkedIn: AI Edge 2026 Better Loops 二手摘要,引用来源可追溯至已收录论文
8 Nidly: ML vs AI Engineer 职业边界分析,缺工程实践含量;上午简报已覆盖

分类标签

agentic-systems software-engineering mlops-boundary ollama qwen-codex colpali document-intelligence benchmark-evoclaw swe-bench moe-hybrid state-space-models rag-production agent-security tool-poisoning pydantic-ai-v2


建议写入路径

/shared/research-kb/inbox/jay/2026-07-06-1500-engineering-filter-agentic-se-llmops-2026.md

后续行动建议

  1. 精读:「arXiv 2606.05608」全文(EvoClaw 数值、Stage III/IV 基准细节)
  2. 核实:SARC、KAOS、Kompute GitHub 仓库活跃度与 Stars
  3. 主题页:「Agentic SW Engineering」建议新增;或并入现有「Agent Systems」页
  4. 核验:IEEE-CAI 2026 ColPali Tutorial 代码仓库(需外部访问)
  5. 精读:Raschka Newsletter 中 GLM-5(Agentic Engineering 首篇)和 Nemotron 3 Super(MoE+SSM 混合推理优化)

写入确认

  • 实际写入路径/shared/research-kb/inbox/jay/2026-07-06-1500-engineering-filter-agentic-se-llmops-2026.md
  • 写入状态:✅ 已完成
  • 本轮无写入时说明:不适用,本轮已写入文件