实验表明,OpenComputer 的硬编码验证器比 LLM-as-judge 评估更贴合人类裁定,尤其当任务成败取决于细粒度应用状态时。Experiments show that OpenComputer's hard-coded verifiers align more closely with human adjudication than LLM-as-judge evaluation, especially when success depends on fine-grained application state.
论文
50 张论文卡片 · Agent 智能体 · 应用落地 · OA 绿色
本文提出三种协议级原语以填补Model Context Protocol的空白:身份传递、自适应工具预算与结构化错误语义,并提出Structured Error Recovery Framework (SERF),提供机器可读的失败语义以支持确定性的Agent自校正。Three protocol-level primitives are proposed to fill gaps in the Model Context Protocol: identity propagation, adaptive tool budgeting, and structured error semantics, and the Structured Error Recovery Framework (SERF), which provides machine-readable failure semantics that enable deterministic agent self-correction.
Agent Lightning v1.0 是一个轻量级的可控 Agentic RL 框架,约 3500 行代码实现,支持任意 Agent harness,并作为研究 retokenization、样本合并、优势计算、损失归一化与后端调度等挑战的实用测试平台。Agent Lightning v1.0 is presented, a lightweight framework for harnessed agentic RL implemented in approximately 3,500 lines of code that supports arbitrary agent harnesses and serves as a practical testbed for studying challenges in retokenization, sample merging, advantage calculation, loss normalization, and backend scheduling.
研究表明矛盾解析本质上是写入时并发控制,并将缺失的契约——一个在隔离性、模式与来源维度上被证明正确的写入时正确性规范——显式化,固定了每个生产启发式都默认假设、却没有任何已部署系统显式给出的保证。It is shown that contradiction resolution is write-time concurrency control and make the missing contract explicit, a write-time correctness specification, proved sound across isolation, schema, and provenance, pinning the guarantee every production heuristic assumes but no deployed system makes explicit.
行为拓扑更多由部署 harness 决定而非 LLM 本身,为安全审计和运行时监控提供一种与模型无关的结构化 primitive,并同时满足两类预测目标。Behavioral topology is shaped more by the deployment harness than by the LLM, providing a model-agnostic structural primitive for safety auditing and runtime monitoring, and addresses both prediction goals.
CrabOS 将人机交替主导复杂任务的支持从依赖桥接的应用层方案提升为原生操作系统能力,为开发与运行 AI agent 提供了新基础。CrabOS elevates support for complex tasks with alternating human and AI leadership from bridge-dependent application-level solutions to native operating-system capabilities, which provide a new foundation for developing and running AI agents.
该工作提出 Super Library Agent 问题:Agent 顺序生成 N 个相关应用组成的组合,同时维护一个共享的 Super Library 用于跨应用可复用组件,并解决了基于候选引导的代码块摘要抽取、抽取前的代码库整合,以及利用抽取 trace 和调用图信息的上下文感知迁移。This work introduces the Super Library Agent problem, where an agent sequentially generates a portfolio of N related applications while maintaining a shared Super Library of reusable cross-application components, and addresses candidate-guided extraction over code chunk summaries, pre-extraction codebase consolidation, and context-aware migration using extraction traces and call-graph information.
为解决 LLM 与非语言 Agent 的协作问题,本文提出 latent state internalization,将子 Agent 的连续表示直接投射到 LLM 的 token 流中作为习得的状态 token,并随着动作推进环境状态而进行动态重编码。To solve LLM collaboration with non-language agents, latent state internalization is introduced, which projects the subagent's continuous representations directly into the LLM's token stream as learned state tokens, with dynamic re-encoding as actions advance the environment state.
本文介绍 SCHEMEARENA,一个面向可扩展谋划行为压力测试的 400 场景基准,通过覆盖多种安全相关工具领域、工具性目标、监管条件与压力机制的因子化场景合成框架构建。This work introduces SCHEMEARENA, a 400-scenario benchmark for scalable scheming stress testing, constructed through a factorized scenario synthesis framework spanning diverse safety-relevant tool domains, instrumental goals, oversight conditions, and pressure mechanisms.
本文规定 EBL-Core,一个执行边界合规配置文件,用于判定一个规范且完全实例化的 AI 生成候选对象在明确条件下是否可获得操作范围的执行权限。EBL-Core, an execution-boundary conformance profile for deciding whether one canonical, fully materialized AI-generated candidate may receive action-scoped execution authority under explicit conditions, is specified.
本文探讨 multi-agent system,并指出当前尚未被充分解决的问题,同时探索了 multi-agent system 在区块链系统中的潜在应用,为其在真实分布式系统中的未来发展与落地提供启示。This paper explores multi-agent systems and identifies challenges that remain inadequately addressed, and explores potential applications of multi-agent systems in blockchain systems to shed light on their future development and application in real-world distributed systems.
提出 Continual Search,一个迭代框架:在多轮对话中持续推动判别器搜索尚未解决的诊断证据,在多个基准测试套件和模型系列上一致提升归因性能Continual Search is introduced, an iterative framework that nudges the judge, over successive turns, to keep searching for unresolved diagnostic evidence, which consistently improves attribution performance across multiple benchmark suites and model families.
结果表明,模型级对齐并不具备可组合性:单独能力强且看似安全的 agent 在组成系统后,可能随着 AI 的持续性与互联化而产生性质上不同的失效模式。The results suggest that model-level alignment is not compositional: individually capable and apparently safe agents can form systems with qualitatively different failure modes as AI becomes persistent and interconnected.
本文提出 XConf (eXperiential Confidence):与模型累积经验一起估计置信度,并将经验式置信度估计视为未来通用置信度估计的新范式。This work proposes XConf (eXperiential Confidence): estimating confidence together with the model's accumulated experience, and sees experiential confidence estimation as a new paradigm for future general-purpose confidence estimation.
本文提出 GPT-Policy,一个用于上下文机器人学习的通用 Agent 框架,集成了一个保留任务相关视觉过渡的 context compiler、一个提出机器人-工具动作的 VLM,以及一个验证并执行每个动作并报告结果的 constrained controller。GPT-Policy is introduced, a general-agent framework for in-context robot learning that integrates a context compiler that preserves task-relevant visual transitions, a VLM that proposes robot-tool actions, and a constrained controller that verifies and executes each action and reports its outcome.
该工作提出 EvoSkill-GUI,一个免训练框架,其中每个 skill 都是一个结构化的多文件包,包含检索元数据、可执行计划、备份定位、故障恢复规则、可访问性工具以及失败案例。This work proposes EvoSkill-GUI, a training-free framework in which each skill is a structured multi-file package containing retrieval metadata, executable plans, backup localization, failure-recovery rules, accessibility utilities, and failure cases.
本文提出 ActObs,对每条轨迹中已有的观测 token 也进行监督,并将这一差异归因于 SFT:动作与观测梯度迅速趋于正交,而仅训练动作会留下较大的残留观测梯度,并将环境预测能力拉低至基座模型之下。This work introduces ActObs, which also supervises the observation tokens already present in each trajectory, and traces this difference to SFT: action and observation gradients rapidly become orthogonal, while action-only training leaves a large residual observation gradient and degrades environment prediction below the base model.
总体而言,研究结果表明,长期交互会以产生安全风险的方式重塑 Agent 的协作模式;限制 Agent 可获取的交互历史数量与范围能够减少串通行为。Overall, the findings show that long-horizon interaction can reshape how agents coordinate in ways that create safety risks, and restricting the amount and scope of interaction history available to agents reduces collusion.
本文论证代理需要一个操作系统级基底,为身份、输入中介、内存治理与执行控制提供强制且不可绕过的服务,并提出 AgentKernel,一个以安全为一流设计约束为前提、面向信任的代理操作系统。This work argues that agents need an operating-system substrate providing mandatory, non-bypassable services for identity, input mediation, memory governance, and execution control, and introduces AgentKernel, a trust-native agent operating system built around the premise that security must be a first-class design constraint.
本文对从 GitHub 收集的 Jev 项目进行了大规模、数据驱动的分析,发现其公共生态在早期增长迅速,新项目不断涌现并被集成到既有仓库中;研究提示 Jev 充当一种可复用的决策组件,其功能随所嵌入的工作流而变化。A large-scale, data-driven analysis of Jev projects collected from GitHub finds rapid early growth in Jev's public ecosystem, with both new projects and integration into existing repositories, and suggests that Jev serves as a reusable decision component whose functionality varies with the surrounding workflow.
EngramRAG 是一种自适应记忆架构,将低延迟的 Waking State 反射与异步后台 Dreaming State 整合周期耦合,并引入融合密集向量、BM25 与 U-PPR 的三源混合检索,通过动态 Reciprocal Rank Fusion (RRF) 实现。The proposed EngramRAG is an adaptive memory architecture coupling a low-latency Waking State reflex with an asynchronous background Dreaming State consolidation cycle, and introduces triple-source hybrid retrieval fusing dense vectors, BM25, and U-PPR via dynamic Reciprocal Rank Fusion (RRF).
提出反事实树上的注意力搜索(ASCT),将训练时的多步搜索转化为局部动作信用,将反事实评估与策略学习相连接,同时部署时只需使用 actor。Attentive Search over Counterfactual Trees (ASCT) is introduced, a framework that turns training-time multi-step search into local action credit and connects counterfactual evaluation to policy learning while deploying the actor alone.
本文提出 EVOKE,一种后训练方法,通过在固定状态下以目标多样性对直接决策施加监督,来施加压力以激发模型内化的、可迁移动作的世界知识,并提供了一种通过直接决策监督激发内化世界知识以获得可迁移动作的新视角。EVOKE is introduced, a post-training method that supplies pressure on eliciting internalized world knowledge for transferable action through direct decision supervision through goal diversity at fixed states, and offers a new perspective on eliciting internalized world knowledge for transferable action through direct decision supervision.
将 Jev 面向决策的 API 集成到边缘服务编排中,在保持服务完成度的同时降低开销,并支持在有界契约下针对延迟受限的准入进行决策模型替换。Jev's decision-oriented application programming interface (API) is integrated into edge service orchestration to reduce overhead while retaining service completion, and decision-model substitution for latency-bound admission on bounded contracts is supported.
本文将多智能体系统(MAS)协调建模为数据管理问题,提出用智能体的状态足迹来刻画它们:即在其自身局部上下文与状态、以及编排器和外部系统状态上的读写行为。This work frames MAS coordination as a data management problem and proposes to describe agents by their state footprint: the state they read and write across their own local context and state, as well as the state of the orchestrator and external systems.
结果表明现有科学软件能够为终端 Agent 提供可扩展且经过行为验证的监督,并提出 software-in-the-loop reconstruction,一种自监督框架,从现有软件工作流(即把结构化输入映射为输出的可执行程序)中获取参考输出与验证目标。Results indicate that existing scientific software can provide scalable, behaviorally verified supervision for terminal agents, and introduces software-in-the-loop reconstruction, a self-supervised framework that obtains reference outputs and verification targets from existing software workflows, executable programs mapping structured inputs to outputs.
在资源受限环境中,专业癫痫专家稀缺,使基于 LLM 的决策支持对管理纵向治疗的一线临床医生具有吸引力。此类系统必须适应当地处方实践并知道何时转诊。我们在乌干达儿科癫痫诊疗中研究该问题,基于纵向非结构化门诊记录预测抗癫痫用药方案。标准提示与医生处方取得了一定程度的一致性,但神经科医生审查显示许多错误反映的是分布失校的处方默认值而非失败。Specialist epilepsy expertise is scarce in resource-constrained settings, making LLM-based decision support attractive for frontline clinicians managing longitudinal treatment. Such systems must adapt to local prescribing practice and know when to defer. We study this problem in Ugandan pediatric epilepsy care, predicting anti-seizure medication regimens from longitudinal unstructured clinic notes. Standard prompting achieves non-trivial agreement with physician prescriptions, but neurologist review shows that many errors reflect distribution-miscalibrated prescribing defaults rather than fail
实验结果表明,所得参数集可生成可区分的个性化换道行为,同时 RAG 始终提升偏好理解效果,尤其对隐式指令效果显著,表明将基于 LLM 的自然语言交互与 Apollo 集成以支持个性化换道行为生成具有潜力。Experimental results show that the derived parameter sets generate distinguishable personalized lane-change behaviors, while RAG consistently improves preference interpretation, particularly for implicit commands, indicating the potential of integrating LLM-based natural-language interaction with Apollo to support personalized lane-change behavior generation.
本文提出 DuoMem,一种双空间蒸馏框架,可将程序化问题求解能力从大型教师模型迁移到紧凑学生模型,并适用于实时边缘部署,而这一点对教师模型而言颇具挑战。DuoMem is introduced, a dual-space distillation framework that transfers procedural problem-solving ability from a large teacher model to compact student models and is viable for real-time edge deployment, which would be challenging for the teacher.
提出 Transparent Two-Pass Execution,一种在推理时将工具执行与 schema 约束响应生成解耦的策略;实验结果表明该方法无需模型重新训练即可恢复工具调用能力,同时保持结构化输出保证。Transparent Two-Pass Execution is proposed, an inference-time strategy that decouples tool execution from schema-constrained response generation and experimental results show that this approach restores tool invocation while preserving structured output guarantees without requiring model retraining.
AOHP 的核心设计原则是将 Agent 视为 OS 中的一等公民,从而支持自适应用户界面以及对 Agent 友好的运行时环境;在任务完成度、执行成本和安全策略合规性方面均展现出明显优势。The core design principle of AOHP is to treat agents as first-class OS actors, enabling adaptive user interfaces and agent-friendly runtime environments, and shows clear advantages in task completion, execution cost, and security-policy compliance.
本研究收集了 9,041 个使用流行 AI Agent 开发的开源应用,审计了 200 个公开部署的应用,发现 1,186 个漏洞,并对 vibe-coded 应用的安全态势提供了实证分析。This study collects 9,041 open-source applications developed using popular AI agents, audits 200 publicly deployed applications, uncovering 1,186 vulnerabilities, and provides an empirical understanding of the security landscape of vibe-coded applications.
提出一种多 Agent 框架,通过以确定性编排约束替代 "LLM-as-a-judge" 路由,解决可能在到达患者前未被发现的过早诊断交接与静默临床幻觉问题;观察到 OLDCARTS 完整度与语义熵之间存在统计显著的负相关,提示结构化信息采集与诊断不确定性降低相关。A multi-agent framework that addresses premature diagnostic handoff and silent clinical hallucinations that may go undetected before reaching the patient by replacing ``LLM-as-a-judge''routing with deterministic orchestration constraints is proposed and observes a statistically significant negative correlation between OLDCARTS completeness and semantic entropy, suggesting that structured information gathering is associated with reduced diagnostic uncertainty.
本文引入 trace-economic underwriting,将工具调用 trace 映射为客户风险敞口与可索赔损失,并以此表示用于定价、控制与风险转移,使用确定性经济标签而非 LLM 评判器。T trace-economic underwriting is introduced, which maps tool-use traces to customer exposure and claimable loss, then uses this representation for pricing, control, and risk transfer, and uses deterministic economic labels rather than an LLM judge.
PalmClaw 是一个开源 Agent 框架,原生运行于手机端,直接在设备上管理 session、memory、Skill、工具以及 agent loop,使 Agent 能够直接调用移动端能力,同时保证每一步操作的显式与可控。PalmClaw is an open-source agent framework that runs natively on mobile phones and manages the sessions, memory, skills, tools, and agent loop directly on the device, allowing agents to use mobile capabilities directly while keeping each action explicit and controlled.
SPEAR 是一个 Python 库,可通过模块化插件架构连接任意 Unreal Engine 应用并对其进行编程化控制;同时引入一种表达力强的高层编程模型,使用户能够以任意数据依赖关系指定复杂的 UE 工作图,并在单个 UE 帧内确定性执行这些图。SPEAR is a Python library that can connect to, and programmatically control, any Unreal Engine application via a modular plugin architecture, and introduces an expressive high-level programming model that enables users to specify complex graphs of UE work with arbitrary data dependencies among work items, and to execute these graphs deterministically within a single UE frame.