研究表明矛盾解析本质上是写入时并发控制,并将缺失的契约——一个在隔离性、模式与来源维度上被证明正确的写入时正确性规范——显式化,固定了每个生产启发式都默认假设、却没有任何已部署系统显式给出的保证。It is shown that contradiction resolution is write-time concurrency control and make the missing contract explicit, a write-time correctness specification, proved sound across isolation, schema, and provenance, pinning the guarantee every production heuristic assumes but no deployed system makes explicit.
论文
404 张论文卡片 · Agent 智能体 · OA 绿色
本文提出了 LIFE-RL 框架,在不修改内部控制流的前提下,将原生编码 Agent harness 与可扩展的策略梯度优化相连接,并通过 GSPO 在三个原生编码 Agent harness 上训练稀疏 MoE 模型 Qwen3.5-35B-A3B 对其进行了评估。LIFE-RL, a framework that bridges native coding-agent harnesses with scalable policy-gradient optimization without modifying their internal control flow, is presented and evaluated by training the sparse MoE model Qwen3.5-35B-A3B with GSPO across three native coding-agent harnesses.
本文提出 Co-RL,一种由多个解耦模型组成的框架,这些模型不共享参数,通过基于彼此输出奖励的强化学习同时进行优化,并表明无监督推理可以通过协作式多智能体训练涌现This work introduces Co-RL, a framework in which multiple decoupled models, sharing no parameters, are simultaneously optimized through RL using rewards derived from their peers, and shows that unsupervised reasoning can emerge through cooperative multi-agent training.
论文认为,agentic CTI 调查的瓶颈在于该 substrate 而非模型能力,并提出了面向 Agent 的语料库脚手架 CTIFoundry。It is argued that this substrate, not model capability, is the bottleneck on agentic CTI investigation, and CTIFoundry, an agent-native corpus scaffold, is presented.
SkillGate 将 9B 策略的成功率从 40.8% 提升至 53.2%,显著优于将相同预算仅用于 outcome reward 的方案,同时将误导性候选的暴露减少三分之二,并读取更少的 skill。SkillGate lifts a 9B policy from 40.8% to 53.2% trial success, well ahead of the identical budget spent on outcome reward alone, while cutting exposure to misleading candidates by two thirds and reading fewer skills.
受损模型评估测试在第一个合法 tool call 之后插入 ground-truth 攻击调用,从而独立于模型行为测试 APC,证明了 APC 实现的 Blast Radius 单调性与组合可靠性。The compromised-model evaluation tests APC independently of model behavior by inserting the ground-truth attack call after the first legitimate tool call, which proves Blast Radius Monotonicity and Composition Soundness for APC implementations and proves Blast Radius Monotonicity and Composition Soundness for APC implementations.
提出 Repo0,一个面向零到全代码生成的持续结构演化框架,维护显式的架构状态,实例化为双有向无环图(Dual-DAG),由需求级 DAG、组件级 DAG 及其对齐关系组成。Repo0 is presented, a continuous structural evolution framework for zero-to-all code generation that maintains an explicit architectural state instantiated as a Dual-Directed-Acyclic-Graph (Dual-DAG), consisting of a requirement-level DAG, a component-level DAG, and their alignment relation.
一项配对消融实验在保留仓库与可执行工程上下文的同时移除显式科学指导,表明科学知识并非一律有益:可靠信息能约束修复、提升平均表现与 token 效率,而错位指导则会诱发锚定,不必然提升精确修复成功率。A paired ablation that removes explicit scientific guidance while preserving the repository and executable engineering context shows that scientific knowledge is not uniformly beneficial: well-grounded information can constrain repair and improve average performance and token efficiency, whereas poorly aligned guidance can induce anchoring and does not necessarily improve exact repair success.
本文指出,持续技能演化的关键瓶颈既非编辑能力亦非迭代轮数,而在于评估反馈能否持续提供可信的演化梯度;据此提出 SkillEvo,由可信反馈生成梯度,由可控治理约束方向。This work argues that the binding constraint on sustained skill evolution is neither editing capability nor the number of iterations, but whether the evaluation feedback keeps supplying trustworthy evolution gradients, and introduces SkillEvo, in which trustworthy feedback generates the gradient and controllable governance constrains its direction.
泄露带来两类现实攻击:一个训练好的分类器可从常规自然语言输出中推断用户记忆的语义谓词;一个由 RL 训练的对抗者可从生产级风格的 Agent 中完整提取社会安全号码。Leakage enables two practical attacks: a trained classifier that infers semantic predicates about user memories from routine natural-language outputs, and an RL-trained adversary that extracts full Social Security Numbers from a production-style agent.
Bespoke-Card 在传统通用估计器与学习型估计器架构之外开辟了一条新的基数估计路径,它是一个 Agent 驱动的系统,将面向特定工作负载的基数估计器合成为可执行代码。Bespoke-Card is opening a new avenue for cardinality estimation next to classical generic estimators and learned estimator architectures, an agent-driven system that synthesizes workload-specific cardinality estimators as executable code.
FlowEvo 是一个免训练框架,在推理时让工作流与技能协同进化:它将成功的工作流编译为可调用技能,存入持久化技能库,并通过直接执行或作为上下文来检索使用这些技能,以构建新的工作流。FlowEvo, a training-free framework in which workflows and skills co-evolve at inference time, compiles successful workflows into callable skills, stores them in a persistent bank, and uses retrieved skills either through direct execution or as context for constructing new workflows.
对指令下发 Agent 的评估应报告模型配置、生成契约、执行路径、工作点及终态校验器,而不应将匹配分数视为模型的内在属性。Evaluations of command-issuing agents should report the model configuration, generation contract, execution path, operating point, and final-state validator rather than treat a matched score as an intrinsic model property.
本文提出 AID-Guard,一种有状态的授权到效果(authorization-to-effect)闭包协议:在提交(commit)时重新校验已批准的请求与提供方状态;在歧义下仅保留一个预订;仅在收到终态结果或经认证的“无效果”并设置投递栅栏(delivery fence)后,才允许释放或生成一个后继。This work presents AID-Guard, a stateful authorization-to-effect closure protocol that revalidates the approved request and provider state at commit, retains one reservation under ambiguity, and permits release or one successor only after a terminal result or certified no effect with a delivery fence.
结果显示,规约规模本身并不能预测实现质量,跨 Agent 迁移可能导致显著的、依赖具体 Agent 的性能下降;因此在异构 SDD 工作流中,不应将规约默认视为与 Agent 无关的工件(artifact)。The results show that specification size alone does not predict implementation quality and that cross-agent transfer can produce substantial agent-dependent degradation, and suggest that specifications in heterogeneous SDD workflows should not automatically be treated as agent-neutral artifacts.
PhysCaP 在 code-as-policy 框架上增加了物理信息驱动的探索层,使其能够通过交互进行显式的信息获取,并引入免训练的物理属性提取模块,仅凭机器人本体感知即可估计物体质量与刚度,无需额外传感器。PhysCaP augments code-as-policy frameworks with a physics-informed exploration layer that enables explicit information-seeking through interaction, and introduces training-free physical property extraction modules that estimate object mass and stiffness from robot proprioception without additional sensors.
在所测试的同伴排序信息流下,鲁棒的结论是词法层面的趋同,而非对一般意见的捕获,也不是合成 LLM-Agent 群体中的一般性协调优势;该结果未估计其对人类或真实生产平台的影响。The robust result is lexical convergence under the tested peer-ranked feed, not general opinion capture or a general coordination advantage in the synthetic LLM-agent populations; it does not estimate effects on people or production platforms.
本文提出 GameXpert-Bench,将游戏开发的三个生命周期阶段以 coding agent 操作化为三条互补的基准轨道,并发现现有 agent 在生成可玩基础框架和实现显式需求方面更可靠,而在发现缺陷、验证运行时行为以及跨变更保持功能一致性方面能力较弱。GameXpert-Bench is introduced, which operationalizes the three lifecycle stages of game development with a coding agent as three complementary benchmark tracks, and finds current agents are more reliable at producing playable foundations and implementing explicit requirements than at discovering defects, verifying runtime behavior, and preserving functionality across changes.
本文提出 ARC(Advantage Regularization via Conditioning),一种通过策略条件化 rollout 分组来恢复更公平的相对比较、并结合混合奖励与熵正则化的训练方法。The proposed ARC (Advantage Regularization via Conditioning), a training recipe that restores fairer relative comparison through strategy-conditioned rollout grouping, together with hybrid rewards and entropy regularization, is proposed.
本文提出 Knowledge Triage 框架,对 agent 知识库的每一行按类型分类,并为每种类型配置独立的保留策略,同时开源发布 AgentArtifactCorpus 数据集、分类器及参考实现。Knowledge Triage, a framework that classifies each line of an agent's knowledge base by type and routes each type through its own retention policy, is addressed, and AgentArtifactCorpus, the classifier, and the reference implementation are released.
介绍 AutoResearch,一个连接 Idea Generation 与 Idea Execution 的两阶段系统,分别解决研究思路如何形成与如何通过实验可靠验证的问题,展示「实验前先夯实洞见、接受前先夯实结论」的研究流程。AutoResearch is introduced, a two-stage system that connects Idea Generation with Idea Execution to address both how research ideas are formed and how they are reliably established through experimentation to demonstrate a research process in which meaningful insight is grounded before experimentation and conclusions are grounded before acceptance.
AtlasNav 减少了 Evidence Blindness,更早实现完整证据,在 PhantomWiki 上对语料结构与规模变化保持鲁棒,并在异构企业数据上取得领先性能。AtlasNav reduces Evidence Blindness, reduces Evidence Blindness, realizes complete evidence earlier, remains robust to corpus-structure and scale shifts on PhantomWiki, and achieves leading performance on heterogeneous enterprise data.
介绍 CyberFactory,一个统一的开源框架,贯通 PoC 生成、漏洞修补与 CyberQA 三大任务中的数据构建、轨迹合成与模型训练。CyberFactory is introduced, a unified open-source framework that connects data construction, trajectory synthesis, and model training across proof-of-concept (PoC) generation, vulnerability patching, and cybersecurity question answering (CyberQA).
介绍 DREAM,一种自主优化控制架构:在不替换现有流水线的前提下叠加感知可感知、可编排、可审计的策略层,支持将 agentic meta-control 作为工业推荐的一种可行范式。This work presents DREAM (Developing Recommender Engine with Agentic Methods), an autonomous optimization control architecture that adds a perception-aware, orchestrable, and auditable policy layer atop existing pipelines without replacing them, supporting agentic meta-control as a viable paradigm for industrial recommendation.
本文指出 LLM agent 中信息抽取与动作之间的 state-transmission 失效,并展示 handoff 变换如何在保留状态内容的同时削弱其对下游动作的约束。This work identifies a state-transmission failure between information extraction and action in large language model agents, and shows how handoff transformations can retain state content while weakening its constraints on downstream action.
AgentRoom 是一种面向并发编码 Agent 的实时协同编辑协议,通过在 CRDT 合并的共享文件系统上将文件级 claim、status 和 broadcast 暴露为 MCP 工具,且运行间的差异小于 CLI-stable 模型。AgentRoom is a realtime collaborative editing protocol for concurrent coding agents that exposes file-level claim, status, and broadcast as MCP tools on a CRDT-merged shared filesystem and has less run-to-run variation than CLI-stable models.
行为拓扑更多由部署 harness 决定而非 LLM 本身,为安全审计和运行时监控提供一种与模型无关的结构化 primitive,并同时满足两类预测目标。Behavioral topology is shaped more by the deployment harness than by the LLM, providing a model-agnostic structural primitive for safety auditing and runtime monitoring, and addresses both prediction goals.
The Station 被评估为一个开放世界多 agent 环境,不同模型家族的 AI agent 在其中无需中心协调器或脚本化流程即可共同追求同一研究目标,并提供发现产生过程的透明记录。The Station is evaluated, an open-world multi-agent environment in which AI agents from different model families pursue a shared research goal without a central coordinator or scripted pipeline, providing a transparent record of how discoveries emerged.
本文提出 Secure On-Policy Distillation (SecOPD),提供 token 级反馈以指导防御性微调,并能泛化到训练中完全未见过的领域。This paper proposes Secure On-Policy Distillation (SecOPD) that provides token-level feedback to guide defensive fine-tuning, and generalizes to domains completely unseen in training.
本文变化 handoff 方向、时机与接口,对比保留仓库状态下的全轨迹传输、压缩与轨迹移除,发现偏好接口随方向反转:减少 LC-model 轨迹信息可提升 escalation 质量,而移除 HC-model 轨迹则会降低 downshift 质量。This work varies handoff direction, timing, and interface, comparing full-trajectory transfer, compaction, and trajectory removal while preserving the repository state, and finds that the preferred interface also reverses with direction: reducing LC-model trajectory information improves escalation quality, whereas removing the HC-model trajectory reduces downshift quality.
本文发布 SWE Refactor Bench,一个包含 20 项全仓库迁移的基准,涵盖 4 类技术债务,作为开发面向可靠全仓库迁移的编码 Agent 的严格测试平台。SWE Refactor Bench is introduced, a benchmark comprising 20 whole-repository migrations, covering 4 kinds of technical debt, and SWE Refactor Bench is positioned as a rigorous testbed for developing coding agents for reliable whole-repository migrations.
本文提出一个模块化的医学影像 Agent,用于轴位 CT 切片中的二元空间关系验证,采用显式模块化空间验证阶段,并表明显式模块化空间验证可作为未来面向报告的医学影像 Agent 的有前景的构建模块。This work presents a modular medical imaging agent for binary spatial relation verification in axial CT slices using explicit modular spatial verification stages, and suggests that explicit modular spatial verification can serve as a promising building block for future report-oriented medical imaging agents.
本工作提出 QBugLM,一个多 Agent 框架,可自动化量子软件调试流水线,覆盖基于分类法的缺陷注入、基于 LLM 的检测与修复,直至基于仿真的验证,框架无关地支持 OpenQASM 3.0 程序。This work proposes QBugLM, a multi-agent framework that automates the quantum software debugging pipeline, from taxonomy-driven bug injection to LLM-based detection and repair, and finally to simulation-based validation, for framework-agnostic OpenQASM 3.0 programs.
本文提出 PILOT,一种通过两种耦合机制实现实时自我改进的 supervisor-worker 框架:(1) live steering 允许独立的 supervisor 在执行期间重定向或中止当前 worker;(2) live self-evolution 将执行中发现的过程与失败模式提炼为可复用的 skills 与记忆。PILOT is presented, a supervisor-worker harness for live self-improvement through two coupled mechanisms: (1) live steering lets a separate supervisor redirect or abort the active worker during execution; and (2) live self-evolution distils procedures and failure modes revealed during execution into reusable skills and memory.
本文探索"以代码表示 3D 形状"的范式,利用并放大 LLM 的编码能力进行 3D 建模,并提出 Procedura 框架——通过编写一个由命名零件构成、并通过类型化、可机器校验的连接关系装配而成的参数化程序来建模对象。The paradigm of 3D shape as code is explored, leveraging and scaling the coding ability of an LLM for 3D modeling, and Procedura, a novel 3D modeling agent framework that writes an object as a procedural assembly, a parametric program whose named parts are joined by typed, machine-checkable mates.
本文提出 Reinforcement Learning with Human-Engine Verification(RLHEV),一种结合稠密引擎信号与开发过程中隐式人类接受反馈的后训练范式,用于支持强化学习后训练。This work proposes Reinforcement Learning with Human-Engine Verification (RLHEV), a post-training paradigm that combines dense engine signals with implicit human acceptance feedback from the development process to support RL post-training.