本文提出 OSWorld-Science,这是一个结合科学意义任务、基于 artifact 的评估以及高效 agent harness 的基准与评估环境,用于研究科学领域的计算机使用,从而在科学工作流中系统评估 agent 能力与 harness 设计。OSWorld-Science is introduced, a benchmark and evaluation environment that combines scientifically meaningful tasks, artifact-based evaluation, and an efficient agent harness for studying computer use in the scientific domain, enabling systematic evaluation of both agent capabilities and harness design in scientific workflows.
论文
1640 张论文卡片 · OA 绿色
本文提出 Loop Scaling Laws,是首个将 recurrence 和 sparsity 与模型规模、数据联合建模的 scaling law,为在算力和显存约束下设计 looped MoE 模型提供了原则性基础。Loop Scaling Laws are introduced, the first scaling law to jointly model recurrence and sparsity alongside model size and data, and provide a principled foundation for designing looped MoE models under compute and memory constraints.
SMART,面向长篇字幕翻译的自演化多 Agent 系统,构建持久化的剧集级 memory,并通过动态 router 和 Mixture-of-Agents 层翻译部分句子,配合术语验证、字幕约束校验和上下文检索工具。SMART, a Self-evolving Multi-Agent system for long-foRm subtitle Translation, builds persistent series-level memory and translates a subset of sentences through a dynamic router and Mixture-of-Agents layer with tools for terminology verification, subtitle constraint validation, and contextual retrieval.
本文提出 EVOKE,一种后训练方法,通过在固定状态下以目标多样性对直接决策施加监督,来施加压力以激发模型内化的、可迁移动作的世界知识,并提供了一种通过直接决策监督激发内化世界知识以获得可迁移动作的新视角。EVOKE is introduced, a post-training method that supplies pressure on eliciting internalized world knowledge for transferable action through direct decision supervision through goal diversity at fixed states, and offers a new perspective on eliciting internalized world knowledge for transferable action through direct decision supervision.
转移层面的结果表明,state adaptation 选择性而非统一地应用时最为有效,且 adaptation 的价值取决于策略反转的频率与幅度。The transition-level results suggest that state adaptation is most useful when applied selectively rather than uniformly, and that adaptation value depends on both the frequency and magnitude of strategy reversals.
RAGScope 是一个泄漏受控的评估协议,用于评估仅使用任务输入、检索上下文和答案文本的本地证据门,结合了上下文分组划分、折范围预处理、组自举区间、部署工作点、端到端运行时以及显式的源偏移压力测试。RAGScope, a leakage-controlled protocol for evaluating local evidence gates that use only the task input, retrieved context, and answer text is presented, which combines context-grouped splits, fold-scoped preprocessing, group bootstrap intervals, deployment operating points, end-to-end runtime, and explicit source-shift stress tests.
基于 RoPE 的语言模型出现长上下文失效,根源在于 RoPE 在维持稳定 token 偏好与区分相近位置之间存在固有权衡。要判断应处理哪种弱点以及如何处理,需要更精确地刻画 RoPE 在不同上下文长度下训练后模型中的行为。我们通过允许 RoPE 各频率下 query-key 尺度不一致,弥补了先前理论的一个关键局限,使之与实际经验观测高度吻合。我们的理论使得上述脆弱性对单个注意力头与输入可测量,并量化了高频Long-context failures of RoPE-based language models can arise from RoPE's intrinsic tradeoff between maintaining stable token preferences and distinguishing nearby positions. Determining which weakness to address, and how, requires a more precise characterization of RoPE's behavior in trained models across context lengths. We address a key limitation of prior theory by allowing unequal query-key scales across RoPE frequencies, which aligns well with practical empirical observations. Our theory makes both vulnerabilities measurable for individual heads and inputs, and quantifies how high-freque
本文提出一个 Agentic RAG 框架,使 LLM 能够使用逻辑表达式构建检索意图,同时将检索后端简化为基于倒排索引的系统,并表明将检索过程锚定在逻辑查询上可显著降低生成响应中的幻觉。This paper proposes an agentic RAG framework that enables LLMs to formulate retrieval intents using logical expressions while simplifying the retrieval backend to an inverted-index-based system, and shows that anchoring the retrieval process in logical queries substantially reduces hallucinations in generated responses.
DyRAD 使用静态背景反射器和运动追踪的动态点反射器建模动态驾驶场景,渲染完整的距离-方位-多普勒 (RAD) 张量,并通过从雷达信号处理链推导出的固定解析点扩散函数渲染反射器,避免传感器引起的扩散被烘焙到场景表示中。DyRAD is presented, which models dynamic driving scenes using static background reflectors and motion-tracked dynamic point reflectors to render complete range-azimuth-Doppler (RAD) tensors, and renders reflectors through a fixed analytic point-spread function derived from the radar's signal-processing chain, preventing sensor-induced spread from being baked into the scene representation.
即使是良性的链接训练也会在底层安全对齐 agent 不变的情况下,相对于基于文本的通信增加有害合规性;安全对齐需要将多智能体系统作为整体来考虑。This work shows that even benign link training can increase harmful compliance relative to text-based communication while the underlying safety-aligned agents remain unchanged, and shows that safety alignment requires considering the multi-agent system as a whole.
本文提出 ATLAS(Aligned Transport of Latent Structure),一种在显式保持关系几何结构的同时校准全局潜空间分布的训练目标,通过一维 Wasserstein-2 传输进行 Wasserstein 嵌入匹配来校准其边缘分布。This work introduces Aligned Transport of Latent Structure (ATLAS), a training objective that explicitly preserves relational geometry while calibrating the global latent distribution and uses Wasserstein embedding matching to calibrate its marginal through one-dimensional Wasserstein-2 transport.
本文提出 MIST(Misleading-Image Stress Test):200 个英文句子,每句围绕一个可作比喻或字面理解的短语,配以对齐图像(描绘其读法)、误导图像(描绘相反读法)或无图像三种条件。MIST, the Misleading-Image Stress Test: 200 English sentences, each built around a phrase readable either figuratively or literally and shown with an aligned image depicting its reading, a misleading image depicting the opposite, or no image at all is introduced.
我们提出 Endless Exam,一个涵盖十四个参数化数学构造问题族的 benchmark,具有可验证的分数,能够区分在已发表数学前沿之前与之后的进展。每个提交的对象会被自动检验有效性,并依据已发表前沿或构造基线获得相对质量分数,不将改进上限设为 1。该 benchmark 从开放性数学问题中汲取长期挑战,并通过改变参数生成更大规模的实例。紧凑证书使得大型构造能被快速验证We introduce the Endless Exam, a benchmark spanning fourteen parameterised families of mathematical construction problems, with verifiable scores that distinguish progress before and beyond published mathematical frontiers. Each submitted object is checked automatically for validity and assigned a relative quality score against a published frontier or construction baseline, without capping improvements at 1. The benchmark draws long-term challenges from open mathematical problems and generates larger instances by varying their parameters. Compact certificates allow large constructions to be ve
本评估刻画了前沿 agent 如何结合源代码级执行、应用截图与图形交互来生成经过验证的软件变更,考察了跨领域与不同任务信息需求下的表现,以及与成功修复相关的开发行为。This evaluation characterizes how frontier agents combine source-level execution with application screenshots and graphical interaction to produce verified software changes, and examines performance across domains and task information requirements, alongside the development behaviors associated with successful repairs.
本文提出一种方法,通过多层感知机的一次前馈传播生成独立样本(与近期提出的 GAN 类似),并使用 MMD 学习生成可被解码为样本的 codes。This work forms a method that generates an independent sample via a single feedforward pass through a multilayer perceptron, as in the recently proposed generative adversarial networks, using MMD to learn to generate codes that can then be decoded to produce samples.
实验表明,屏蔽高熵偏移位置相比朴素的 SD 提升了分布外泛化能力,由此得到的自蒸馏 judges 在所评估的主观子类别上比基于结果监督 RL 训练的 judges 高出 2-9 个百分点,同时在客观子类别上保持竞争力。Experiments show that masking higher-entropy-shift positions improves out-of-distribution generalization over naive SD, and the resulting self-distilled judges outperform judges trained with outcome-supervised RL by 2-9 percentage points on the evaluated subjective subcategories, while remaining competitive on objective ones.
DAGent 是一个基于 DAG 的多智能体框架,采用先评估再生长的增量规划:Orchestrator 逐批扩展任务图,每一步扩展都以已完成节点的置信度与不确定性信号为条件;证据条件化规划在更低的每任务 token、工具调用和步骤开销下达到了比 Plan-then-Patch 更高的准确率。DAGent, a DAG-based multi-agent framework with Evaluate-then-Grow incremental planning: an Orchestrator grows the task graph one batch at a time, conditioning each expansion on confidence and uncertainty signals from completed nodes, which shows that evidence-conditioned planning reaches higher accuracy at lower per-task token, tool-call, and step footprints than its Plan-then-Patch counterpart.
本文提出 MILO(Meta-evolutionary Island Orchestration),一个共同演化 agent harness 及其发现策略的框架,使用前沿模型(Opus 4.8)与开源权重模型(gpt-oss-120b)超越了八个 SOTA harness 与六种搜索方法。This work introduces MILO (Meta-evolutionary Island Orchestration), a framework that co-evolves agent harnesses and the strategy used to discover them and outperforms eight state-of-the-art harnesses and six search methods using frontier (Opus 4.8) and open-weight (gpt-oss-120b) models.
本文对一个自 2026 年 3 月起持续运行的个人助理 Agent 运行时中的静默失败进行纵向研究,该系统包含约 40 个定时任务、8 个 LLM 提供商、一个工具治理代理以及一个知识库记忆层,由 4,286 个单元测试和 827 项治理检查守护。A longitudinal study of silent failures in a personal-assistant agent runtime in continuous production since March 2026, with roughly 40 scheduled jobs, 8 LLM providers, a tool-governance proxy, and a knowledge-base memory plane, defended by 4,286 unit tests and 827 governance checks is presented.
Tacit-TTS 是一个从 IndexTTS2 蒸馏而来的高效无需转录的零样本语音克隆系统,用掩码非自回归生成替换自回归的文本到语义解码,引入无需训练的声学长度估计,并通过 ReFlow 蒸馏加速流匹配渲染器。Tacit-TTS is presented, an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2 that replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, introduces training-free acoustic length estimation, and accelerates the flow-matching renderer through ReFlow distillation.
统一多模态模型(UMM)通常依赖独立的视觉表示分别完成理解与生成,这增加了视觉上下文长度,并使其难以与既有视觉-语言预训练流程集成。近期 pixel-space modeling 的进展提供了一种无编码器的替代方案,但将该范式从图像扩展到视频并非易事:视频理解与生成采用不同的时间表示,统一视觉接口的设计仍是开放问题。本文提出 PixelUMM,一种用于统一图像与视频理解的无编码器模型……Unified Multimodal Models (UMMs) often rely on separate visual representations for understanding and generation, increasing visual context length and complicating integration with established vision-language pretraining pipelines. Recent advances in pixel-space modeling offer an encoder-free alternative, but extending this paradigm from images to videos is non-trivial: video understanding and generation adopt different temporal representations, leaving the design of a unified visual interface an open question. We present PixelUMM, an encoder-free model for unified image and video understanding
本文提出 BIABench,一个由 16 个从已发表生物研究重建的任务组成的基准,保留了其科学问题、成像数据与真值标注,为评估并最终训练可靠的、面向长程的生物图像分析 agent 提供可验证的框架。BIABench is introduced, a benchmark of 16 tasks reconstructed from published biological studies that retain their scientific questions, imaging data and ground truth, and provides a verifiable framework for evaluating, and eventually training, agents for reliable long-horizon bioimage analysis.
理论分析表明 RWTD 的不动点分布在参考策略的 off-policy 奖励倾斜与当前模型的 on-policy 倾斜之间插值,提供了一种在奖励适配与保留先验知识之间取得平衡的原则性方法。Theoretical analysis shows that the fixed-point distributions of RWTD interpolate between off-policy reward tilting of the reference and on-policy tilting of the current model, providing a principled approach to balancing reward adaptation with retention of prior knowledge.
提出一种自适应、自动化的数据提取攻击流程,在黑盒设置下针对 MRAG(其中检索到的视觉产物本身就是答案)发起攻击,表明亟需专门面向多模态数据设计的安全防护。An adaptive and automatic data extraction attack procedure operating in a black box setting against MRAG, a configuration in which the retrieved visual artifact is itself the response, and shows the urgent need for safeguards specifically designed for multimodal data.
InterEvolve 提出了一个物体感知的前向-后向行为基础模型,其在冻结身体先验上的物体残差可在测试时将关于身体或物体的奖励转化为 loco-manipulation 行为,并将任务以奖励程序的形式指定:带完成条件与可调常数的分阶段奖励。InterEvolve develops an object-aware forward-backward behavioral foundation model, whose object residuals on a frozen body prior turn a new reward about the body or objects into loco-manipulation behavior at test time, and specifies tasks as reward programs: staged rewards with completion conditions and tunable constants.
本文提出 HIDE,一个用于在部分可观测条件下评估操作记忆的基准,并提出 SEEK 框架,结合三种互补的记忆机制来保留历史证据并追踪执行状态。This work introduces HIDE, a benchmark for evaluating manipulation memory under partial observability, and proposes $SEEK, a framework combining three complementary memory mechanisms to retain historical evidence and track execution state, and proposes a framework combining three complementary memory mechanisms to retain historical evidence and track execution state.
结果表明,有效的拒绝能够保留任务结构,同时限制与参考策略的耦合,并且较小的冻结模型可以低成本地提供此类拒绝能力。The results suggest that effective rejects preserve task structure while limiting coupling to the reference policy, and that smaller frozen models can provide them at low cost.
结果表明,纯视觉设置会降低准确率并增加 token 成本,因为 Agent 缺乏足够的符号化细节,需通过重复的视觉查询进行补偿;研究指向一种面向下一代编码 Agent 的实用文本与视觉混合设计。The results show that a strictly vision-only setup degrades accuracy and increases token cost, because agents lack sufficient symbolic detail and compensate with repeated visual queries, and point to a practical hybrid text-and-vision design for next-generation coding agents.
结果表明,HeteroFold 能够在接收端无需预填充的前提下实现高效的跨系列 KV 复用,并在全部四个长上下文基准和大多数短上下文设置上取得最佳的 cache 迁移性能。Results show that HeteroFold enables efficient cross-family KV reuse without receiver prefill, and achieves the best cache-transfer performance on all four long-context benchmarks and most short-context settings.
DataMagic 通过声明式多 Agent 编排,从原始表格数据自动生成数据可视化视频,在完全自动化与细粒度人工控制之间架起桥梁,并提升了创作效率、降低了感知认知负荷。DataMagic, which authors data videos from raw tabular data through declarative multi-agent orchestration, bridging full automation with fine-grained human control, and improves creation efficiency and reduces perceived cognitive load.
本文提出稳定性感知的泛化目标(SAGO),一个用于衡量模型在同一输入上面对不同扰动和基准时行为变化程度的评估框架,涵盖生成一致性、内部激活、置信度以及响应镜像等多个维度的变异性。The Stability-Aware Generalization Objective (SAGO) is introduced, a framework that measures how much model behavior changes for the same input under different variations and benchmarks, capturing variability across several dimensions including generation consistency, internal activations, confidence, and response mirroring.
本文提出 LOCI,一种混合的空间记忆架构,同时维护键值缓存与循环记忆两种表示,在重访内容的复现上比代表性世界模型以及同配置的 full-softmax 模型都更为忠实。LOCI is introduced, a hybrid spatial-memory architecture that keeps both representations of key-value caches and recurrent memory that reproduces revisited content more faithfully than representative world models and a same-recipe full-softmax model.
PhysVista 是一个旨在通过借鉴人类“感知-推理-评估”过程的认知闭环框架来评测 VLM 物理智能的基准,揭示了视觉识别与真实物理理解之间持续存在的差距,并为面向物理基础的多模态智能设计提供了更具原则性的方向。PhysVista is a benchmark designed to evaluate physical intelligence in VLMs through a closed cognitive loop framework inspired by the human seeing-reasoning-assessment process, highlighting a persistent gap between visual recognition and genuine physical understanding, and pointing toward more principled designs for physically grounded multimodal intelligence.
本文提出 AntPlan,一个包含 505 张真实专业建筑平面图、覆盖 92 类物体和十类住宅房间且具有密集家具标注的精选数据集,以及 Architect-Ant,一个用于生成家具布局的框架,可在不依赖高成本迭代式 Agent 推理的情况下直接进行约束感知的布局生成。AntPlan, a curated dataset of 505 real professional architectural floor plans with dense furniture annotations spanning 92 object classes and ten residential room categories, and Architect-Ant, a framework for generating furniture layouts are introduced, enabling direct constraint-aware layout generation without relying on costly iterative agentic inference.
本文提出 JevSpawn,一种将自然语言任务规范连接到有限概率探索的组合策略,并将 JevSpawn 确立为结构化 Agent 推理的一种有前景的方法,在任务性能和导航速度上均有所提升。This work introduces JevSpawn, a compositional policy that connects natural language task specifications to finite probabilistic exploration, and establishes JevSpawn as a promising approach to structured agentic inference, with improved task performance and faster navigation.
SAKIKO 是一个审计框架,通过定向错误发现、路由器条件干预、目标解析验证以及前瞻性冻结统计许可来形式化表示修复,并确立了在声明内部修复之前必须进行结果解析裁定的必要性。SAKIKO is presented, an auditing framework that formalizes representation repair via directional error discovery, router-conditioned intervention, destination-resolved verification, and prospectively frozen statistical licensing and establishes the necessity of outcome-resolved adjudication before claiming internal repair.