提出 MIRROR——一个统一的跨表层框架,在显式新颖性约束下以检索到的上下文为条件生成候选,并执行记忆引导的蒙特卡洛树搜索,使检索可影响搜索先验,同时避免提示词级别的复制。MIRROR is presented, a unified cross-surface framework that performs memory-guided Monte Carlo tree search while conditioning candidate generation on retrieved context under an explicit novelty constraint, allowing retrieval to inform search priors without enabling prompt copying.
论文
274 张论文卡片 · 评测集 · OA 绿色
本工作将带外防御组织为经典完整性保护、引用监控与最小权限的具体实例,对它们覆盖与未覆盖的内容进行结构化对比;与该假设一致但尚未被证实的是:确定性的带外强制执行相比带内检测,是更难被自适应攻击者攻破的目标。This work organizes out-of-band defenses as instances of classical integrity protection, reference monitoring, and least privilege, yielding a structured comparison of what they do and do not cover, consistent with, but not established, the hypothesis that deterministic out-of-band enforcement is a harder target for an adaptive attacker than in-band detection.
所得到的模型在大多数数据集-预测步长组合上优于先前的线性预测器,并在八个基准中的六个上超越 Transformer、MLP 和 CNN 基线;同时它还可作为对数据本身的诊断工具,揭示那些被更大模型默默吸收进其学习参数中的结构。The resulting models beat prior linear forecasters on most dataset-horizon entries and exceed Transformer, MLP, and CNN baselines on six of eight benchmarks, and serve as a diagnostic on the data itself, revealing structures that larger models absorb silently into their learned parameters.
提出 GauntletBench,一个用于评估 Agent 在挑战性场景中泛化能力的 Web 基准,聚焦于三种被低估的能力(时间感知、图形理解与 3D 推理),揭示了当前 Agent 能力与复杂真实场景所需能力之间的巨大差距。GauntletBench, a web-based benchmark for evaluating agent generalisation in challenging scenarios, focusing on three underexplored capabilities (temporal perception, graphical understanding, and 3D reasoning), is introduced, revealing the substantial gap between current agent capabilities and those required for complex real-world scenarios.
本文展示了用于 IT-Grundschutz(IT-GS)认证部分自动化的多 Agent 系统(MAS)架构结合混合检索增强生成(HybridRAG)的技术实现与实证评估,并为强化合规严谨性引入两项新的 MAS 架构技术贡献。This paper presents the technical implementation and empirical evaluation of a Multi-Agent System (MAS) architecture combined with Hybrid Retrieval Augmented Generation (HybridRAG) for the partial automation of IT-GS certification and introduces two novel technical contributions to the MAS architecture to enforce the compliance rigor.
结果表明,工具使用评测应从函数调用准确率转向不可靠工具环境下的任务完成度,并建议工具使用评测应从函数调用准确率转向不可靠工具环境下的任务完成度。(注:原文末句疑似重复)Results suggest that tool-use evaluation should move beyond function-call accuracy toward task completion under unreliable tool environments, and suggest that tool-use evaluation should move beyond function-call accuracy toward task completion under unreliable tool environments.
本文表明强化学习(RL)后训练已具备实现有效 step-level 评分所需的要素,从而完全无需额外的奖励模型训练,并在通用随机 Markov 决策过程下推导出一种隐式 advantage,称为 progress advantage。This work shows that reinforcement learning (RL) post-training already provides the ingredients for effective step-level scoring, eliminating the need for dedicated reward model training altogether, and derives an implicit advantage under a general stochastic Markov decision process, which is term progress advantage.
HAKARI-Bench 是一个轻量级基准,将现有检索套件重建为统一格式的小型数据集(Nano-sets),支持在同一条件下对五类检索方法及其效率变体进行与模型无关的对比。HAKARI-Bench is a lightweight benchmark that reconstructs existing retrieval suites into small datasets (Nano-sets) in a unified format, enabling same-condition, model-agnostic comparison of five retrieval families and their efficiency variants.
一个包含 382 个真实企业任务、覆盖 6 种专业角色和 22 项程序性技能的基准,用于评估技能在任务、角色和模型骨干间的迁移能力,发现部分技能可在任务和模型间广泛泛化,而另一些则专化为角色特定工作流,在迁移时失去效力。A benchmark of 382 realistic enterprise tasks spanning six professional roles and 22 procedural skills, designed to evaluate how skills transfer across tasks, roles, and model backbones finds that some skills generalize broadly across tasks and models, whereas others become specialized to role-specific workflows and lose effectiveness under transfer.
ABACUS 提出三项贡献:与基于多头自注意力分解得到的目标性图相结合的密度感知自适应缩放,用于在空间上锚定计数预测;通过 GRPO 训练的边界感知计数策略,配合嵌套的局部、边界与全局奖励,以消除裁剪边界处的过度与不足计数。ABACUS introduces three contributions: density-aware adaptive zooming paired with an objectness map from multi-head self-attention decomposition to spatially ground count predictions; a boundary-aware count policy trained via GRPO with nested local, boundary, and global rewards to eliminate over- and undercounting at crop boundaries.
本文命名并测量了推测性查询的检索收敛到包含答案结果时的输入流位置——"工具意图稳定化":即推测性查询的检索收敛到包含答案结果的输入流位置。This work names and measures the point in the input stream at which a speculative query's retrieval converges on the answer-bearing result, tool-intent stabilization: the point in the input stream at which a speculative query's retrieval converges on the answer-bearing result.
提出 GateMem,一个面向多主体共享内存 Agent 的基准,联合评估合法长程请求及其状态更新的效用、跨上下文授权边界的访问控制,以及 Agent 在收到显式删除请求后的主动遗忘能力。GateMem is introduced, a benchmark for multi-principal shared-memory agents that jointly evaluates utility for legitimate long-horizon requests with state updates, access control across contextual authorization boundaries, and agent-facing active forgetting after explicit deletion requests.
对涵盖前沿闭源与开源模型、共 22 个模型在多个规模上的评估发现,所有模型家族均存在不可忽略的隐私泄露,且指令遵循能力与泄露率呈正相关。Evaluating 22 models spanning frontier proprietary and open-source models at multiple scales, it is found that all model families exhibit non-trivial leakage, and that instruction- following ability correlates with leakage rate.
结果表明仅评估求解器或仅评估答案是不足的:Agent 的差异不仅体现在发现关键 contingency 上,还体现在验证预算的使用、显式提交、类型强制、重复验证、基于证据的报告以及缓解行为等方面。The results show why solver-only or answer-only evaluation is insufficient: agents are distinguished not only by top-contingency discovery, but also by validation-budget use, explicit submission, type coercions, duplicate validations, evidence-backed reporting, and mitigation behavior.
结果表明模型倾向于选择有害场景,在中性预订选项上的表现低于随机猜测水平,其中 Claude 4.8 取得最高分 64.7%。The results indicate that models tend to prefer harmful scenarios, performing below the random chance rate for selecting a neutral booking option, with Claude $4.8$ achieving the highest performance at $64.7\%$.
新加坡 AI Safety Institute 与韩国 AI Safety Institute 联合评估了涵盖客服、DevOps、网页自动化以及企业与个人生产力场景下 12 项真实非对抗任务中的 Agent 数据泄露问题,表明操作性数据泄露是与对抗性数据外泄不同的一阶 Agent 安全问题。A joint evaluation by the Singapore AI Safety Institute and the Korea AI Safety Institute examining agent data leakage in 12 realistic, non-adversarial tasks spanning customer support, DevOps, web automation, and enterprise and personal productivity indicates that operational data leakage is a first-order agent-safety concern distinct from adversarial exfiltration.
对 Agent 行为的分析揭示了长视野经济交互中的显著差异:表现更好的模型与其他企业的沟通更为活跃,而 Claude Haiku 4.5 则表现出 idle-drift 失效模式,在生成连贯评估与规划的同时仍反复选择不行动。Analysis of agent behavior reveals substantial differences in long-horizon economic interaction: higher-performing models communicate more actively with other firms, whereas Claude~Haiku~4.5 exhibits an idle-drift failure mode, repeatedly choosing inaction despite producing coherent assessments and plans.
本文提出了一种动作分级伤害评估量表,按七级有序量表对智能体的工具调用轨迹进行打分,评判依据包括所执行动作的可逆性、是否越界涉及其他方以及是否扩大了权限。An action-graded harm rubric is introduced that scores an agent's tool-call trajectory on a seven-level ordinal scale according to whether the executed action was reversible, whether it crossed scope to reach another party, and whether it expanded privilege.
提出 AgentLens,一个面向交互式代码 Agent 的生产级评估基准,将形式化验证(在存在客观检查时)与 LLM 编写的轨迹评审及并排比较相结合,使每次运行都能给出关于分数为何如此的可读解释。This work presents AgentLens, a production-assessed benchmark for interactive code agents that pairs formal verification, where an objective check exists, with LLM-written trajectory reviews and side-by-side comparisons, so that each run yields a readable explanation of why the score is what it is.
提出首个面向动态真实世界场景评估主动 agent 的能力驱动基准,并在模型与框架层面的全面比较表明,基模型能力与 agent 框架设计共同决定了真实世界环境中的性能表现。The first capability-driven benchmark designed to evaluate proactive agents in dynamic, real-world settings is introduced, and comprehensive comparisons across both models and frameworks show how base model capabilities and agent framework designs jointly shape performance in real-world environments.
该审计揭示现有基准样本中 55% 可在无视觉输入或时序上下文的情况下被解决,并提出 Video-Oasis,一个用于系统性审计现有视频理解基准的可持续诊断套件。This audit reveals that 55\% of existing benchmark samples are solvable without visual input or temporal context, and introduces Video-Oasis, a sustainable diagnostic suite for systematically auditing existing video understanding benchmarks.
Soofi S 30B-A3B 是一个主权开源的 MoE 混合 Mamba Transformer 德英基础模型,在作者的对比中超越了所有欧洲主权基线,包括活跃参数量远超自身的模型。Soofi S 30B-A3B is a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English that outperforms every European sovereign baseline in the authors' comparison, including ones far larger in active parameters.
该工作提出 LongMedBench,一个基于真实 EHR 的长程临床决策基准,并设计了一套包含三个评估维度的分类体系:事实型问答、时序推理、长程决策。This work introduces LongMedBench, a real-world EHR-based benchmark for long-horizon clinical decision-making, and proposes an evaluation taxonomy with three suites: fact-based QA, temporal reasoning, and long-horizon decision-making.
该工作提出 Long-Horizon-Terminal-Bench,一个涵盖九大类共 46 个长程任务的终端基准,包括实验复现、软件工程、多模态分析、交互游戏与科学计算,并分析失败模式与错误规律,以推动长程终端 Agent 的后续研究。This work introduces Long-Horizon-Terminal-Bench, a terminal benchmark of 46 long-horizon tasks spanning nine categories, including experiment reproduction, software engineering, multimodal analysis, interactive games, and scientific computing, and analyzes failure modes and error patterns to support future progress on long-horizon terminal agents.
论文提出 Flow-ERD,一个同时追求真实性与多样性的多 Agent 仿真器,在 WOSAC 测试基准上排名第一,并在可复现基线中主导真实性-多样性 Pareto 前沿。Flow-ERD is introduced, a multi-agent simulator that pursues realism and diversity jointly and ranks first on the WOSAC test benchmark and dominates the realism--diversity Pareto front among reproducible baselines.
本工作提出 AdvancedMathBench,用于评估 LLM 在高级数学证明上的推理能力;同时引入 VerifierBench,包含 888 条由模型生成的证明轨迹及专家真值,用于评估模型能否正确判断证明有效性并给出合理的验证依据。This work introduces AdvancedMathBench, a benchmark suite designed to evaluate the reasoning capabilities of LLMs on advanced mathematical proofs, and introduces VerifierBench, consisting of 888 model-generated proof trajectories paired with expert ground truth, to evaluate whether models can correctly judge proof validity and provide sound verification rationales.
本文提出 MCLASH,一个多语言道德决策 benchmark,用于捕捉跨语言的文化情境化道德直觉和社会规范;并提出 MET(Multilingual Ethics with Theory-grounded reasoning),一种基于心理学与哲学专家策划的理论依据的两步提示方法。This work introduces MCLASH, a multilingual moral decision-making benchmark to capture culturally situated moral intuitions and social norms across languages, and proposes MET (Multilingual Ethics with Theory-grounded reasoning), a two-step prompting method built on expert-curated, theory-based grounds drawn from psychology and philosophy.
开发了自动化评分流水线,用于评估多种模型,包括开源权重模型、闭源语言模型、视觉语言模型和图像生成模型,结果显示没有单一模型在所有任务类型上占优,且某些任务对所有评估模型仍具挑战性。An automated grading pipeline is developed to evaluate a wide range of models, including open-weight and closed-source language, vision-language, and image-generation models, and shows that no single model dominates across all task types, and that some tasks remain challenging for all evaluated models.
提出 LakeQuest,一个经人工校验的基准,包含 9,846 个 QA 对,用于在真实数据湖场景下评估端到端检索与综合 pipeline,并揭示现代问答系统中的关键失效模式。LakeQuest is introduced, a human-validated benchmark of 9,846 QA pairs designed to evaluate the end-to-end retrieve-and-synthesize pipeline over realistic data lakes and exposes critical failure modes in modern QA systems.
提出 SDABench,一个围绕六项能力(描述性、探索性、推断性、推断性、预测性、因果性、机制性)跨五大领域(生物、化学、环境、地理、物理)重新组织评估的基准。SDABench is introduced, a benchmark that reorganizes evaluation around six capabilities (descriptive, exploratory, inferential, inferential, predictive, causal, and mechanistic) across five domains (Biology, Chemistry, Environment, Geography, Physics).
视觉语言模型(VLMs)在 DocVQA、ChartQA、MMLongBench-Doc 等视觉文档理解基准上表现强劲。但真实文档融合长度、布局复杂度、模态、问题难度等多因素,难以将模型失败归因于具体原因。我们提出 SynthDocBench,一个全合成的长上下文视觉文档理解基准,系统性控制文档长度、布局结构、模态构成与问题类型等因子。该基准的构建……Vision language models (VLMs) have achieved strong performance on visual document understanding benchmarks such as DocVQA, ChartQA, and MMLongBench-Doc. However, real-world documents combine multiple factors such as length, layout complexity, modality, and question difficulty, which makes it difficult to attribute model failures to specific causes. We introduce SynthDocBench, a fully synthetic benchmark for long-context visual document understanding that systematically controls factors including document length, layout structure, modality composition, and question type. The benchmark is constr
论文提出 PolicyShiftGuard,一个紧凑的策略条件护栏,采用结合随机策略 SFT(RP-SFT)与边界对策略适配(BP-Adapt)的两阶段训练方案,并验证匹配的通过/拒绝边界对是稳定策略适配的关键。This work proposes PolicyShiftGuard, a compact policy-conditioned guardrail trained with a two-stage training recipe that combines Randomized Policy SFT (RP-SFT) with Boundary-Pair Policy Adaptation (BP-Adapt), and confirms that matched pass/block boundary pairs are essential for stable policy adaptation.
AgentCompass 被提出,它是一个开源、轻量且可扩展的面向 LLM-based Agent 的评估基础设施,将评估流程围绕三个独立组件组织,从而在不重新实现复杂执行逻辑的前提下支持灵活配置。AgentCompass is introduced, an open-source, lightweight, and extensible infrastructure for evaluating LLM-based agents that organizes the evaluation process around three independent components, thereby enabling flexible configurations without requiring the reimplementation of complex execution logic.
本文提出一种实用评估协议,将评估重点从任务完成转向经过验证的漏洞发现,可在涵盖多种攻击面与漏洞类型的足够复杂目标上开展评估,并结合结构化真值标注与基于 LLM 的语义匹配来识别漏洞。This paper presents a practical evaluation protocol that shifts assessment from task completion to validated vulnerability discovery, allowing evaluation in sufficiently complex targets spanning multiple attack surfaces and vulnerability classes, and combines structured ground-truth with LLM-based semantic matching to identify vulnerabilities.
本文提出 SIS-Bench,一个在统一 self-in-space 表述下评估 UAV 场景具身空间智能的基准,并探索了一种融合光流与视觉特征的运动感知表征,以纳入与自身相关的动态信息。SIS-Bench is introduced, a benchmark for evaluating embodied spatial intelligence in UAV scenarios under a unified self-in-space formulation, and a motion-aware representation that incorporates self-related dynamics through optical flow and visual feature fusion is explored.
本文将关键帧执行拆解为存在性、保真度、时序顺序、定位、持续性与唯一性六个互补指标,并通过结合专用感知模型的、基于证据的 MLLM 判断来评估整体视频质量。This work decomposes keyframe execution into six complementary metrics covering presence, fidelity, temporal ordering, localization, persistence, and uniqueness, while assessing overall video quality through evidence-grounded MLLM judgments augmented with specialized perception models.