EMBODIEDSWE-GEN 将 coding agent 的单一解决方案扩展为大规模多样化轨迹用于训练 VLA,并表明仅在 coding agent 生成的仿真演示上微调的 VLA,即可在真实机器人上完成长时任务。EMBODIEDSWE-GEN, which expands a single solution from coding agent into large diverse trajectories for training a VLA, and shows that a VLA finetuned solely on coding-agent-generated simulation demonstrations completes a long-horizon task on real robot.
论文
286 张论文卡片 · 评测集
介绍 HappyWorld-Bench,一个用于评估生成世界在 Agent 交互过程中是否保持可靠的综合基准,并强调对世界模型的评估不仅应看视觉质量,还应考察状态一致性以及其对动作和干预响应的正确性。HappyWorld-Bench is introduced, a comprehensive benchmark that evaluates whether generated worlds remain reliable as agents interact with them, and highlights the need to evaluate world models not only by visual quality, but also by state consistency and the correctness of their responses to actions and interventions.
本文主张每个 NLP 子领域应将主要性能指标与一个校准分数配对,呼吁将校准视为每个模型的基本属性而非边缘话题。It is argued that each NLP subfield should pair its main performance metric with a calibration score and call for treating calibration as an essential property of every model rather than a niche topic.
本文提出 OmniEchoBench,一个面向空间音视频感知与音-视-语言导航的统一基准,以及一个空间感知的全模态模型,该模型在预训练语义音频通路之外引入 FOA 空间编码器。OmniEchoBench, a unified benchmark for spatial audio-visual perception and audio-vision-language navigation, and a spatially aware omni-modal model, which introduces an FOA spatial encoder alongside a pretrained semantic audio pathway.
本文提出 RLCDAlignBench,在十类对齐失败上对 Jev 进行基准测试:谄媚、越狱、欺骗、提示注入、幻觉、隐私侵犯、社会偏见、奖励黑客、不确定性隐瞒与权力寻求。RLCDAlignBench is presented, which benchmarks Jev on ten alignment failures: sycophancy, jailbreaks, deception, prompt injection, hallucination, privacy violation, social bias, reward hacking, concealing uncertainty, and power seeking.
本文提出一种新颖的 score-purified fusion (SPF) 方法,在所有评测 benchmark 上均达到 SOTA 性能,验证了该方法在利用高质量多模态信息进行 RGB-D 语义分割任务中的有效性。A novel score-purified fusion (SPF) method is proposed, which achieves state-of-the-art performance across all evaluated benchmarks, demonstrating the effectiveness of the approach in leveraging high-quality multimodal information for RGB-D semantic segmentation.
为在传统二元任务成功之外量化协作有效性,本文提出 Causal Collaboration Effectiveness (CCE),一种基于图的指标,用于追踪 agent 动作之间的因果依赖,并度量团队投入中实际促成最终结果的比例。To quantify collaboration effectiveness in addition to conventional binary task success, Causal Collaboration Effectiveness (CCE) is proposed, a graph-based metric that traces causal dependencies between agent actions and measures what fraction of a team's effort actually contributed to the outcome.
本工作提出 AlphaEval,一个基于真实生产环境的基准,包含来自七家在其核心业务中部署 AI Agent 的公司的 94 个任务,覆盖六个 O*NET (Occupational Information Network) 领域;并贡献了一套从需求到基准的构建框架,将从需求到评估的完整流程标准化。This work presents AlphaEval, a production-grounded benchmark of 94 tasks sourced from seven companies deploying AI agents in their core business, spanning six O*NET (Occupational Information Network) domains, and contributes a requirement-to-benchmark construction framework that standardizes the entire pipeline from requirement to evaluation.
[TLDR中文] 双重差分(DID)研究被广泛用于评估气候政策,但评估支持其识别假设的证据仍具挑战。我们提出 ARGUS,一个结构化的语言模型流水线,针对十一维的假设—含义—证据评估标准对所报告的证据进行审计,并在无法检索到相关证据时选择弃答。我们通过注入缺陷、经济学论文以及一个使用经协调标签的小规模试点评估 ARGUS。在 11 类缺陷基准上,ARGUS 检测出 73% 的植入缺陷,而基于关键词的流水线仅能检测出 18%。Difference-in-differences (DID) studies are widely used to evaluate climate policy, but assessing the evidence supporting their identification assumptions remains challenging. We introduce ARGUS, a structured language-model pipeline that audits reported evidence against an eleven-dimension assumption-implication-evidence rubric and abstains when relevant evidence cannot be retrieved. We evaluate ARGUS using injected flaws, economics papers, and a small pilot with reconciled labels. On the 11-flaw benchmark, ARGUS detects 73% of planted flaws, compared with 18% for a keyword-based pipeline. Acr
本文提出 IndicBankBench,一个涵盖五个业务领域、一个能力/拒答领域以及二十个主轴、共 799 个案例的零售银行 benchmark,并提供 case 级诊断分析,以区分那些提出不必要追问的系统与那些采取行动却未能调和客户上下文或完整解决诉求的系统。IndicBankBench is introduced, a 799-case benchmark for Indian retail banking spanning five operational domains, a capability/refusal domain, and twenty primary axes, and a case-level diagnostics that distinguish systems that ask unnecessary questions from those that act but fail to reconcile customer context or fully resolve the request.
本文勾勒了一幅计算机视觉版图:在其中,越来越复杂的视觉任务可通过通用接口访问,而精确且对保真度敏感的感知仍是重要前沿。A changing landscape of computer vision is mapped in which increasingly sophisticated visual tasks are accessible through a general-purpose interface, while precise and fidelity-sensitive perception remains an important frontier.
APM-Bench 将真实场景的流式交互重构为多会话生命轨迹,并揭示了一个清晰的效用—延迟—存储权衡:现有方法仍难以同时实现可靠的长程记忆、低开销以及跨会话有效的主动协助。APM-Bench is introduced, which reformulates real-world streaming interaction as multi-session life trajectories and reveals a clear utility--latency--storage trade-off: existing methods still struggle to simultaneously achieve reliable long-term recall, low overhead, and effective proactive assistance across sessions.
本文提出 DashboardQA,这是首个明确设计用于评估视觉-语言 GUI Agent 对真实世界仪表板理解与交互能力的基准,结果表明交互式仪表板推理对所有受评估的 VLM 而言都是一项具有挑战性的任务。DashboardQA is introduced, the first benchmark explicitly designed to assess how vision-language GUI agents comprehend and interact with real-world dashboards, and indicates that interactive dashboard reasoning is a challenging task overall for all the VLMs evaluated.
本文提出 LibraryDesignBench,这是一个两阶段基准,Agent 根据一份定义所需能力和潜在用例但不规定具体设计的规范,实现一个功能完备的 library。LibraryDesignBench is introduced, a two-phase benchmark in which an agent implements a full-featured library from a specification that defines required capabilities and potential use cases without prescribing the design.
本文提出 OSWorld-Science,这是一个结合科学意义任务、基于 artifact 的评估以及高效 agent harness 的基准与评估环境,用于研究科学领域的计算机使用,从而在科学工作流中系统评估 agent 能力与 harness 设计。OSWorld-Science is introduced, a benchmark and evaluation environment that combines scientifically meaningful tasks, artifact-based evaluation, and an efficient agent harness for studying computer use in the scientific domain, enabling systematic evaluation of both agent capabilities and harness design in scientific workflows.
DyRAD 使用静态背景反射器和运动追踪的动态点反射器建模动态驾驶场景,渲染完整的距离-方位-多普勒 (RAD) 张量,并通过从雷达信号处理链推导出的固定解析点扩散函数渲染反射器,避免传感器引起的扩散被烘焙到场景表示中。DyRAD is presented, which models dynamic driving scenes using static background reflectors and motion-tracked dynamic point reflectors to render complete range-azimuth-Doppler (RAD) tensors, and renders reflectors through a fixed analytic point-spread function derived from the radar's signal-processing chain, preventing sensor-induced spread from being baked into the scene representation.
我们提出 Endless Exam,一个涵盖十四个参数化数学构造问题族的 benchmark,具有可验证的分数,能够区分在已发表数学前沿之前与之后的进展。每个提交的对象会被自动检验有效性,并依据已发表前沿或构造基线获得相对质量分数,不将改进上限设为 1。该 benchmark 从开放性数学问题中汲取长期挑战,并通过改变参数生成更大规模的实例。紧凑证书使得大型构造能被快速验证We introduce the Endless Exam, a benchmark spanning fourteen parameterised families of mathematical construction problems, with verifiable scores that distinguish progress before and beyond published mathematical frontiers. Each submitted object is checked automatically for validity and assigned a relative quality score against a published frontier or construction baseline, without capping improvements at 1. The benchmark draws long-term challenges from open mathematical problems and generates larger instances by varying their parameters. Compact certificates allow large constructions to be ve
本评估刻画了前沿 agent 如何结合源代码级执行、应用截图与图形交互来生成经过验证的软件变更,考察了跨领域与不同任务信息需求下的表现,以及与成功修复相关的开发行为。This evaluation characterizes how frontier agents combine source-level execution with application screenshots and graphical interaction to produce verified software changes, and examines performance across domains and task information requirements, alongside the development behaviors associated with successful repairs.
本文提出 BIABench,一个由 16 个从已发表生物研究重建的任务组成的基准,保留了其科学问题、成像数据与真值标注,为评估并最终训练可靠的、面向长程的生物图像分析 agent 提供可验证的框架。BIABench is introduced, a benchmark of 16 tasks reconstructed from published biological studies that retain their scientific questions, imaging data and ground truth, and provides a verifiable framework for evaluating, and eventually training, agents for reliable long-horizon bioimage analysis.
本文提出 HIDE,一个用于在部分可观测条件下评估操作记忆的基准,并提出 SEEK 框架,结合三种互补的记忆机制来保留历史证据并追踪执行状态。This work introduces HIDE, a benchmark for evaluating manipulation memory under partial observability, and proposes $SEEK, a framework combining three complementary memory mechanisms to retain historical evidence and track execution state, and proposes a framework combining three complementary memory mechanisms to retain historical evidence and track execution state.
本文提出稳定性感知的泛化目标(SAGO),一个用于衡量模型在同一输入上面对不同扰动和基准时行为变化程度的评估框架,涵盖生成一致性、内部激活、置信度以及响应镜像等多个维度的变异性。The Stability-Aware Generalization Objective (SAGO) is introduced, a framework that measures how much model behavior changes for the same input under different variations and benchmarks, capturing variability across several dimensions including generation consistency, internal activations, confidence, and response mirroring.
PhysVista 是一个旨在通过借鉴人类“感知-推理-评估”过程的认知闭环框架来评测 VLM 物理智能的基准,揭示了视觉识别与真实物理理解之间持续存在的差距,并为面向物理基础的多模态智能设计提供了更具原则性的方向。PhysVista is a benchmark designed to evaluate physical intelligence in VLMs through a closed cognitive loop framework inspired by the human seeing-reasoning-assessment process, highlighting a persistent gap between visual recognition and genuine physical understanding, and pointing toward more principled designs for physically grounded multimodal intelligence.
提出 Ego2Act:一个面向目标的基准,包含来自 110 个真实日常任务的 2.640 段视频,覆盖不同的物体杂乱度与多步复杂度;同时提出 Ego2ActJudge,一种无参考评估流水线,在任务完成度与物理合理性评估上与人类共识的对齐效果优于相关基线。Ego2Act is introduced, a goal-directed benchmark featuring 2,640 videos from 110 real-world tasks across day-to-day settings, varying object clutter and multi-step complexity, and Ego2ActJudge, a reference-free evaluation pipeline that achieves better task completion and physics plausibility evaluation alignment with human consensus compared to relevant baselines.
提出了 OpenTumorBoard,一个包含 611 个患者案例、19,157 轮讨论、涵盖十个专科角色的基准,源自 YouTube 上 12,534 分钟公开肿瘤委员会会议录音的转写,并发布了自动化整理流水线,以支持多学科、个性化癌症决策中 LLM 的开发与评估。OpenTumorBoard is introduced, a benchmark with 611 patient cases and 19,157 discussion turns across ten specialist roles, transcribed from 12,534 minutes of publicly available tumor board recordings on YouTube, and automated curation pipeline is released to support the development and evaluation of LLMs for multidisciplinary, personalized cancer decision-making.
提出了 DyadMem,一个双领域、全流程的记忆基准,伴随大量标注工作,并提出了新定义——用户条件关系型智能体记忆(URAM),以推动该领域发展。DyadMem is a dual-domain, full-pipeline memory benchmark with extensive annotation efforts for advancing the domain's development and is proposed with the proposed new definition User-conditioned Relational Agent Memory (URAM).
提出了 World Embedding Benchmark,包含来自 80 个族的 8,000 个受控仿真案例,涵盖流体力学、固体力学、动力学以及光学与电磁学,旨在强调需要联合评估物理一致性与物理属性可恢复性,并展示物理表示对改进视频生成的效用。The World Embedding Benchmark is introduced, comprising 8,000 controlled simulation cases from 80 families spanning fluid mechanics, solid mechanics, dynamics, and optics&electromagnetism, to highlight the need to evaluate physical alignment and property recoverability jointly and demonstrate the utility of physical representations for improving video generation.
提出了 HyperBrowseComp,一个多语言、多模态的浏览基准,包含由母语或高水平使用者编写并经人工验证的、跨 13 种语言的 423 个问题,难度源于在开放网络上发现并关联证据。HyperBrowseComp is introduced, a multilingual and multimodal browsing benchmark comprising 423 manually authored and human-validated questions across 13 languages, written by native or highly proficient speakers, with difficulty arising from discovering and connecting evidence on the open web.
SimuVerity 为评估智能体在可执行 Simulink 模型生成中的工程能力与诊断失败提供了系统性基础,并表明结构相似性是衡量工程性能的一个糟糕代理指标。SimuVerity provides a systematic basis for assessing agents' engineering capabilities and diagnosing failures in executable Simulink model generation, and shows that structural similarity is a poor proxy for engineering performance.
提出 4DCodeBench,一个面向 4D 逆向图形(通过代码生成)的基准,Agent 以可执行图形程序的形式从视频重建动态场景,结果表明强大的静态重建能力尚未转化为对复杂动态的可靠重建。4DCodeBench is introduced, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs, finding that strong static reconstruction capabilities do not yet translate into reliable reconstruction of complex dynamics.
提出 ACG-Bench,一个面向逐臂组合泛化(arm-wise Compositional Generalization)的基准,为超越固定训练流程的双臂策略设计提供实证指导,考察了 arm-token grouping、技能专属 LoRA 适配器(SkillLoRA)以及 arm-wise attention(AWA),凸显了技能条件化参数与注意力结构的互补性。This work introduces ACG-Bench, a benchmark for arm-wise Compositional Generalization that provides empirical guidance for designing dual-arm policies that generalize beyond fixed training routines, and examines arm-token grouping, skill-specific LoRA adapters (SkillLoRA), and arm-wise attention (AWA), highlighting the complementarity of skill-conditioned parameters and attention structure.
引入 Mathematical Primitive 概念来探查结构性数学理解,并提出 \hlei{},一个沿发现(Discovery)、生成(Generation)、消化(Digestion)和执行(Execution)四个维度评估数学推理的新型基准。This paper introduces the notion of Mathematical Primitive to probe structural mathematical understanding and proposes \hlei{}, a novel benchmark that evaluates mathematical reasoning along four distinct dimensions: Discovery, Generation, Digestion, and Execution.
提出 UndoBench,一个覆盖 8 个企业领域、36 个基础工作流与 36 个故障场景的基准测试,通过相同种子下的反事实配对试验以及链路级效应历史与环境状态预言机,将任务能力与恢复能力解耦。UndoBench is introduced, a benchmark spanning 36 base workflows and 36 fault scenarios across 8 enterprise domains, decoupling task competence from recovery capability via counterfactual paired trials under identical seeds alongside wire-level effect-history and environment-state oracles.
检索增强生成(RAG)系统易遭受嵌入在检索内容中的提示注入攻击。我们提出 RAG-PIBench,一个面向 RAG 风格提示注入检测的基准,包含跨冻结训练、验证与受保护测试划分的 4,876 个上下文示例。基于泄露感知构建流水线与严格评估协议,我们对比了基于关键词、语义引用、TF-IDF 与 Transformer 的检测器。DistilBERT 在受保护测试集上取得最佳性能(F1=0.896,PR-AUC=0.968),而 TF-IDF SVM 与逻辑回归仍具竞争力...Retrieval-Augmented Generation (RAG) systems are vulnerable to prompt-injection attacks embedded in retrieved content. We introduce RAG-PIBench, a benchmark for RAG-style prompt-injection detection containing 4,876 contextual examples across frozen train, validation, and protected-test splits. Using a leakage-aware construction pipeline and strict evaluation protocol, we compare keyword-based, semantic-reference, TF-IDF, and transformer-based detectors. DistilBERT achieves the best protected-test performance (F1 = 0.896, PR-AUC = 0.968), while TF-IDF SVM and logistic regression remain competit
场地特异性施肥推荐系统会根据地点、土壤属性、作物类型和生产目标调整养分建议,但当推荐功能主要通过交互式界面访问、输出未做版本管理、训练后的近似模型无法独立加载或基准测试时,其科学复用性受到限制。本技术报告介绍 Turba 施肥机器学习技术栈——面向摩洛哥场地特异性施肥推荐的可复现三层开源实现,其中 turba-client 提供对公开……Site-specific fertilizer recommendation systems adapt nutrient advice to location, soil properties, crop type, and production targets, but scientific reuse is constrained when recommendation functions remain accessible mainly through interactive interfaces, outputs are not versioned, and trained approximations cannot be independently loaded or benchmarked. This technical report presents the Turba fertilizer machine learning stack, a three-layer open-source implementation for reproducible site-specific fertilizer recommendation in Morocco. turba-client provides programmatic access to publicly a
提出 RoboDojo,一个面向通用机器人操作策略综合评估的仿真-真机统一基准,将 30 种策略集成到 XPolicyLab 并在 RoboDojo 上进行评测,建立了公开的排行榜与系统性的策略性能分析。RoboDojo is introduced, a unified sim-and-real benchmark for comprehensive evaluation of generalist robot manipulation policies that integrates 30 policies into XPolicyLab and evaluates them on RoboDojo, establishing a public leaderboard and systematic analysis of current policy performance.
多模态大语言模型(MLLMs)正快速发展以实现持续的音视频推理,这迫切需要能够揭示其能力上限的评估。音视频描述是一项理想的诊断任务,但现有 benchmark 面临耦合的权衡:整句评分覆盖全面但缺乏定位,局部探针定位精确但缺乏覆盖,且无约束的 LLM 评判器带来不稳定性。我们提出 OmniCapBench(Omni-Video Caption Benchmark),将音视频描述评估重构为深度结构化的诊断框架Multimodal large language models (MLLMs) are rapidly evolving toward continuous audio--visual reasoning, creating an urgent need for evaluations that expose their capability limits. Audio--visual captioning is an ideal diagnostic task, yet current benchmarks face a coupled trade-off: whole-caption scores provide coverage without localization, local probes provide localization without coverage, and unconstrained LLM judges introduce instability. We introduce OmniCapBench (Omni-Video Caption Benchmark), a benchmark that reframes audio--visual caption evaluation as a deep-structured diagnostic fr