研究库 论文知识库
Papers · organized/paper_cards

论文

274 张论文卡片 · 评测集 · OA 绿色

开放获取 全部 绿色 · 1640
RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation
RGBD20K:面向 RGB-D 语义分割的大规模基准
arXiv:2609.29028 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一种新颖的 score-purified fusion (SPF) 方法,在所有评测 benchmark 上均达到 SOTA 性能,验证了该方法在利用高质量多模态信息进行 RGB-D 语义分割任务中的有效性。A novel score-purified fusion (SPF) method is proposed, which achieves state-of-the-art performance across all evaluated benchmarks, demonstrating the effectiveness of the approach in leveraging high-quality multimodal information for RGB-D semantic segmentation.

AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs
AgentWorld:多 Agent LLM 长期协作基准
arXiv:2609.31590 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

为在传统二元任务成功之外量化协作有效性,本文提出 Causal Collaboration Effectiveness (CCE),一种基于图的指标,用于追踪 agent 动作之间的因果依赖,并度量团队投入中实际促成最终结果的比例。To quantify collaboration effectiveness in addition to conventional binary task success, Causal Collaboration Effectiveness (CCE) is proposed, a graph-based metric that traces causal dependencies between agent actions and measures what fraction of a team's effort actually contributed to the outcome.

1. AlphaEval: Evaluating Agents in Production
AlphaEval: 在生产环境中评估 Agent
arXiv:2604.12162 评测基准 评测集 OA · 绿色 被引 1 · S2

本工作提出 AlphaEval,一个基于真实生产环境的基准,包含来自七家在其核心业务中部署 AI Agent 的公司的 94 个任务,覆盖六个 O*NET (Occupational Information Network) 领域;并贡献了一套从需求到基准的构建框架,将从需求到评估的完整流程标准化。This work presents AlphaEval, a production-grounded benchmark of 94 tasks sourced from seven companies deploying AI agents in their core business, spanning six O*NET (Occupational Information Network) domains, and contributes a requirement-to-benchmark construction framework that standardizes the entire pipeline from requirement to evaluation.

Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations
[标题中文] 气候政策因果评估中识别假设的循证审计
arXiv:2609.30867 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

[TLDR中文] 双重差分(DID)研究被广泛用于评估气候政策,但评估支持其识别假设的证据仍具挑战。我们提出 ARGUS,一个结构化的语言模型流水线,针对十一维的假设—含义—证据评估标准对所报告的证据进行审计,并在无法检索到相关证据时选择弃答。我们通过注入缺陷、经济学论文以及一个使用经协调标签的小规模试点评估 ARGUS。在 11 类缺陷基准上,ARGUS 检测出 73% 的植入缺陷,而基于关键词的流水线仅能检测出 18%。Difference-in-differences (DID) studies are widely used to evaluate climate policy, but assessing the evidence supporting their identification assumptions remains challenging. We introduce ARGUS, a structured language-model pipeline that audits reported evidence against an eleven-dimension assumption-implication-evidence rubric and abstains when relevant evidence cannot be retrieved. We evaluate ARGUS using injected flaws, economics papers, and a small pilot with reconciled labels. On the 11-flaw benchmark, ARGUS detects 73% of planted flaws, compared with 18% for a keyword-based pipeline. Acr

IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking
[标题中文] IndicBankBench:评估语言模型助手在印度零售银行中的安全性与可靠性
arXiv:2609.29167 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 IndicBankBench,一个涵盖五个业务领域、一个能力/拒答领域以及二十个主轴、共 799 个案例的零售银行 benchmark,并提供 case 级诊断分析,以区分那些提出不必要追问的系统与那些采取行动却未能调和客户上下文或完整解决诉求的系统。IndicBankBench is introduced, a 799-case benchmark for Indian retail banking spanning five operational domains, a capability/refusal domain, and twenty primary axes, and a case-level diagnostics that distinguish systems that ask unnecessary questions from those that act but fail to reconcile customer context or fully resolve the request.

Hard Vision, Easy Vision: What GPT-6 Astra Reveals Across Computer Vision
难视觉,易视觉:GPT-6 Astra 揭示的计算机视觉全貌
arXiv:2609.35718 多模态 评测集 OA · 绿色 被引 1 · S2

本文勾勒了一幅计算机视觉版图:在其中,越来越复杂的视觉任务可通过通用接口访问,而精确且对保真度敏感的感知仍是重要前沿。A changing landscape of computer vision is mapped in which increasingly sophisticated visual tasks are accessible through a general-purpose interface, while precise and fidelity-sensitive perception remains an important frontier.

APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants
[标题中文] APM-Bench:面向自我中心流式视频助手的跨会话持久记忆基准
arXiv:2609.37559 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

APM-Bench 将真实场景的流式交互重构为多会话生命轨迹,并揭示了一个清晰的效用—延迟—存储权衡:现有方法仍难以同时实现可靠的长程记忆、低开销以及跨会话有效的主动协助。APM-Bench is introduced, which reformulates real-world streaming interaction as multi-session life trajectories and reveals a clear utility--latency--storage trade-off: existing methods still struggle to simultaneously achieve reliable long-term recall, low overhead, and effective proactive assistance across sessions.

Can Agents Design Libraries for Agents?
Agent 能否为 Agent 设计类库?
arXiv:2609.36730 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 LibraryDesignBench,这是一个两阶段基准,Agent 根据一份定义所需能力和潜在用例但不规定具体设计的规范,实现一个功能完备的 library。LibraryDesignBench is introduced, a two-phase benchmark in which an agent implements a full-featured library from a specification that defines required capabilities and potential use cases without prescribing the design.

OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software
OSWorld-Science:面向科学与科研软件学习与使用的计算机操作 Agent 基准
arXiv:2609.39903 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 OSWorld-Science,这是一个结合科学意义任务、基于 artifact 的评估以及高效 agent harness 的基准与评估环境,用于研究科学领域的计算机使用,从而在科学工作流中系统评估 agent 能力与 harness 设计。OSWorld-Science is introduced, a benchmark and evaluation environment that combines scientifically meaningful tasks, artifact-based evaluation, and an efficient agent harness for studying computer use in the scientific domain, enabling systematic evaluation of both agent capabilities and harness design in scientific workflows.

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
DyRAD:面向动态驾驶场景的雷达新视角合成
arXiv:2609.39841 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

DyRAD 使用静态背景反射器和运动追踪的动态点反射器建模动态驾驶场景,渲染完整的距离-方位-多普勒 (RAD) 张量,并通过从雷达信号处理链推导出的固定解析点扩散函数渲染反射器,避免传感器引起的扩散被烘焙到场景表示中。DyRAD is presented, which models dynamic driving scenes using static background reflectors and motion-tracked dynamic point reflectors to render complete range-azimuth-Doppler (RAD) tensors, and renders reflectors through a fixed analytic point-spread function derived from the radar's signal-processing chain, preventing sensor-induced spread from being baked into the scene representation.

The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence
无尽之试:当代模型迈向超智能的数学构造
arXiv:2609.24555 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

我们提出 Endless Exam,一个涵盖十四个参数化数学构造问题族的 benchmark,具有可验证的分数,能够区分在已发表数学前沿之前与之后的进展。每个提交的对象会被自动检验有效性,并依据已发表前沿或构造基线获得相对质量分数,不将改进上限设为 1。该 benchmark 从开放性数学问题中汲取长期挑战,并通过改变参数生成更大规模的实例。紧凑证书使得大型构造能被快速验证We introduce the Endless Exam, a benchmark spanning fourteen parameterised families of mathematical construction problems, with verifiable scores that distinguish progress before and beyond published mathematical frontiers. Each submitted object is checked automatically for validity and assigned a relative quality score against a published frontier or construction baseline, without capping improvements at 1. The benchmark draws long-term challenges from open mathematical problems and generates larger instances by varying their parameters. Compact certificates allow large constructions to be ve

CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering
CUA-SWE:当 Computer-Use Agents 遇见可视化软件工程
arXiv:2609.32600 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本评估刻画了前沿 agent 如何结合源代码级执行、应用截图与图形交互来生成经过验证的软件变更,考察了跨领域与不同任务信息需求下的表现,以及与成功修复相关的开发行为。This evaluation characterizes how frontier agents combine source-level execution with application screenshots and graphical interaction to produce verified software changes, and examines performance across domains and task information requirements, alongside the development behaviors associated with successful repairs.

BIABench: Evaluating AI agents on real-world bioimage analysis tasks
BIABench: 在真实生物图像分析任务上评估 AI agent
arXiv:2609.34274 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 BIABench,一个由 16 个从已发表生物研究重建的任务组成的基准,保留了其科学问题、成像数据与真值标注,为评估并最终训练可靠的、面向长程的生物图像分析 agent 提供可验证的框架。BIABench is introduced, a benchmark of 16 tasks reconstructed from published biological studies that retain their scientific questions, imaging data and ground truth, and provides a verifiable framework for evaluating, and eventually training, agents for reliable long-horizon bioimage analysis.

Benchmarking and Enhancing Skill-Level Memory for Partially Observable Robotic Manipulation
在部分可观测机器人操作任务上对技能级记忆的基准评测与增强
arXiv:2609.38886 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 HIDE,一个用于在部分可观测条件下评估操作记忆的基准,并提出 SEEK 框架,结合三种互补的记忆机制来保留历史证据并追踪执行状态。This work introduces HIDE, a benchmark for evaluating manipulation memory under partial observability, and proposes $SEEK, a framework combining three complementary memory mechanisms to retain historical evidence and track execution state, and proposes a framework combining three complementary memory mechanisms to retain historical evidence and track execution state.

Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs
泛化即稳定性,而非准确率:LLM 的多轴评估
arXiv:2610.01428 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出稳定性感知的泛化目标(SAGO),一个用于衡量模型在同一输入上面对不同扰动和基准时行为变化程度的评估框架,涵盖生成一致性、内部激活、置信度以及响应镜像等多个维度的变异性。The Stability-Aware Generalization Objective (SAGO) is introduced, a framework that measures how much model behavior changes for the same input under different variations and benchmarks, capturing variability across several dimensions including generation consistency, internal activations, confidence, and response mirroring.

PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop
PhysVista:通过感知-推理-评估闭环评测 VLM 的物理智能
arXiv:2610.00559 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

PhysVista 是一个旨在通过借鉴人类“感知-推理-评估”过程的认知闭环框架来评测 VLM 物理智能的基准,揭示了视觉识别与真实物理理解之间持续存在的差距,并为面向物理基础的多模态智能设计提供了更具原则性的方向。PhysVista is a benchmark designed to evaluate physical intelligence in VLMs through a closed cognitive loop framework inspired by the human seeing-reasoning-assessment process, highlighting a persistent gap between visual recognition and genuine physical understanding, and pointing toward more principled designs for physically grounded multimodal intelligence.

Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation
Ego2Act:评估自我中心视频生成中的目标导向操作
arXiv:2610.01092 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Ego2Act:一个面向目标的基准,包含来自 110 个真实日常任务的 2.640 段视频,覆盖不同的物体杂乱度与多步复杂度;同时提出 Ego2ActJudge,一种无参考评估流水线,在任务完成度与物理合理性评估上与人类共识的对齐效果优于相关基线。Ego2Act is introduced, a goal-directed benchmark featuring 2,640 videos from 110 real-world tasks across day-to-day settings, varying object clutter and multi-step complexity, and Ego2ActJudge, a reference-free evaluation pipeline that achieves better task completion and physics plausibility evaluation alignment with human consensus compared to relevant baselines.

OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories
OpenTumorBoard:多学科肿瘤委员会讨论轨迹的真实世界基准
arXiv:2609.32810 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出了 OpenTumorBoard,一个包含 611 个患者案例、19,157 轮讨论、涵盖十个专科角色的基准,源自 YouTube 上 12,534 分钟公开肿瘤委员会会议录音的转写,并发布了自动化整理流水线,以支持多学科、个性化癌症决策中 LLM 的开发与评估。OpenTumorBoard is introduced, a benchmark with 611 patient cases and 19,157 discussion turns across ten specialist roles, transcribed from 12,534 minutes of publicly available tumor board recordings on YouTube, and automated curation pipeline is released to support the development and evaluation of LLMs for multidisciplinary, personalized cancer decision-making.

DyadMem: A Long-Term Memory Benchmark of How Agents Work with Users
DyadMem:关于 Agent 如何与用户协作的长期记忆基准
arXiv:2610.03020 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出了 DyadMem,一个双领域、全流程的记忆基准,伴随大量标注工作,并提出了新定义——用户条件关系型智能体记忆(URAM),以推动该领域发展。DyadMem is a dual-domain, full-pipeline memory benchmark with extensive annotation efforts for advancing the domain's development and is proposed with the proposed new definition User-conditioned Relational Agent Memory (URAM).

World Embedding Benchmark
World Embedding Benchmark(世界嵌入基准)
arXiv:2610.03632 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出了 World Embedding Benchmark,包含来自 80 个族的 8,000 个受控仿真案例,涵盖流体力学、固体力学、动力学以及光学与电磁学,旨在强调需要联合评估物理一致性与物理属性可恢复性,并展示物理表示对改进视频生成的效用。The World Embedding Benchmark is introduced, comprising 8,000 controlled simulation cases from 80 families spanning fluid mechanics, solid mechanics, dynamics, and optics&electromagnetism, to highlight the need to evaluate physical alignment and property recoverability jointly and demonstrate the utility of physical representations for improving video generation.

4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes
4DCodeBench:动态场景逆向图形中的 Agent 基准测试
arXiv:2610.03715 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 4DCodeBench,一个面向 4D 逆向图形(通过代码生成)的基准,Agent 以可执行图形程序的形式从视频重建动态场景,结果表明强大的静态重建能力尚未转化为对复杂动态的可靠重建。4DCodeBench is introduced, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs, finding that strong static reconstruction capabilities do not yet translate into reliable reconstruction of complex dynamics.

Arm-wise Compositional Generalization in Dual-Arm Vision-Language-Action Models
双臂视觉-语言-动作模型中的按臂组合泛化
arXiv:2610.06184 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 ACG-Bench,一个面向逐臂组合泛化(arm-wise Compositional Generalization)的基准,为超越固定训练流程的双臂策略设计提供实证指导,考察了 arm-token grouping、技能专属 LoRA 适配器(SkillLoRA)以及 arm-wise attention(AWA),凸显了技能条件化参数与注意力结构的互补性。This work introduces ACG-Bench, a benchmark for arm-wise Compositional Generalization that provides empirical guidance for designing dual-arm policies that generalize beyond fixed training routines, and examines arm-token grouping, skill-specific LoRA adapters (SkillLoRA), and arm-wise attention (AWA), highlighting the complementarity of skill-conditioned parameters and attention structure.

The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models
缺失的原语:诊断与修复大语言模型中的数学推理
arXiv:2610.02191 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

引入 Mathematical Primitive 概念来探查结构性数学理解,并提出 \hlei{},一个沿发现(Discovery)、生成(Generation)、消化(Digestion)和执行(Execution)四个维度评估数学推理的新型基准。This paper introduces the notion of Mathematical Primitive to probe structural mathematical understanding and proposes \hlei{}, a novel benchmark that evaluates mathematical reasoning along four distinct dimensions: Discovery, Generation, Digestion, and Execution.

UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents
UndoBench:在使用工具的 AI Agent 中分离任务能力与恢复能力
arXiv:2610.05622 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 UndoBench,一个覆盖 8 个企业领域、36 个基础工作流与 36 个故障场景的基准测试,通过相同种子下的反事实配对试验以及链路级效应历史与环境状态预言机,将任务能力与恢复能力解耦。UndoBench is introduced, a benchmark spanning 36 base workflows and 36 fault scenarios across 8 enterprise domains, decoupling task competence from recovery capability via counterfactual paired trials under identical seeds alongside wire-level effect-history and environment-state oracles.

Technical Report on the Turba Fertilizer Machine Learning Stack in Morocco
摩洛哥 Turba 施肥机器学习技术报告
arXiv:2610.05949 评测基准 评测集 OA · 绿色 被引 0 · OpenAlex

场地特异性施肥推荐系统会根据地点、土壤属性、作物类型和生产目标调整养分建议,但当推荐功能主要通过交互式界面访问、输出未做版本管理、训练后的近似模型无法独立加载或基准测试时,其科学复用性受到限制。本技术报告介绍 Turba 施肥机器学习技术栈——面向摩洛哥场地特异性施肥推荐的可复现三层开源实现,其中 turba-client 提供对公开……Site-specific fertilizer recommendation systems adapt nutrient advice to location, soil properties, crop type, and production targets, but scientific reuse is constrained when recommendation functions remain accessible mainly through interactive interfaces, outputs are not versioned, and trained approximations cannot be independently loaded or benchmarked. This technical report presents the Turba fertilizer machine learning stack, a three-layer open-source implementation for reproducible site-specific fertilizer recommendation in Morocco. turba-client provides programmatic access to publicly a

RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies
RoboDojo:面向通用机器人操作策略综合评估的仿真-真机统一基准
arXiv:2607.04434 评测基准 评测集 OA · 绿色 被引 44 · S2

提出 RoboDojo,一个面向通用机器人操作策略综合评估的仿真-真机统一基准,将 30 种策略集成到 XPolicyLab 并在 RoboDojo 上进行评测,建立了公开的排行榜与系统性的策略性能分析。RoboDojo is introduced, a unified sim-and-real benchmark for comprehensive evaluation of generalist robot manipulation policies that integrates 30 policies into XPolicyLab and evaluates them on RoboDojo, establishing a public leaderboard and systematic analysis of current policy performance.

PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages
PluraMath:将数学推理评估拓展至丰富资源语言之外
arXiv:2607.05992 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

确认了丰富资源语言与代表性不足语言之间在数学推理性能上存在持续差距,性能更优主要与更强的指令遵循能力相关,并提出了完全开源的数据集、数据采集流程与评估框架。A persistent gap in mathematical reasoning performance between high-resource and underrepresented languages is confirmed, with stronger results largely associated with better instruction-following ability, and a fully open-source dataset, data acquisition pipeline, and evaluation framework is introduced.

HETERQA: Benchmarking Record Retrieval over Multiple Heterogeneous Sources
HETERQA:跨多个异构来源的记录检索基准
arXiv:2607.03028 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出 HETERQA,一个包含 857 个 QA 对的综合性基准,涵盖五个异构来源的记录检索,并表明 HETERQA 为异构来源下的记录检索提供了有效的测试平台,为未来检索方法留下了显著空间。This work introduces HETERQA, a comprehensive benchmark with 857 QA pairs for record retrieval over five heterogeneous sources and indicates that HETERQA provides an effective testbed for record retrieval over heterogeneous sources and leaves substantial room for future retrieval methods.

Taste-aware music retrieval from audio embeddings
基于音频嵌入的品味感知音乐检索
arXiv:2607.03296 RAG 检索增强 评测集 OA · 绿色 被引 2 · S2

将预测的味觉空间作为基于内容的检索索引,对 309 项条目池的排序比 CLAP-text 基线(处于随机水平)忠实得多;ridge probes 与 audio-bandstop knockout 在已记载的声-味对应关系上读出了最强表征。Operationalised as a content-based retrieval index, the predicted taste space ranks a 309-item pool far more faithfully than a CLAP-text baseline, which sits at chance; ridge probes and an audio-bandstop knockout read the strongest representations against documented sound-taste correspondences.

DataComp-VLM: Improved Open Datasets for Vision-Language Models
DataComp-VLM:面向视觉-语言模型的改进开源数据集
arXiv:2606.28551 多模态 评测集 MPG.PuRe (Max Planck Society) OA · 绿色 被引 1 · S2

数据混合(而非过滤)是构建高质量训练数据集的关键:以指令型数据为主的混合在扩展时优于以描述型数据为主的混合,且规模越大优势越明显。It is found that data mixing, not filtering, is key to a high-quality training dataset: instruction-heavy mixtures scale better than caption-heavy ones, with gains widening at larger scales.

AGVBench: A Reliability-Oriented Benchmark of Data Augmentation for Vein Recognition
AGVBench:面向可靠性的静脉识别数据增强基准
arXiv:2607.02271 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

静脉识别是一种安全生物特征技术,常受限于标注数据稀缺与成像差异;而面向自然图像设计的增强策略可能破坏其关键的细粒度拓扑与纹理。本文提出 AGVBench,在 5 个公开掌/指静脉数据集、7 种骨干网络(含经典 CNN、视觉 Transformer 及静脉专用模型)上评测 30 种代表性增强策略。结果显示,多图混合类方法(如 MixUp、PuzzleMix……Vein recognition is a secure biometric technology often constrained by limited annotated data and imaging variations. While data augmentation mitigates this, strategies designed for natural images may disrupt the fine-grained topology and textures essential for identity discrimination. We present AGVBench, which evaluates 30 representative augmentation strategies on five public palm- and finger-vein datasets with seven backbone architectures, covering classic CNNs, vision transformers, and vein-specific recognition models. Our results show that multi-image mixing methods (e.g., MixUp, PuzzleMi

EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments
EvoPolicyGym:在交互式环境中评估自主策略演化
arXiv:2607.02440 评测基准 评测集 OA · 绿色 被引 2 · S2

提出"自主策略演化"评估范式:在固定交互预算下,由 harness-model Agent 反复编辑可执行策略系统;并在 EvoPolicyGym 中实例化,该基准基于一组紧凑型交互式 RL 环境构建,用于评测 Agent 如何迭代改进已探索策略。This work introduces Autonomous Policy Evolution, a controlled evaluation setting in which a harness-model agent repeatedly edits an executable policy system under a fixed interaction budget, and instantiates this setting in EvoPolicyGym, a benchmark built from compact interactive RL environments that evaluates how agents iteratively improve explored policies.

Discrete Diffusion Language Models for Interactive Radiology Report Drafting
用于交互式放射学报告起草的离散扩散语言模型
arXiv:2607.01436 多模态 评测集 OA · 绿色 被引 1 · S2

本文适配了一款专家混合扩散语言模型 DiffusionGemma-26B,并在医学视觉问答数据集上,使用相同的 LoRA 配置将其与同规模的自回归模型 Gemma-4-26B 进行基准对比,由对冗长度鲁棒的 LLM 裁判打分。This work adapts a mixture-of-experts diffusion language model, DiffusionGemma-26B, and benchmark it against its same-size AR sibling Gemma-4-26B under an identical LoRA recipe on medical visual question answering datasets, scored by a verbosity-robust LLM judge.

Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning
打破失败级联:面向医学多模态推理的步骤感知强化学习
arXiv:2606.31825 多模态 评测集 OA · 绿色 被引 2 · S2

在四个多模态 LLM backbone 上,MRPO 始终优于标准 GRPO 与一项最新 RL baseline,并在 Qwen3-VL-8B-Thinking 上以 4.59 分超越规模远大于它的医学 MLLM(如 HuatuoGPT-Vision-34B)。Across four multimodal LLM backbones, MRPO consistently outperforms standard GRPO and a recent RL baseline, and on Qwen3-VL-8B-Thinking even surpasses substantially larger medical MLLMs such as HuatuoGPT-Vision-34B by 4.59 points.

Beyond IID: How General Are Tabular Foundation Models, Really?
超越 IID:表格基础模型的泛化能力究竟如何?
arXiv:2606.30410 评测基准 评测集 OA · 绿色 被引 11 · S2

BeyondArena 是首个面向表格数据的统一整体基准,支持多种任务类型(IID、时序、分组),覆盖样本量与特征维度的不同尺度,并涵盖来自广泛学科的多样化特征类型。BeyondArena is the first unified holistic benchmark for tabular data that supports diverse task types (IID, temporal, grouped), across sample size and feature dimensionality scales, with diverse feature types from a broad range of disciplines.

PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception
PerceptionRubrics:将多模态评估校准至人类感知
arXiv:2606.28322 评测基准 评测集 OA · 绿色 被引 2 · S2

提出 PerceptionRubrics——一个基于评分量表的评估框架,旨在弥合饱和的基准分数与真实场景脆弱性之间的差距,并验证了严格的感知保真是可靠生成的前提。PerceptionRubrics is introduced, a rubric-based evaluation framework that addresses the gap between saturated benchmark scores and real-world brittleness, validating that strict perceptual fidelity is the prerequisite for reliable generation.