研究库 论文知识库
Papers · organized/paper_cards

论文

224 张论文卡片 · 评测基准

开放获取 全部 绿色 · 1640
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
MILO:通过编排式多 Agent 进化实现自动化 Harness 发现
arXiv:2609.38349 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 MILO(Meta-evolutionary Island Orchestration),一个共同演化 agent harness 及其发现策略的框架,使用前沿模型(Opus 4.8)与开源权重模型(gpt-oss-120b)超越了八个 SOTA harness 与六种搜索方法。This work introduces MILO (Meta-evolutionary Island Orchestration), a framework that co-evolves agent harnesses and the strategy used to discover them and outperforms eight state-of-the-art harnesses and six search methods using frontier (Opus 4.8) and open-weight (gpt-oss-120b) models.

Aligning One-Step Generative Models with Reward-Weighted Transport Distillation
用奖励加权传输蒸馏对齐单步生成模型
arXiv:2609.30840 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

理论分析表明 RWTD 的不动点分布在参考策略的 off-policy 奖励倾斜与当前模型的 on-policy 倾斜之间插值,提供了一种在奖励适配与保留先验知识之间取得平衡的原则性方法。Theoretical analysis shows that the fixed-point distributions of RWTD interpolate between off-policy reward tilting of the reference and on-policy tilting of the current model, providing a principled approach to balancing reward adaptation with retention of prior knowledge.

Benchmarking and Enhancing Skill-Level Memory for Partially Observable Robotic Manipulation
在部分可观测机器人操作任务上对技能级记忆的基准评测与增强
arXiv:2609.38886 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 HIDE,一个用于在部分可观测条件下评估操作记忆的基准,并提出 SEEK 框架,结合三种互补的记忆机制来保留历史证据并追踪执行状态。This work introduces HIDE, a benchmark for evaluating manipulation memory under partial observability, and proposes $SEEK, a framework combining three complementary memory mechanisms to retain historical evidence and track execution state, and proposes a framework combining three complementary memory mechanisms to retain historical evidence and track execution state.

Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs
泛化即稳定性,而非准确率:LLM 的多轴评估
arXiv:2610.01428 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出稳定性感知的泛化目标(SAGO),一个用于衡量模型在同一输入上面对不同扰动和基准时行为变化程度的评估框架,涵盖生成一致性、内部激活、置信度以及响应镜像等多个维度的变异性。The Stability-Aware Generalization Objective (SAGO) is introduced, a framework that measures how much model behavior changes for the same input under different variations and benchmarks, capturing variability across several dimensions including generation consistency, internal activations, confidence, and response mirroring.

PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop
PhysVista:通过感知-推理-评估闭环评测 VLM 的物理智能
arXiv:2610.00559 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

PhysVista 是一个旨在通过借鉴人类“感知-推理-评估”过程的认知闭环框架来评测 VLM 物理智能的基准,揭示了视觉识别与真实物理理解之间持续存在的差距,并为面向物理基础的多模态智能设计提供了更具原则性的方向。PhysVista is a benchmark designed to evaluate physical intelligence in VLMs through a closed cognitive loop framework inspired by the human seeing-reasoning-assessment process, highlighting a persistent gap between visual recognition and genuine physical understanding, and pointing toward more principled designs for physically grounded multimodal intelligence.

OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories
OpenTumorBoard:多学科肿瘤委员会讨论轨迹的真实世界基准
arXiv:2609.32810 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出了 OpenTumorBoard,一个包含 611 个患者案例、19,157 轮讨论、涵盖十个专科角色的基准,源自 YouTube 上 12,534 分钟公开肿瘤委员会会议录音的转写,并发布了自动化整理流水线,以支持多学科、个性化癌症决策中 LLM 的开发与评估。OpenTumorBoard is introduced, a benchmark with 611 patient cases and 19,157 discussion turns across ten specialist roles, transcribed from 12,534 minutes of publicly available tumor board recordings on YouTube, and automated curation pipeline is released to support the development and evaluation of LLMs for multidisciplinary, personalized cancer decision-making.

"I think this is the most disruptive technology": Exploring Sentiments of ChatGPT Early Adopters using Twitter Data
"我认为这是最具颠覆性的技术":基于 Twitter 数据探索 ChatGPT 早期采用者的情感
arXiv:2212.05856 评测基准 应用落地 OA · 绿色 被引 284 · S2

基于 10,732 条早期 ChatGPT 用户推文的混合方法研究,对每个主题进行深入定性情感分析,结果显示大多数早期采用者在软件开发颠覆性、娱乐与创意发挥等主题上表达了压倒性的积极情感。A mixed-method study using 10,732 tweets from early ChatGPT users to conduct an in-depth qualitative sentiment analysis of each topic, showing that the majority of the early adopters have expressed overwhelmingly positive sentiments related to topics such as Disruptions to software development, Entertainment and exercising creativity.

SimuVerity: Benchmarking Agents for Engineering-Grade Simulink Model Generation
SimuVerity:面向工程级 Simulink 模型生成的智能体基准
arXiv:2610.02304 评测基准 评测集 被引 0 · S2

SimuVerity 为评估智能体在可执行 Simulink 模型生成中的工程能力与诊断失败提供了系统性基础,并表明结构相似性是衡量工程性能的一个糟糕代理指标。SimuVerity provides a systematic basis for assessing agents' engineering capabilities and diagnosing failures in executable Simulink model generation, and shows that structural similarity is a poor proxy for engineering performance.

LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models
LexReward:面向法律语言模型的分类驱动奖励框架
arXiv:2609.39071 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

实验表明,基于评分细则的奖励能可靠地区分不同质量的法律回答,且在该偏好数据上进行 DPO 训练在所有三个维度上均提升了性能;逐维度分析进一步支持了所提分类法与奖励构建的有效性。Experiments show that the rubric-based rewards reliably distinguish legal responses of different quality and that DPO training on the preference data improves performance across all three dimensions, and Dimension-wise analyses further support the effectiveness of the proposed taxonomy and reward construction.

Multilingual GSM-Symbolic: What determines capability transfer across languages?
多语言 GSM-Symbolic:跨语言能力迁移由什么决定?
arXiv:2610.03367 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Multilingual GSM-Symbolic,一个可扩展的多语言数学数据集,包含 30,000 对题项匹配的问答对,覆盖 15 种语言;并将能力差异的最大决定因素量化为模型规模、语言资源水平、推理能力与类型学距离。This work introduces Multilingual GSM-Symbolic, an extensible multilingual mathematical dataset covering 30,000 item-matched question-answer pairs and spanning 15 languages and quantifies the largest determinants of capability as model size, language resource level, reasoning, and typological distance.

4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes
4DCodeBench:动态场景逆向图形中的 Agent 基准测试
arXiv:2610.03715 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 4DCodeBench,一个面向 4D 逆向图形(通过代码生成)的基准,Agent 以可执行图形程序的形式从视频重建动态场景,结果表明强大的静态重建能力尚未转化为对复杂动态的可靠重建。4DCodeBench is introduced, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs, finding that strong static reconstruction capabilities do not yet translate into reliable reconstruction of complex dynamics.

Tail-Influence Sampling for CVaR Policy Evaluation
面向 CVaR 策略评估的尾部影响采样
arXiv:2609.38096 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

为每个可查询的条件分布推导出一个尾部影响度量,聚合其不确定性对每一次 Bellman 复用中 CVaR 的影响,并刻画了尾部最优与均值最优分配重合的精确网格(exact-grid)机制。A tail influence for each queryable conditional law that aggregates how its uncertainty affects CVaR across every Bellman reuse is derived, and an exact-grid regime in which tail- and mean-optimal allocations coincide is characterized.

The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models
缺失的原语:诊断与修复大语言模型中的数学推理
arXiv:2610.02191 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

引入 Mathematical Primitive 概念来探查结构性数学理解,并提出 \hlei{},一个沿发现(Discovery)、生成(Generation)、消化(Digestion)和执行(Execution)四个维度评估数学推理的新型基准。This paper introduces the notion of Mathematical Primitive to probe structural mathematical understanding and proposes \hlei{}, a novel benchmark that evaluates mathematical reasoning along four distinct dimensions: Discovery, Generation, Digestion, and Execution.

EMHO: EMbodied Agent Harness Optimization via Experience Traces
EMHO:基于经验轨迹的具身 Agent Harness 优化
arXiv:2610.08432 评测基准 方法 被引 0 · S2

定性分析表明,EMHO 超越了从失败与无效动作中恢复的范畴,重塑了具身 agent 对环境的解释与交互方式,并在 EmbodiedBench 的导航和操作任务上进行了评估。Qualitative analysis shows that EMHO goes beyond recovering from failures and unproductive actions to reshape how the embodied agent interprets and interacts with its environment, and is evaluated on EmbodiedBench across navigation and manipulation tasks.

Technical Report on the Turba Fertilizer Machine Learning Stack in Morocco
摩洛哥 Turba 施肥机器学习技术报告
arXiv:2610.05949 评测基准 评测集 OA · 绿色 被引 0 · OpenAlex

场地特异性施肥推荐系统会根据地点、土壤属性、作物类型和生产目标调整养分建议,但当推荐功能主要通过交互式界面访问、输出未做版本管理、训练后的近似模型无法独立加载或基准测试时,其科学复用性受到限制。本技术报告介绍 Turba 施肥机器学习技术栈——面向摩洛哥场地特异性施肥推荐的可复现三层开源实现,其中 turba-client 提供对公开……Site-specific fertilizer recommendation systems adapt nutrient advice to location, soil properties, crop type, and production targets, but scientific reuse is constrained when recommendation functions remain accessible mainly through interactive interfaces, outputs are not versioned, and trained approximations cannot be independently loaded or benchmarked. This technical report presents the Turba fertilizer machine learning stack, a three-layer open-source implementation for reproducible site-specific fertilizer recommendation in Morocco. turba-client provides programmatic access to publicly a

RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies
RoboDojo:面向通用机器人操作策略综合评估的仿真-真机统一基准
arXiv:2607.04434 评测基准 评测集 OA · 绿色 被引 44 · S2

提出 RoboDojo,一个面向通用机器人操作策略综合评估的仿真-真机统一基准,将 30 种策略集成到 XPolicyLab 并在 RoboDojo 上进行评测,建立了公开的排行榜与系统性的策略性能分析。RoboDojo is introduced, a unified sim-and-real benchmark for comprehensive evaluation of generalist robot manipulation policies that integrates 30 policies into XPolicyLab and evaluates them on RoboDojo, establishing a public leaderboard and systematic analysis of current policy performance.

PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages
PluraMath:将数学推理评估拓展至丰富资源语言之外
arXiv:2607.05992 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

确认了丰富资源语言与代表性不足语言之间在数学推理性能上存在持续差距,性能更优主要与更强的指令遵循能力相关,并提出了完全开源的数据集、数据采集流程与评估框架。A persistent gap in mathematical reasoning performance between high-resource and underrepresented languages is confirmed, with stronger results largely associated with better instruction-following ability, and a fully open-source dataset, data acquisition pipeline, and evaluation framework is introduced.

HETERQA: Benchmarking Record Retrieval over Multiple Heterogeneous Sources
HETERQA:跨多个异构来源的记录检索基准
arXiv:2607.03028 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出 HETERQA,一个包含 857 个 QA 对的综合性基准,涵盖五个异构来源的记录检索,并表明 HETERQA 为异构来源下的记录检索提供了有效的测试平台,为未来检索方法留下了显著空间。This work introduces HETERQA, a comprehensive benchmark with 857 QA pairs for record retrieval over five heterogeneous sources and indicates that HETERQA provides an effective testbed for record retrieval over heterogeneous sources and leaves substantial room for future retrieval methods.

Measuring the Gap Between Human and LLM Research Ideas
衡量人类与 LLM 研究思路之间的差距
arXiv:2607.01233 评测基准 方法 OA · 绿色 被引 9 · S2

结果表明,强大的 LLM 能够产出一系列合理思路,但其范围仍比人类研究品味更窄,并存在系统性偏移。It is suggested that strong LLMs can produce a range of reasonable ideas, but that range remains narrower than, and systematically shifted relative to, human research taste.

AGVBench: A Reliability-Oriented Benchmark of Data Augmentation for Vein Recognition
AGVBench:面向可靠性的静脉识别数据增强基准
arXiv:2607.02271 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

静脉识别是一种安全生物特征技术,常受限于标注数据稀缺与成像差异;而面向自然图像设计的增强策略可能破坏其关键的细粒度拓扑与纹理。本文提出 AGVBench,在 5 个公开掌/指静脉数据集、7 种骨干网络(含经典 CNN、视觉 Transformer 及静脉专用模型)上评测 30 种代表性增强策略。结果显示,多图混合类方法(如 MixUp、PuzzleMix……Vein recognition is a secure biometric technology often constrained by limited annotated data and imaging variations. While data augmentation mitigates this, strategies designed for natural images may disrupt the fine-grained topology and textures essential for identity discrimination. We present AGVBench, which evaluates 30 representative augmentation strategies on five public palm- and finger-vein datasets with seven backbone architectures, covering classic CNNs, vision transformers, and vein-specific recognition models. Our results show that multi-image mixing methods (e.g., MixUp, PuzzleMi

EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments
EvoPolicyGym:在交互式环境中评估自主策略演化
arXiv:2607.02440 评测基准 评测集 OA · 绿色 被引 2 · S2

提出"自主策略演化"评估范式:在固定交互预算下,由 harness-model Agent 反复编辑可执行策略系统;并在 EvoPolicyGym 中实例化,该基准基于一组紧凑型交互式 RL 环境构建,用于评测 Agent 如何迭代改进已探索策略。This work introduces Autonomous Policy Evolution, a controlled evaluation setting in which a harness-model agent repeatedly edits an executable policy system under a fixed interaction budget, and instantiates this setting in EvoPolicyGym, a benchmark built from compact interactive RL environments that evaluates how agents iteratively improve explored policies.

Cross-Domain Generalization Failure in Lightweight Intrusion Detection Models for IIoT Networks
面向 IIoT 网络的轻量级入侵检测模型的跨域泛化失效
arXiv:2607.00553 评测基准 应用落地 OA · 绿色 被引 2 · S2

应在真实类别分布下使用跨网络评估来判断部署就绪度,而非仅依赖域内准确率。Deployment readiness should be assessed using cross-network evaluation under realistic class distributions, rather than within-domain accuracy alone, to suggest deployment readiness should be assessed using cross-network evaluation under realistic class distributions, rather than within-domain accuracy alone.

Beyond IID: How General Are Tabular Foundation Models, Really?
超越 IID:表格基础模型的泛化能力究竟如何?
arXiv:2606.30410 评测基准 评测集 OA · 绿色 被引 11 · S2

BeyondArena 是首个面向表格数据的统一整体基准,支持多种任务类型(IID、时序、分组),覆盖样本量与特征维度的不同尺度,并涵盖来自广泛学科的多样化特征类型。BeyondArena is the first unified holistic benchmark for tabular data that supports diverse task types (IID, temporal, grouped), across sample size and feature dimensionality scales, with diverse feature types from a broad range of disciplines.

PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception
PerceptionRubrics:将多模态评估校准至人类感知
arXiv:2606.28322 评测基准 评测集 OA · 绿色 被引 2 · S2

提出 PerceptionRubrics——一个基于评分量表的评估框架,旨在弥合饱和的基准分数与真实场景脆弱性之间的差距,并验证了严格的感知保真是可靠生成的前提。PerceptionRubrics is introduced, a rubric-based evaluation framework that addresses the gap between saturated benchmark scores and real-world brittleness, validating that strict perceptual fidelity is the prerequisite for reliable generation.

How Good Can Linear Models Be for Time-Series Forecasting?
线性模型在时间序列预测中能做到多好?
arXiv:2606.27282 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

所得到的模型在大多数数据集-预测步长组合上优于先前的线性预测器,并在八个基准中的六个上超越 Transformer、MLP 和 CNN 基线;同时它还可作为对数据本身的诊断工具,揭示那些被更大模型默默吸收进其学习参数中的结构。The resulting models beat prior linear forecasters on most dataset-horizon entries and exceed Transformer, MLP, and CNN baselines on six of eight benchmarks, and serve as a diagnostic on the data itself, revealing structures that larger models absorb silently into their learned parameters.

Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments
穿越 gauntlet:重新评估 Agent 在熟悉环境之外的能力
arXiv:2606.14397 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 GauntletBench,一个用于评估 Agent 在挑战性场景中泛化能力的 Web 基准,聚焦于三种被低估的能力(时间感知、图形理解与 3D 推理),揭示了当前 Agent 能力与复杂真实场景所需能力之间的巨大差距。GauntletBench, a web-based benchmark for evaluating agent generalisation in challenging scenarios, focusing on three underexplored capabilities (temporal perception, graphical understanding, and 3D reasoning), is introduced, revealing the substantial gap between current agent capabilities and those required for complex real-world scenarios.

The Galaxy's Guide to the Tokenizer: A Benchmark for Scientific Foundation Models
The Galaxy's Guide to the Tokenizer: A Benchmark for Scientific Foundation Models
arXiv:2606.25610 评测基准 综述 OA · 绿色 被引 1 · S2

研究发现重建质量与表示质量是解耦的;在所考察的任务中,没有任何单一方法能够在所有任务上稳定取得最优表现。It is found that reconstruction and representation quality are decoupled, and no single method consistently performs best across the tasks considered here and no single method consistently performs best across the tasks considered here.

Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation
管理 LLM Agent 中的程序性记忆:控制、适应与评估
arXiv:2606.23127 评测基准 评测集 OA · 绿色 被引 9 · S2

一个包含 382 个真实企业任务、覆盖 6 种专业角色和 22 项程序性技能的基准,用于评估技能在任务、角色和模型骨干间的迁移能力,发现部分技能可在任务和模型间广泛泛化,而另一些则专化为角色特定工作流,在迁移时失去效力。A benchmark of 382 realistic enterprise tasks spanning six professional roles and 22 procedural skills, designed to evaluate how skills transfer across tasks, roles, and model backbones finds that some skills generalize broadly across tasks and models, whereas others become specialized to role-specific workflows and lose effectiveness under transfer.

PowerAgentBench-SS: A Benchmark for Agentic AI in Power System Steady-State Studies
PowerAgentBench-SS:面向电力系统稳态研究的 Agentic AI 基准
arXiv:2606.18789 评测基准 评测集 OA · 绿色 被引 4 · S2

结果表明仅评估求解器或仅评估答案是不足的:Agent 的差异不仅体现在发现关键 contingency 上,还体现在验证预算的使用、显式提交、类型强制、重复验证、基于证据的报告以及缓解行为等方面。The results show why solver-only or answer-only evaluation is insufficient: agents are distinguished not only by top-contingency discovery, but also by validation-budget use, explicit submission, type coercions, duplicate validations, evidence-backed reporting, and mitigation behavior.

AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification
AdvancedMathBench:一个面向高等数学证明生成与验证的基准测试套件
arXiv:2607.11849 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本工作提出 AdvancedMathBench,用于评估 LLM 在高级数学证明上的推理能力;同时引入 VerifierBench,包含 888 条由模型生成的证明轨迹及专家真值,用于评估模型能否正确判断证明有效性并给出合理的验证依据。This work introduces AdvancedMathBench, a benchmark suite designed to evaluate the reasoning capabilities of LLMs on advanced mathematical proofs, and introduces VerifierBench, consisting of 888 model-generated proof trajectories paired with expert ground truth, to evaluate whether models can correctly judge proof validity and provide sound verification rationales.

MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning
MET:理论支撑且文化感知的多语言道德推理
arXiv:2607.11736 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 MCLASH,一个多语言道德决策 benchmark,用于捕捉跨语言的文化情境化道德直觉和社会规范;并提出 MET(Multilingual Ethics with Theory-grounded reasoning),一种基于心理学与哲学专家策划的理论依据的两步提示方法。This work introduces MCLASH, a multilingual moral decision-making benchmark to capture culturally situated moral intuitions and social norms across languages, and proposes MET (Multilingual Ethics with Theory-grounded reasoning), a two-step prompting method built on expert-curated, theory-based grounds drawn from psychology and philosophy.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models
Blind-Spots-Bench:评估多模态模型中的盲点
arXiv:2607.08317 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

开发了自动化评分流水线,用于评估多种模型,包括开源权重模型、闭源语言模型、视觉语言模型和图像生成模型,结果显示没有单一模型在所有任务类型上占优,且某些任务对所有评估模型仍具挑战性。An automated grading pipeline is developed to evaluate a wide range of models, including open-weight and closed-source language, vision-language, and image-generation models, and shows that no single model dominates across all task types, and that some tasks remain challenging for all evaluated models.

Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms
深度强化学习评估与设计范式的原理性分析
arXiv:2607.07769 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

介绍其潜在原因的理论基础,阐明强化学习算法的渐近性能在性能排名与数据规模之间不存在单调关系。The theoretical foundations of the underlying causes outlining that the asymptotic performance of reinforcement learning algorithms does not have a monotone relationship between performance rankings and data-regimes are introduced.

LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes
LakeQuest:面向跨数据湖根植式问答的三领域基准
arXiv:2607.12310 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 LakeQuest,一个经人工校验的基准,包含 9,846 个 QA 对,用于在真实数据湖场景下评估端到端检索与综合 pipeline,并揭示现代问答系统中的关键失效模式。LakeQuest is introduced, a human-validated benchmark of 9,846 QA pairs designed to evaluate the end-to-end retrieve-and-synthesize pipeline over realistic data lakes and exposes critical failure modes in modern QA systems.

Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists
LLMs 是否准备好进行科学发现?面向 AI 科学家的能力导向基准
arXiv:2607.11079 评测基准 评测集 OA · 绿色 被引 1 · S2

提出 SDABench,一个围绕六项能力(描述性、探索性、推断性、推断性、预测性、因果性、机制性)跨五大领域(生物、化学、环境、地理、物理)重新组织评估的基准。SDABench is introduced, a benchmark that reorganizes evaluation around six capabilities (descriptive, exploratory, inferential, inferential, predictive, causal, and mechanistic) across five domains (Biology, Chemistry, Environment, Geography, Physics).

SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding
SynthDocBench:面向长上下文视觉文档理解的受控基准
arXiv:2607.10400 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

视觉语言模型(VLMs)在 DocVQA、ChartQA、MMLongBench-Doc 等视觉文档理解基准上表现强劲。但真实文档融合长度、布局复杂度、模态、问题难度等多因素,难以将模型失败归因于具体原因。我们提出 SynthDocBench,一个全合成的长上下文视觉文档理解基准,系统性控制文档长度、布局结构、模态构成与问题类型等因子。该基准的构建……Vision language models (VLMs) have achieved strong performance on visual document understanding benchmarks such as DocVQA, ChartQA, and MMLongBench-Doc. However, real-world documents combine multiple factors such as length, layout complexity, modality, and question difficulty, which makes it difficult to attribute model failures to specific causes. We introduce SynthDocBench, a fully synthetic benchmark for long-context visual document understanding that systematically controls factors including document length, layout structure, modality composition, and question type. The benchmark is constr