研究库 论文知识库
Papers · organized/paper_cards

论文

217 张论文卡片 · 评测基准 · OA 绿色

开放获取 全部 绿色 · 1640
ContextBias: Controlled Evaluation of Bias Persistence Under Context Shift in Text-to-Image Models
ContextBias:受控评估文本到图像模型在上下文偏移下的偏见持续性
arXiv:2608.29847 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

评估四个 SOTA 模型发现,将角色置于语义无关的上下文中并不会抑制该角色的关联属性;相反,跨角色的属性集中度会上升(合并 BI $+0.047$)。Evaluating four state-of-the-art models finds that placing a role in a semantically unrelated context does not suppress role-linked attributes; instead, cross-role attribute concentration increases (pooled BI $+0.047$).

EvoGenUI-Bench: Evaluating LLMs as Multi-Turn Generative UI Assistants
EvoGenUI-Bench:评估 LLM 作为多轮生成式 UI 助手
arXiv:2608.29387 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 EvoGenUI-Bench,一个面向多轮界面维护的基准,包含 150 个五轮任务,共计 750 轮,覆盖三种场景:信息呈现、可执行交互和工具驱动的外部状态。EvoGenUI-Bench is introduced, a benchmark for multi-turn interface maintenance comprising 150 five-turn tasks and 750 turns across three scenarios: information presentation, executable interaction, and tool-grounded external state.

Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement
Harness-of-Harness:具备持续改进能力的多日自主软件开发
arXiv:2609.01481 评测基准 方法 OA · 绿色 被引 2 · S2

在持续多天、超过 70 轮迭代的部署中,HoH 自主开发出一款第一人称射击游戏,具备完整的主线剧情、完整实现的核心机制、可供人类游玩的体验、精美的画面与集成的音效。In a multi-day deployment with more than 70 iterations, HoH autonomously develops a first-person-shooter game, featuring a coherent storyline, fully implemented core mechanics, human-playable experience, polished visuals and integrated audio.

DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation
DramaChain Bench:面向短剧生成的端到端基准
arXiv:2609.00646 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 DramaChain Bench,首个覆盖完整生产链各阶段的短剧基准,并验证最终剧集质量并非仅由视频生成决定。DramaChain Bench is presented, the first short-drama benchmark that evaluates every stage of the complete production chain, and confirms that final episode quality is not governed by video generation alone.

Does Imitation Learning Preserve Temporal Robustness in Dexterous Manipulation? An Expert-Learner Comparison Across Task Execution Speeds
模仿学习是否保留灵巧操作的时间鲁棒性?跨任务执行速度的专家-学习者对比
arXiv:2609.01453 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

任务名义成功率相同并不意味着跨执行速度保留专家性能;本文在相同任务条件、初始条件采样与加速倍率下对比了专家与学习者。Equal nominal task success does not imply preservation of expert performance across execution speeds, and an expert and learner under the same task conditions, initial-condition draws, and speedup factors is compared.

AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling
AgentJudgeBench:一个用于评估 LLM 法官在智能体工具调用方面表现的多难度基准
arXiv:2608.26623 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

首个系统性研究 LLM-as-a-judge 在工作流 DAG 上对 Agentic 工具调用评判可靠性的基准,区别于面向开放式文本或偏好的通用 LLM-as-a-judge 任务;揭示了当前 LLM judge 的根本局限,并给出面向 Agentic 系统可靠评估的实践指南。The first benchmark to systematically study LLM-as-a-judge reliability for agentic tool-calling over workflow DAGs, as distinct from the broader LLM-as-a-judge task of open-ended text or preference evaluation, exposes fundamental limitations of current LLM judges and yields practical guidelines for reliable evaluation in agentic systems.

HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
HarnessDev:LLM 能创建并演进自身的 Agent Harness 吗?
arXiv:2609.01437 评测基准 评测集 OA · 绿色 被引 5 · S2

研究发现,生成的 harness 在代码与搜索/研究任务上仍显著落后于成熟的人工参考方案,而在写作与机器学习实验任务上达到或超过所选参考方案,且执行成本差异巨大。It is found that generated harnesses remain substantially behind mature human-engineered references on code and on search and research, while matching or exceeding the selected references on writing and machine-learning experimentation, with large variation in execution cost.

Aspire: Can Models Self-Evolve from Vague Goals?
Aspire:模型能否从模糊目标中自我演进?
arXiv:2608.31111 评测基准 评测集 OA · 绿色 被引 1 · S2

本文提出 ASPIRE,面向模糊目标驱动自我演化的基准,表明模糊目标会将搜索资源导向目标解释阶段;并在涵盖 6 类目标的 520 题隐藏专家评测集上评估所得系统。This work introduces ASPIRE, a benchmark for vague-goal-driven self-evolution and shows that vague goals redirect search effort toward goal interpretation, and evaluates the resulting systems on a hidden, expert-authored set of 520 items spanning six goals.

S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?
S3Gym:LLM 能将自测试与自评判转化为自我提升吗?
arXiv:2608.31100 评测基准 评测集 OA · 绿色 被引 1 · S2

这些发现表明,仅识别成功动作并不足够;Agent 还需将反馈转化为可执行且可迁移的策略。本文给出统一框架以诊断该过程,并定位阻碍 Agent 将交互经验转化为可靠自我提升的瓶颈。These findings show that recognizing successful actions is insufficient; agents must also transform feedback into executable and transferable policies, and provide a unified framework for diagnosing this process and identifying the bottlenecks that prevent agents from translating interaction experience into reliable self-improvement.

WHALE: A Simple Recipe for Joint Harness-Weight Optimization
WHALE:联合 Harness-权重优化的简洁方案
arXiv:2609.00196 评测基准 方法 OA · 绿色 被引 8 · S2

本文提出 Weight-Harness Alternating LEarning (WHALE),一种简单的方法,交替进行两个阶段:先在当前 harness 下更新模型,再在更新后的模型下通过在线拒绝采样微调与 Meta-Harness 搜索更优的 harness。Weight-Harness Alternating LEarning (WHALE), a simple recipe that alternates two phases: updating the model under the current harness, then searching for a better harness under the updated model with online rejection-sampling fine-tuning and Meta-Harness, is proposed.

Small Language Models as Judges for Rubric-Based Reinforcement Learning
小语言模型作为基于评分标准的强化学习评判器
arXiv:2608.30005 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文研究较小的语言模型能否作为高效且可靠的 rubric 评分器,并比较了从小模型中提取逐项判断的三种方式:生成式判定、Yes/No Logprob 边际以及探针判别器。This work studies whether smaller language models can serve as efficient and reliable rubric-based judges, and compares three ways of extracting criterion-level judgments from small models: Generative verdicts, Yes/No Logprob margins, and Probe judges.

Last Translation Benchmark
终极翻译基准(Last Translation Benchmark)
arXiv:2609.04173 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 The Last Translation Benchmark,一组由人工编写并经同行评审的样本,可打破领先的机器翻译模型,并提出新评估方法:每个样本附带手工编写的验证规则,描述该样本上的具体失败案例,从而支持可靠且可操作的未来评估。The Last Translation Benchmark is introduced, a collection of human-authored and peer-reviewed examples that break leading machine translation models and a new evaluation approach: each example comes with handcrafted verification rules describing concrete failure cases on that example, therefore allowing reliable and actionable future evaluation.

Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization
优化之前先提问:交互式优化中的动态预形式化澄清。
arXiv:2609.05258 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出了 OR-Clarify,一个用于预表述澄清的 benchmark,并提出了 Interactive Optimization (InterOPT),这是一个两阶段框架,能识别未解决的、影响表述的关键缺口,并据此引导系统决定是提出下一个问题还是停止提问。This work introduces OR-Clarify, a benchmark for pre-formulation clarification and proposes Interactive Optimization (InterOPT), a two-stage framework that identifies unresolved formulation-critical gaps and uses them to guide whether to ask the next question or to stop.

When Models Edit Too Much: On the Fidelity of Minimal Code Edits
模型过度编辑时:最小代码编辑的保真度问题
arXiv:2609.04061 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作通过向参考解中注入受控的 AST 级 corruption,从 400 个 BigCodeBench 问题构建了一个评估框架,为每个修复任务赋予已知的最小 patch,并将 edit fidelity 定位为 code-repair 质量的一个独立维度,表明其可被度量与学习。This work constructs an evaluation framework from 400 BigCodeBench problems by injecting controlled AST-level corruptions into reference solutions, giving each repair task a known minimal patch, and positions edit fidelity as a distinct axis of code-repair quality and shows that it can be measured and learned.

Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs
思维链表象之下:LLM 中推理操作的机制化解读
arXiv:2609.04753 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作发现 reasoning operations 在 held-out 表征中是可分的,且 separability 在中间层达到峰值,并验证该结构无法由词汇或位置混淆因素解释。This work finds that reasoning operations are separable in held-out representations, with separability peaking in middle layers, and verify that this structure is not explained by lexical or positional confounds.

Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models
知道不该回答什么:视觉语言模型中的选择性不遵从
arXiv:2609.04720 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 KoNA,一个评估 VLM 选择性不遵从的基准,覆盖五类情形:False Premise、Visual Inaccessibility、Universal Unknown、Task Feasibility 与 Safety。实验结果显示,微调后的模型能够区分可回答部分与需要不遵从的部分,并以符合任务要求的方式作答。KoNA is introduced, a benchmark for evaluating selective non-compliance in VLMs across five categories: False Premise, Visual Inaccessibility, Universal Unknown, Task Feasibility, and Safety, and results suggest that the fine-tuned models can distinguish between answerable components and those requiring non-compliance and respond in a task-appropriate manner.

Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image Detection
借助 CLIP 与 DINO:面向可泛化深度伪造图像检测的不确定性感知级联融合网络
arXiv:2609.07670 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 UCF-Net,一个不确定性感知级联融合网络,利用 CLIP 语言对齐的语义先验与 DINO 自监督视觉结构先验,在域内与跨域评估中均取得最优平均 AUC。UCF-Net is proposed, an uncertainty-aware cascaded fusion network that harnesses CLIP's language-aligned semantic priors and DINO's self-supervised visual-structure priors and achieves the best mean AUC among the evaluated methods in both in-domain and cross-domain evaluations.

VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification
VDiff-Bench:面向细粒度图像差异识别的挑战性 benchmark
arXiv:2609.06245 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

VDiff-Bench 提供一个针对性诊断基准,用于评估 MLLM 的比较视觉理解能力,揭示标准单图视觉-语言任务无法捕捉的失败模式。VDiff-Bench provides a targeted diagnostic for evaluating comparative visual understanding in MLLMs, exposing failures that are not captured by standard single-image vision-language tasks.

EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?
EVOHARNESSBENCH:你的 Agent 能否跟上不断演进的 Harness?
arXiv:2609.04280 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 EvoHarnessBench,一个在工具、技能与 Agent 三个维度上对可控 harness 演化条件下的 Agent 进行评测的 benchmark,并将 harness 演化确立为一项独立挑战:Agent 需要在持续演化的 harness 下保持原有有效行为EvoHarnessBench is introduced, a benchmark for evaluating agents under controlled harness evolution across three axes (tools, skills, and agents), and establishes harness evolution as a distinct challenge for building agents that can keep pace with an evolving harness while preserving previously effective behavior.

AgentAudit: An Open, Extensible Framework for Full-Lifecycle Trust Evaluation of AI Agents
AgentAudit:一个面向 AI Agent 全生命周期信任评估的开放式可扩展框架
arXiv:2609.09875 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

AgentAudit 沿能力、接地性、安全与行为四个方面的十个维度评估完整执行轨迹——即指令完整性、规划器、记忆、工具选择、工具调用、工具正确性、对齐、工具忠实性、安全性与执行完整性——以精确定位导致观察到的失败的具体阶段。AgentAudit evaluates the entire execution trace across ten capability, grounding, security and behavioural dimensions, namely instruction integrity, planner, memory, tool selection, tool invocation, tool correctness, alignment, tool faithfulness, security and execution integrity, to pinpoint the exact stage responsible for an observed failure.

Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails
协同进化的 Harness 与模型:On-policy 修正助力弱模型在模仿失效处迎头赶上
arXiv:2609.09134 评测基准 方法 OA · 绿色 被引 3 · S2

开发一条 on-policy 专家修正流水线,由元层级 MLE Agent 自动化,在弱模型自身的 rollout 中定位失败回合,并请专家仅重写该回合,从而保留模型的规划风格,融合 Harness 进化与模型适配带来的收益。An on-policy expert-correction pipeline is developed, automated by a meta-level MLE agent, that localizes the failing turn in the weaker model's own rollout and asks the expert to rewrite only that turn, which preserves the model's planning style and combines the gains of harness evolution and model adaptation.

Diffs vs. Whole Files: An Empirical Comparison of Iterative Edit-Based and Direct Generation for Flutter/Dart Code Models
Diff 还是整文件:Flutter/Dart 代码模型中迭代编辑式生成与直接生成的实证比较
arXiv:2609.05779 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

识别出 diff 生成能够胜出的一种与架构无关的统一机制:它在短小且空间局部化的编辑上具有竞争力;其类别级优势恰好集中在作者数据集中平均编辑步数最低的两类任务——重构与错误处理/边界用例修复。A single, architecture-independent mechanism behind the conditions where diff-based generation does win is identified: it is competitive on short, spatially localized edits, and its category-level wins concentrate in exactly the two task categories - refactoring and error-handling/edge-case fixes - with the lowest mean edit-step count in the authors' dataset.

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
WearableQA:面向真实可穿戴数据的健康推理基准
arXiv:2609.05405 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

WearableQA 由 200 位真实用户的可穿戴时序数据、血液生物标志物和人口统计信息构建的 4,084 道十选一选择题组成,为评估 LLM 在真实可穿戴数据上的推理能力提供了现实且具诊断性的基准。WearableQA, a benchmark comprising 4,084 10-option multiple-choice questions constructed from the wearable time series, blood biomarkers, and demographics of 200 real users, provides a realistic and diagnostic benchmark for evaluating LLM reasoning over real-world wearable data.

SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents
SWE-Bench Pro Verified:面向软件工程 Agent 的可靠基准
arXiv:2609.08149 评测基准 评测集 OA · 绿色 被引 1 · S2

SWE-Bench Pro Verified 提供了一个更可信的基准来评估软件工程 Agent,其结合了反作弊保护(消除主要泄漏渠道而不干扰正常 Agent 功能)和任务精修(最小限度地修正有缺陷实例中的不一致性)。SWE-Bench Pro Verified offers a more trustworthy benchmark for assessing software engineering agents, which combines anti-hacking safeguards that eliminate major leakage channels without disrupting normal agent functionality, with task refinement that minimally corrects inconsistencies within flawed instances.

StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean
StochBench:面向 Lean 中随机过程的领域专用基准
arXiv:2609.09264 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文介绍 StochBench,一个基于 Lean 4 的基准,包含 450 道覆盖不同抽象层次的研究生随机过程问题,每道题配有其自然语言来源,更能代表领域特定的应用数学,同时对高级证明器仍具挑战。StochBench is introduced, a Lean 4 benchmark of 450 graduate stochastic-processes problems at varying abstraction levels, each paired with its natural-language source that better represents domain-specific applied mathematics while remaining challenging for advanced provers.

EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents
EvoSafeHarness:面向 Agent 安全防护的演进式模型与领域专属 Harness
arXiv:2609.05903 评测基准 应用落地 OA · 绿色 被引 2 · S2

EvoSafeHarness 是一个面向安全性的优化框架,针对目标领域中的冻结模型合成可部署的 harness,由模型行为、领域规范和新鲜上下文对抗审查共同引导,以拒绝基准特定规则。EvoSafeHarness is a safety-specific optimization framework that synthesizes a deployable harness for a frozen model in a target domain, guided by model behavior, domain specifications, and fresh-context adversarial review to reject benchmark-specific rules.

MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes
MetroLLM-Bench:将语言模型作为交通售票机运行时进行评估
arXiv:2609.10016 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

引入 MetroLLM-Bench,一个包含 955 个用例的基准,用于测试语言模型作为交通信息亭策略层的能力,并评估了来自六个厂商的 26 个模型,其中 23 个被排名。This work introduces MetroLLM-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk, and evaluates twenty-six models from six vendors, of which twenty-three are ranked.

2️⃣ arXiv · Generating Leakage-Free Benchmarks for Robust RAG Evaluation(⭐⭐⭐⭐⭐ 必读评测方法论)
arXiv · Generating Leakage-Free Benchmarks for Robust RAG Evaluation(⭐⭐⭐⭐⭐ 必读评测方法论)
arXiv:2605.08838 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出了 SeedRG,一个用于缓解 knowledge leakage 并应对 benchmark aging 问题的半合成 benchmark 生成 pipeline。SeedRG is introduced, a semi-synthetic benchmark generation pipeline that mitigates knowledge leakage and addresses the issue of benchmark aging.

IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications
IdeaAMBIG:研究构想规范中面向实现的关键缺口基准
arXiv:2609.10539 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

研究面向实现的研究方法规范的可编码化就绪度,定义为其是否为胜任的实现者或编码 Agent 提供足够的方法学信息,以在不引入未支持假设的情况下构建预期方法。This work studies the codification readiness of implementation-facing research-method specifications, defined by whether they provide sufficient methodological information for a competent implementer or coding agent to construct the intended method without unsupported assumptions.

Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation
Benchmark Radar:面向 AI 基准与评测的活体数据库与搜索引擎
arXiv:2609.11115 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Benchmark Radar,一个用于 AI 基准检索与发现的活体数据库和搜索引擎,覆盖 LLM 评测、Agent 与工具使用基准、代码、推理、安全及领域评测。Benchmark Radar is presented, a living database and search engine for retrieval and discovery of AI benchmarks, covering LLM evaluation, agentic and tool-use benchmarks, coding, reasoning, safety, and domain-specific evaluations.

2️⃣ arXiv · "Towards Automated Kernel Generation in the Era of LLMs"(Survey)
arXiv · "Towards Automated Kernel Generation in the Era of LLMs"(Survey)⭐⭐⭐⭐
arXiv:2601.15727 评测基准 综述 OA · 绿色 被引 12 · S2

本文聚焦 LLM 驱动的 kernel generation 领域,给出现有方法的结构化综述,涵盖 LLM-based 方法与 agentic optimization workflow,并系统梳理了支撑该领域学习与评测的数据集与 benchmark。This survey addresses the gap in LLM-driven kernel generation by providing a structured overview of existing approaches, spanning LLM-based approaches and agentic optimization workflows, and systematically organizing the datasets and benchmarks that underpin learning and evaluation in this domain.

DataFlex-RL: An Evaluation Platform for RLVR Data Policies
DataFlex-RL:面向 RLVR 数据策略的评测平台
arXiv:2609.06107 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出了 DataFlex-RL,一个在统一 GRPO 方案下比较数据策略的评测平台;研究发现,改变数据策略会显著影响训练过程,但相对于均匀训练并未带来可复现的提升。DataFlex-RL, an evaluation platform for comparing choices under a common GRPO recipe, is introduced, finding that changing the data policy measurably changes the training process but does not produce a reproducible improvement over uniform training.

Feature Recovery for Object Understanding After Irreversible Fire Damage
Feature Recovery for Object Understanding After Irreversible Fire Damage
arXiv:2609.12078 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出即插即用的 Feature Recovery Module (FRM),可在保持宿主网络冻结的前提下,将退化编码器特征映射到与干净图像对齐的表征;该模块提升了场景级检测、CLIP/SigLIP2 特征恢复以及全部四项物体级 VLM 任务,且退化越严重增益越大。The Feature Recovery Module (FRM), a plug-and-play module that maps degraded encoder features to pristine-aligned representations while keeping the host frozen, is proposed, which improves scene-level detection, CLIP/SigLIP2 feature recovery, and all four object-level VLM tasks, with larger gains under more severe degradation.

Towards a Deterministic Math Solver for Clinical Language Models
Towards a Deterministic Math Solver for Clinical Language Models
arXiv:2609.10728 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

即便在公式、变量与笔记访问条件一致时,向 Program-Solve 接口加入 executor 对不同开源权重模型的帮助程度各异,且无论哪种情况都无法替代经过验证的公式或可靠的变量抽取。Adding an executor to the Program-Solve interface helps some open-weight models more than others even under matched formula, variable and note access, and is not a substitute for verified formulas or reliable variable extraction either way.

Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks
Pick Your Poison:学习选择投毒集合以实现更强的 LLM 后门攻击
arXiv:2609.15029 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

将投毒选择形式化为 oracle 预算下的集合优化问题,提出 SAILS (Set-level Audit-Informed Iterative Learned Selection),通过数百次微调-评估运行学习集合打分器,对百万级候选集合排序,仅审计少量入围集合。This work formalizes poison selection as oracle-budgeted set optimization and introduces SAILS (Set-level Audit-Informed Iterative Learned Selection), which learns a set scorer from a few hundred finetune-and-evaluate runs, ranks millions of candidate sets, and audits only a small shortlist.

E2A-Bench: Benchmarking Evidence-to-Action Reliability in Financial Chart Reasoning
E2A-Bench:金融图表推理中证据到行动可靠性的基准测试
arXiv:2609.14302 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

论文提出了 E2A-Bench,一个面向金融图表推理的 969 查询基准,由 323 个 HS300 成分股在三种输入模态下构建,并附带由 OHLCV 确定性派生的证据锚点;结果表明金融 VLM 评估应追溯从证据到决策的完整链路,而非依赖单一幻觉分数。E2A-Bench is introduced, a 969-query benchmark for financial chart reasoning, constructed from 323 HS300 constituents under three input modalities with deterministic OHLCV-derived evidence anchors, and results show that financial VLM evaluation should trace the full evidence-to-action chain rather than rely on a single hallucination score.