研究库 论文知识库
Papers · organized/paper_cards

论文

226 张论文卡片 · 评测基准

开放获取 全部 绿色 · 1640
ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction
ExtractBench:一个面向模式引导的企业文档抽取基准
arXiv:2607.29677 评测基准 评测集 OA · 绿色 被引 1 · S2

LlamaExtract Agentic Plus在三项指标上均排名第一,准确度可与coding agent相媲美而成本仅为其一小部分,是首个同时在大规模下对数值准确性、记录完整性、grounding与实测成本进行打分的方法。LlamaExtract Agentic Plus ranks first on all three metrics, with accuracy comparable to coding agents at a fraction of the cost, and is the first to score value accuracy, record completeness at scale, grounding, and measured cost together.

Evaluation-Verification Reward for Consistent Multi-Reference Image Editing
用于一致多参考图像编辑的评估-验证奖励
arXiv:2607.29025 评测基准 方法 OA · 绿色 被引 1 · S2

多维评估-验证奖励(EVR)将评估分解为独立视觉准则;针对每个准则,MLLM Evaluator生成多个候选假设,Verifier在具体视觉证据中grounding每个claim以接受或拒绝,产生可靠且细粒度的奖励信号。A Multi-dimensional Evaluation-Verification Reward (EVR) decomposes evaluation into distinct visual criteria; for each criterion, an MLLM Evaluator generates multiple candidate hypotheses, and a Verifier grounds each claim in concrete visual evidence to accept or reject it, producing reliable and fine-grained reward signals.

Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants
更少澄清,更优代码:面向编程助手的跨会话个性化歧义自适应基准测试
arXiv:2607.26611 评测基准 评测集 OA · 绿色 被引 1 · S2

CAPA通过六种机制刻画个性化编码歧义,并使用受控的三阶段生成流程将这些机制注入无歧义的可执行任务,为开发长期编码助手奠定基础,使其生成的代码更好地对齐用户意图并减少反复澄清。CAPA, which characterizes personalized coding ambiguity through six mechanisms and injects these mechanisms into unambiguous executable tasks using a controlled three-stage generation pipeline, provides a foundation for developing long-term coding assistants that better align generated code with user intent while reducing repeated clarification.

RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems
RecHarness:用于自演化推荐系统的 Bandit 路由 Agentic Harness
arXiv:2607.29241 评测基准 方法 OA · 绿色 被引 3 · S2

RecHarness 将优化过程分为两步:bandit 路由器根据历史验证反馈选择下一步修改方向,LLM 在选定方向内生成具体优化假设与可执行代码编辑RecHarness separates the optimization process into two steps: a bandit router selects the next modification direction according to historical validation feedback, while the LLM generates a concrete optimization hypothesis and executable code edit within the selected direction.

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures
模型还是 Harness?面向 Agent 失败定位的以交互为中心的分类法
arXiv:2607.28802 评测基准 评测集 OA · 绿色 被引 5 · S2

本文提出以交互为中心的分类法,将失败定位到其起源的交互并识别责任组件,将 41 种失败模式归到两个组件之间的边及指示修复归属的故障侧This work introduces an interaction-centric taxonomy that localizes failures to the interactions in which they originate and identifies the responsible component, and organizes 41 failure modes by assigning each to an edge between two components and a fault side indicating where the repair belongs.

To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing
标题 -> 标题中文:增是机器,删是人工:LLM 代码编辑中删除回避的度量与缓解
arXiv:2607.28887 评测基准 方法 OA · 绿色 被引 1 · S2

在后训练中教授删除操作可减少删除回避行为并提升更广泛的代码编辑性能,表明该行为是训练不足而非不可达成Teaching deletion during post-training reduces deletion avoidance and improves broader code-editing performance, suggesting the behavior is undertrained rather than beyond reach.

When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills
当 Agent 学会成为你:Persona Skill 中隐私泄漏、冒充风险与防御的基准评测
arXiv:2608.03700 评测基准 评测集 OA · 绿色 被引 3 · S2

提出 AntiSkillBench,一个端到端的基准,用于评估 persona-skill 流水线中的风险与防御;实验表明 persona-skill 风险在不同的 agent backbone 和蒸馏协议下持续存在,从显式属性扩展到沟通风格与个性特征。AntiSkillBench is introduced, an end-to-end benchmark for evaluating risks and defenses across the persona-skill pipeline, and experiments show that persona-skill risks persist across agent backbones and distillation protocols, extending from explicit attributes to communication styles and personality traits.

Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements
LLM 的金融推理可信吗?基于长周期财报的真实世界检验
arXiv:2607.28661 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

FinIndices 是一个大规模基准,在未裁剪的财务报表(最长 32K tokens)上评估数据处理保真度,带来显著的零提示增益,验证通过以数据为中心的对齐可部分恢复结构化逻辑。FinIndices is a large-scale benchmark evaluating data-processing fidelity over uncropped financial statements (up to 32K tokens) and yields substantial zero-hint gains, validating that structured logic can be partially restored via data-centric alignment.

ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels
ChronoLens: 跨时间、语言与语言学层面的语言变化测量
arXiv:2608.03507 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 ChronoLens,结合冻结的多语言语言模型、特征对齐的 crosscoder 与事后语言学干预,应用于来自五个议会传统、跨越 1803–2026 年的 4498 万篇文档和约 172 亿 tokens,表明历史语言变化是一个结构化的、多维度的过程。ChronoLens is introduced, a framework that combines frozen multilingual language models, feature-aligned crosscoders, and post-hoc linguistic interventions, and applies it to 44.98 million documents and approximately 17.2 billion tokens from five parliamentary traditions spanning 1803--2026, showing that historical language change is a structured, multidimensional process.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks
GDPevo:面向真实业务任务的 agent 自进化评估
arXiv:2608.03764 评测基准 评测集 OA · 绿色 被引 1 · S2

提出 GDPevo,一种基于 GDP 相关企业工作流的 evolution-native 基准,并配套完全自动化的数据生成流水线;表现最佳的进化 agent 仍远低于全信息 oracle 上限,表明当前 agent 的自进化能力远未充分实现。GDPevo is presented, an evolution-native benchmark grounded in GDP-related enterprise workflows, together with the fully automated data pipeline that generates it, and the best evolved agents remain far below the fully informed oracle ceiling, indicating that the self-evolution ability of current agents remains far from fully realized.

FinanceHarness: Autonomous Financial Deep Research Framework
FinanceHarness:面向金融领域的自主深度研究框架
arXiv:2607.27853 评测基准 评测集 OA · 绿色 被引 1 · S2

FinanceHarness 是一个运行金融工具与从业者引导工作流的框架,端到端自动化金融深度研究:环境与数据构建、Agent 执行循环以及奖励建模;FinanceGym 包含论点驱动的研报问题与评分标准,结合 pre-cutoff 与 post-cutoff 准则。FinanceHarness is presented, a harness that runs finance-oriented tools and practitioner-guided workflows, automating financial deep research end to end: environment and data construction, the agent execution loop, and reward modeling, and FinanceGym, comprising thesis-driven research questions and rubrics that combine pre-cutoff and post-cutoff criteria.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
OSReward:建立跨平台计算机使用 Reward Model 的标准化评估
arXiv:2607.28609 评测基准 方法 OA · 绿色 被引 1 · S2

提出 OSReward,一个面向真实场景的高质量基准,用于评估 VLM 评判器在 CUA 轨迹上的表现;同时发布一个面向 CUA 社区的、带推理标注的开放轨迹判断语料库,以弥补大规模可靠 CUA 奖励的缺口。OSReward is introduced, a realistic, high-quality benchmark that evaluates VLM judges on CUA trajectories, and an open corpus of reasoning-annotated trajectory judgments for the CUA community, to close the gap in reliable CUA reward at scale.

MameLoshnLM: Yiddish Language Model and Evaluation Benchmark
MameLoshnLM:意第绪语语言模型与评估基准
arXiv:2608.05850 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

MameLoshnLM 是首个专为意第绪语构建的 8B 参数开源语言模型,既为意第绪语 NLP 提供基础,也为历史悠久但数字化程度不足的语言建模开发提供可复用的实践模板。MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish, is presented, providing both a foundation for Yiddish NLP and a practical template for language model development in historically rich but digitally underrepresented languages.

DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces
DataSpace:面向异构工作空间可验证分析的数据 Agent 基准
arXiv:2608.03451 评测基准 评测集 OA · 绿色 被引 2 · S2

提出 DataSpace,一个基准,用于评估数据 Agent 在任务本地异构工作空间中生成可验证表格结果的能力,并指出提升数据 Agent 可靠性的关键挑战。DataSpace, a benchmark in which data agents produce verifiable tabular results from task-local heterogeneous workspaces, is introduced and key challenges for improving data-agent reliability are identified.

BERTScore: Evaluating Text Generation with BERT
BERTScore:使用 BERT 评估文本生成
arXiv:1904.09675 评测基准 方法 OA · 绿色 被引 9857 · S2

本工作提出 BERTScore——一种文本生成自动评估指标,与人类判断的相关性更强,并在模型选择性能上优于现有指标。This work proposes BERTScore, an automatic evaluation metric for text generation that correlates better with human judgments and provides stronger model selection performance than existing metrics.

Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding
MLLM 能解读"创意跳跃"吗?面向跨概念理解的 C4 基准
arXiv:2608.06501 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 C4,一个受认知启发的成语跨概念创造力评估框架,揭示了当前 MLLM 在通过跨概念关系解码创造性编码语义方面存在的显著差距。C4, a cognition-inspired evaluation framework for Chengyu (Chinese idiom)-based Cross-Concept Creativity, is introduced, exposing a substantial gap in how current MLLMs decode creatively encoded meaning through cross-concept relations.

Taxonomy-Driven Analysis of Open-Source AI Risk Mitigation Tools
基于分类法的开源 AI 风险缓解工具分析
arXiv:2608.07446 评测基准 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一种结构化协议,通过对开源 LLM 评估与安全工具的分类驱动分析来自动化 AI 风险缓解,并给出一个可同时适用于开源与商用方案的分类驱动框架。This paper proposes a structured protocol to automate AI risk mitigation through a taxonomy-driven analysis of open-source LLM evaluation and security tools, and presents a taxonomy-driven framework applicable to open-source and proprietary solutions.

LitTraceQA: A Benchmark for Multi-Stage Grounding and Verification in Scientific Question Answering
LitTraceQA:科学问答中多阶段定位与验证的基准
arXiv:2608.07370 评测基准 评测集 OA · 绿色 被引 3 · S2

通过分别评估论文检索、证据 grounding 与答案准确性,LitTraceQA 为生成可验证答案、而非无依据摘要的科学 QA 系统提供了测试基准。By evaluating paper retrieval, evidence grounding, and answer accuracy separately, LitTraceQA provides a testbed for scientific QA systems that produce verifiable answers rather than unsupported summaries.

CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation
CodeXGLUE:面向代码理解与生成的机器学习基准数据集
arXiv:2102.04664 评测基准 评测集 OA · 绿色 被引 1628 · S2

本文介绍了 CodeXGLUE,一个基准数据集,旨在推动面向程序理解与生成的机器学习研究,涵盖 14 个数据集上的 10 项任务,并提供模型评估与比较的平台。This paper introduces CodeXGLUE, a benchmark dataset to foster machine learning research for program understanding and generation that includes a collection of 10 tasks across 14 datasets and a platform for model evaluation and comparison.

The Natural Language Decathlon: Multitask Learning as Question Answering
自然语言十项全能:将多任务学习视为问答
arXiv:1806.08730 评测基准 评测集 OA · 绿色 被引 669 · S2

于 2018 年 8 月 28 日中午 12:15 在 Pettit 微电子研究中心 102 A/B 室进行报告。Presented on August 28, 2018 at 12:15 p.m. in the Pettit Microelectronics Research Center, Room 102 A/B.

MatrAIx: Simulating the World with 8.3 Billion Persona Agents
MatrAIx:用 83 亿 Persona Agent 模拟世界
arXiv:2608.04205 评测基准 评测集 OA · 绿色 被引 9 · S2

本文提出 MatrAIx,一个面向异构用户的群体规模模拟用户评估基础设施,为使用多样化模拟人类用户评估 AI 系统和数字产品提供端到端支撑。MatrAIx is introduced, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users and provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.

Complementary Matrix-Gated QKAN Fast-Weight Programmers for Quantum Dynamics Forecasting
用于量子动力学预测的互补矩阵门控 QKAN 快权重编程器
arXiv:2607.27945 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Self-Modulating QKAN-based FWPs,用低秩生成的逐元素调制(作用于新提案分支、有界旧状态分支或两者)替代该广播门,并提出 Complementary Matrix Gating (CMG),在保持标量门控的有界凸更新和仿射前缀扫描结构的同时提供逐坐标的记忆控制。This work introduces Self-Modulating QKAN-based FWPs, which replace this broadcast gate with low-rank-generated element-wise modulation of the new-proposal branch, a bounded old-state branch, or both, and proposes Complementary Matrix Gating (CMG), which provides coordinate-wise memory control while preserving the bounded convex update and affine prefix-scan structure of scalar gating.

Evo-Bench: Can Language Models Improve Agent Harness?
Evo-Bench:语言模型能否改进 Agent Harness?
arXiv:2608.09096 评测基准 评测集 OA · 绿色 被引 6 · S2

Evo-Bench 是首个跨 Search、Office 和 General Agent 领域评估模型内在 harness 演化能力的基准,并暴露了早期饱和等关键时序异常,同时证明所合成的 harness 是高度可迁移的推理结构,能持续提升多样化策略模型。Evo-Bench is the first benchmark designed to evaluate models'intrinsic harness-evolving capabilities across Search, Office, and General agent domains, and exposes critical temporal anomalies like early saturation, while demonstrating that the synthesized harnesses act as highly transferable reasoning structures, consistently boosting diverse policy models.

BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
BDH-CQ:基于循环潜变量推理的上下文学习
arXiv:2608.09888 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 BDH-CQ,一个结合上下文学习与循环潜在推理的推理模型,通过在高维潜在空间中的迭代计算来求解查询,且无需将中间推理外化为语言。BDH-CQ is introduced, a reasoning model that combines in-context learning with recurrent latent reasoning that solves a query through iterative computation in a high-dimensional latent space, without verbalizing its intermediate reasoning.

MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models
MMOOC:面向多模态大语言模型上下文外评估的综合基准
arXiv:2607.27637 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 MMOOC,一个用于评估 MLLMs 拒答与鲁棒回答能力的大规模 benchmark,并引入 LLM-as-a-Judge 指标来衡量模型推理的正确性。This work presents MMOOC, a large-scale benchmark for evaluating refusal and robust answering abilities of MLLMs, and introduces an LLM-as-a-Judge metric to assess the correctness of model reasoning.

Generalized Out-of-Distribution Detection: A Survey
广义分布外检测:综述
arXiv:2110.11334 评测基准 综述 OA · 绿色 被引 1563 · S2

本文针对 OOD 检测领域的近期技术发展空白,提出统一框架 generalized OOD detection(广义 OOD 检测),涵盖上述五类问题,即 AD、ND、OSR、OOD detection 与 OD。This paper addresses the gap in recent technical developments in recent technical developments in the field of OOD detection by presenting a unified framework called generalized OOD detection, which encompasses the five aforementioned problems, i.e.,AD, ND, OSR, OOD detection, and OD.

Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness
Cultivar:用于调查数据污染与本地化鲁棒性的对比式、面向区域的翻译基准
arXiv:2608.09766 评测基准 评测集 OA · 绿色 被引 1 · S2

本文倡导源对比评估,并构建了 Cultivar——FLORES 的本地化子集,可用于特定 locale 的翻译评估;研究发现:MT 专用模型鲁棒性较差,少数模型可能对 FLORES 存在过拟合,且模型普遍更擅长翻译美国 locale 的内容,而非其他 locale,无论语种如何。This work advocates for source-contrastive evaluation and instantiates Cultivar, a localised subset of FLORES, which enables locale-specific translation evaluation and finds that MT-specialised models are less robust, a few models potentially overfit FLORES, and models tend to translate US content better than that of other locales, regardless of language.

Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure
无攻击者博弈:选择压力下 LLM 驱动搜索中的基准指纹化
arXiv:2608.08722 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

面向策略性优化下的可测性设计指南:保留探针仅在不可枚举轴上保持有效性;门控必须衡量保留集性能,而非仅正确性;迁移率只有在附带每类失败机制评级时才可解读。Design guidance for measurement under strategic optimization is distill design guidance for measurement under strategic optimization: held-out probes retain validity only on non-enumerable axes; gates must measure held-out performance, not just correctness; and a transfer rate is interpretable only with per-failure mechanism grades.

TSDS-Toolbox: A Toolbox for Measuring Time-Series Dataset Similarity
TSDS-Toolbox:用于衡量时间序列数据集相似性的工具箱
arXiv:2608.08119 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文实现了对时间序列数据集相似度方法的系统化、可复现比较,提供灵活的扩展性以添加自定义数据集、相似度方法及下游时间序列任务,并通过集成的时间序列数据集归约器对数据集级和序列级相似度方法进行一致评估。This work enables systematic and reproducible comparisons of time-series dataset similarity methods, flexible extensibility for users to add customized datasets, similarity methods, and downstream time-series tasks, and consistent evaluation of both dataset-level and series-level similarity methods through integrated time-series dataset reducers.

Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness
解码级 Taboo:面向 LLM 鲁棒性的诊断式压力测试
arXiv:2608.09900 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出解码层禁忌 (Decoding-Level Taboo),一种零提示的诊断式压力测试,在运行时直接干预 logit 空间,于词边界处动态遮蔽主要候选 token,强制模型进行迂回表达。Decoding-Level Taboo is introduced, a zero-prompt diagnostic stress test that intervenes directly in logit space at runtime, forcing models out of their nominal paths by dynamically masking primary candidate tokens at word boundaries, forcing machine circumlocution.

360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents
360CityArena:面向具身智能体的真实感虚拟城市导航基准
arXiv:2608.08814 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

360CityArena 为真实感城市区域导航与空间推理提供了必要且具有挑战性的测试平台;基于 SOTA LMM 智能体的评估显示,即使是最强的模型 Gemini 2.5 Flash,其表现仍远低于人类水平。360CityArena provides a necessary and challenging testbed for photorealistic urban-district navigation and spatial reasoning, and evaluation using state-of-the-art LMM-based agents shows that even the strongest model, Gemini 2.5 Flash, performs far below human level.

Towards A Rigorous Science of Interpretable Machine Learning
迈向严谨的可解释机器学习科学
arXiv:1702.08608 评测基准 观点 OA · 绿色 被引 5592 · S2

这篇立场论文定义了可解释性,阐述了何时需要(以及何时不需要)可解释性,并提出了一种用于严格评估的分类法,同时指出了迈向更严谨的可解释机器学习科学所面临的开放性问题This position paper defines interpretability and describes when interpretability is needed (and when it is not), and suggests a taxonomy for rigorous evaluation and exposes open questions towards a more rigorous science of interpretable machine learning.

Towards the Systematic Reporting of the Energy and Carbon Footprints of\n Machine Learning
迈向机器学习能耗与碳足迹的系统化报告
arXiv:2002.05651 评测基准 观点 OA · 绿色 被引 781 · S2

引入了一个框架,通过提供简洁接口来跟踪实时能耗与碳排放、生成标准化的在线附录来简化核算,并为节能的强化学习算法建立排行榜以激励负责任的研究A framework is introduced that makes accounting easier by providing a simple interface for tracking realtime energy consumption and carbon emissions, as well as generating standardized online appendices, and creates a leaderboard for energy efficient reinforcement learning algorithms to incentivize responsible research.

Solving Quantitative Reasoning Problems with Language Models
用语言模型解决定量推理问题
arXiv:2206.14858 评测基准 评测集 OA · 绿色 被引 2051 · S2
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives
LLM Agent 能否照本宣科?面向交互式叙事的长程一致性基准
arXiv:2608.08160 评测基准 评测集 OA · 绿色 被引 1 · S2

本文将该挑战形式化为叙事承诺保持 (Narrative Commitment Preservation, NCP),并提出 NCP-Bench:一个基于电影剧情梗概构建的包含 100 个叙事环境的基准,每个环境均提供可在玩家智能体与叙述者智能体交互过程中自动检查的结构化叙事规范。This work forms this challenge as Narrative Commitment Preservation (NCP), and introduces NCP-Bench, a benchmark of 100 narrative environments derived from movie synopses that each environment includes a structured narrative specification that can automatically check throughout the interaction between the player agent and the narrator agent.

LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers
LLMRouter:面向 LLM 路由器开发、评估与部署的统一基础设施
arXiv:2608.06867 评测基准 应用落地 OA · 绿色 被引 1 · S2

本文给出了 LLM routing 的统一形式化,将其刻画为由五个组件构成的序贯决策过程:context 编码器、模型编码器、评分函数、决策规则和学习信号,涵盖单轮、多轮和个性化 routing。This work presents a unified formulation of LLM routing as a sequential decision process characterized by five components: context encoders, model encoders, scoring functions, decision rules, and learning signals, covering single-turn, multi-turn, and personalized routing.