研究库 论文知识库
Papers · organized/paper_cards

论文

287 张论文卡片 · 评测集

开放获取 全部 绿色 · 1640
DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces
DataSpace:面向异构工作空间可验证分析的数据 Agent 基准
arXiv:2608.03451 评测基准 评测集 OA · 绿色 被引 2 · S2

提出 DataSpace,一个基准,用于评估数据 Agent 在任务本地异构工作空间中生成可验证表格结果的能力,并指出提升数据 Agent 可靠性的关键挑战。DataSpace, a benchmark in which data agents produce verifiable tabular results from task-local heterogeneous workspaces, is introduced and key challenges for improving data-agent reliability are identified.

Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events
AI 人格会成长吗?LLM Agent 在生活事件后的人格演化分析与基准测试
arXiv:2608.06485 Agent 智能体 评测集 OA · 绿色 被引 2 · S2

本文研究 11 项重大生活事件引发的人格变化,以大五人格作为心理测量锚点,并将所得轨迹与人类人格心理学的纵向证据进行对照,指出当前 PC-Agents 模拟了人类人格动态的均值,但未能模拟其形态。This work studies event-induced personality change after 11 major life events, using the Big Five traits as a psychometric anchor and interpreting the resulting trajectories against longitudinal evidence from human personality psychology, and suggests that current PC-Agents simulate the mean of human personality dynamics, but not its shape.

PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say
PrivacyPeek:审计 LLM Agent 获取了什么,而不仅仅是说了什么
arXiv:2606.00152 Agent 智能体 评测集 OA · 绿色 被引 5 · S2

实验表明,对敏感信息的不必要获取广泛存在,并观察到任务完成能力与获取阶段泄露之间的相关性,因此对获取阶段隐私进行审计既紧迫也必要。The experiments show that the unnecessary acquisition of sensitive information is widespread, and a correlation between the task-completion capability and acquisition-stage leakage, and auditing acquisition-stage privacy both urgent and necessary is observed.

LitTraceQA: A Benchmark for Multi-Stage Grounding and Verification in Scientific Question Answering
LitTraceQA:科学问答中多阶段定位与验证的基准
arXiv:2608.07370 评测基准 评测集 OA · 绿色 被引 3 · S2

通过分别评估论文检索、证据 grounding 与答案准确性,LitTraceQA 为生成可验证答案、而非无依据摘要的科学 QA 系统提供了测试基准。By evaluating paper retrieval, evidence grounding, and answer accuracy separately, LitTraceQA provides a testbed for scientific QA systems that produce verifiable answers rather than unsupported summaries.

CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation
CodeXGLUE:面向代码理解与生成的机器学习基准数据集
arXiv:2102.04664 评测基准 评测集 OA · 绿色 被引 1628 · S2

本文介绍了 CodeXGLUE,一个基准数据集,旨在推动面向程序理解与生成的机器学习研究,涵盖 14 个数据集上的 10 项任务,并提供模型评估与比较的平台。This paper introduces CodeXGLUE, a benchmark dataset to foster machine learning research for program understanding and generation that includes a collection of 10 tasks across 14 datasets and a platform for model evaluation and comparison.

The Natural Language Decathlon: Multitask Learning as Question Answering
自然语言十项全能:将多任务学习视为问答
arXiv:1806.08730 评测基准 评测集 OA · 绿色 被引 669 · S2

于 2018 年 8 月 28 日中午 12:15 在 Pettit 微电子研究中心 102 A/B 室进行报告。Presented on August 28, 2018 at 12:15 p.m. in the Pettit Microelectronics Research Center, Room 102 A/B.

Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection
多智能体取证推理用于可泛化的深度伪造视频检测
arXiv:2608.06865 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 FaceVid-Forensics-100K,一个大规模深伪视频数据集,包含 100,000 个视频,涵盖 33 种合成方法,覆盖换脸、表情重演与全脸合成;同时提出一个多智能体取证推理框架,由四个领域专家 Agent 分别从四个角度独立分析伪造线索。FaceVid-Forensics-100K is introduced, a large-scale deepfake video dataset comprising 100,000 videos and spanning 33 synthesis methods across face swapping, face reenactment, and entire-face synthesis, and a multi-agent forensic reasoning framework that employs four specialized domain-expert agents to independently analyze forgery cues from four perspectives.

CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models
CLIP-CC-Bench:评估视频-语言模型中的段落级视频描述
arXiv:2608.04302 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

CLIP-CC-Bench 为长视频描述提供了一个实用的评估框架,填补了现有短片段和仅 QA 基准的空白,并通过评分者间一致性(inter-judge agreement)与 bootstrap 排序稳定性量化该协议的内可靠性。CLIP-CC-Bench provides a practical evaluation framework for long-form video description, filling a gap left by existing short-clip and QA-only benchmarks and quantifying the protocol's internal reliability through inter-judge agreement and bootstrap ranking stability.

MatrAIx: Simulating the World with 8.3 Billion Persona Agents
MatrAIx:用 83 亿 Persona Agent 模拟世界
arXiv:2608.04205 评测基准 评测集 OA · 绿色 被引 9 · S2

本文提出 MatrAIx,一个面向异构用户的群体规模模拟用户评估基础设施,为使用多样化模拟人类用户评估 AI 系统和数字产品提供端到端支撑。MatrAIx is introduced, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users and provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.

Evo-Bench: Can Language Models Improve Agent Harness?
Evo-Bench:语言模型能否改进 Agent Harness?
arXiv:2608.09096 评测基准 评测集 OA · 绿色 被引 6 · S2

Evo-Bench 是首个跨 Search、Office 和 General Agent 领域评估模型内在 harness 演化能力的基准,并暴露了早期饱和等关键时序异常,同时证明所合成的 harness 是高度可迁移的推理结构,能持续提升多样化策略模型。Evo-Bench is the first benchmark designed to evaluate models'intrinsic harness-evolving capabilities across Search, Office, and General agent domains, and exposes critical temporal anomalies like early saturation, while demonstrating that the synthesized harnesses act as highly transferable reasoning structures, consistently boosting diverse policy models.

WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks
WeClawArena:人本 Agent 网络中跨用户 Agent 协作与安全的可审计沙箱与基准
arXiv:2608.03499 Agent 智能体 评测集 OA · 绿色 被引 3 · S2

本文提出 WeClawArena,一个面向个人工作空间多参与方 owned-agent 协作的可审计基准与运行时沙盒,基于有界运行时证据审计攻击成功情况,支持任务分解失败、隐私泄露、证据投毒以及权限路径失效等问题的诊断。WeClawArena is introduced, an auditable benchmark and runtime sandbox for multi-party owned-agent collaboration over personal workspaces and audits attack success from bounded runtime evidence, supporting diagnosis of task breakdown, privacy leakage, poisoned evidence, and invalid authority paths.

MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models
MMOOC:面向多模态大语言模型上下文外评估的综合基准
arXiv:2607.27637 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 MMOOC,一个用于评估 MLLMs 拒答与鲁棒回答能力的大规模 benchmark,并引入 LLM-as-a-Judge 指标来衡量模型推理的正确性。This work presents MMOOC, a large-scale benchmark for evaluating refusal and robust answering abilities of MLLMs, and introduces an LLM-as-a-Judge metric to assess the correctness of model reasoning.

Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness
Cultivar:用于调查数据污染与本地化鲁棒性的对比式、面向区域的翻译基准
arXiv:2608.09766 评测基准 评测集 OA · 绿色 被引 1 · S2

本文倡导源对比评估,并构建了 Cultivar——FLORES 的本地化子集,可用于特定 locale 的翻译评估;研究发现:MT 专用模型鲁棒性较差,少数模型可能对 FLORES 存在过拟合,且模型普遍更擅长翻译美国 locale 的内容,而非其他 locale,无论语种如何。This work advocates for source-contrastive evaluation and instantiates Cultivar, a localised subset of FLORES, which enables locale-specific translation evaluation and finds that MT-specialised models are less robust, a few models potentially overfit FLORES, and models tend to translate US content better than that of other locales, regardless of language.

Business Arena: Benchmarking LLM Agents in a Realistic Marketplace
Business Arena:在真实市场环境中基准测试 LLM Agents
arXiv:2608.08621 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

本文提出 Business Arena——一个受控环境,AI agent 在其中经营跨境店铺,在长周期内向供应商采购并向买家销售,迈出了构建面向端到端商业 agent 的真实可信测试床的第一步。Business Arena, a controlled environment where an AI agent runs a cross-border shop, buying from suppliers and selling to buyers over a long horizon, is introduced, taking a first step toward a realistic and trustworthy testbed for evaluating end-to-end business agents.

Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure
无攻击者博弈:选择压力下 LLM 驱动搜索中的基准指纹化
arXiv:2608.08722 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

面向策略性优化下的可测性设计指南:保留探针仅在不可枚举轴上保持有效性;门控必须衡量保留集性能,而非仅正确性;迁移率只有在附带每类失败机制评级时才可解读。Design guidance for measurement under strategic optimization is distill design guidance for measurement under strategic optimization: held-out probes retain validity only on non-enumerable axes; gates must measure held-out performance, not just correctness; and a transfer rate is interpretable only with per-failure mechanism grades.

TSDS-Toolbox: A Toolbox for Measuring Time-Series Dataset Similarity
TSDS-Toolbox:用于衡量时间序列数据集相似性的工具箱
arXiv:2608.08119 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文实现了对时间序列数据集相似度方法的系统化、可复现比较,提供灵活的扩展性以添加自定义数据集、相似度方法及下游时间序列任务,并通过集成的时间序列数据集归约器对数据集级和序列级相似度方法进行一致评估。This work enables systematic and reproducible comparisons of time-series dataset similarity methods, flexible extensibility for users to add customized datasets, similarity methods, and downstream time-series tasks, and consistent evaluation of both dataset-level and series-level similarity methods through integrated time-series dataset reducers.

Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness
解码级 Taboo:面向 LLM 鲁棒性的诊断式压力测试
arXiv:2608.09900 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出解码层禁忌 (Decoding-Level Taboo),一种零提示的诊断式压力测试,在运行时直接干预 logit 空间,于词边界处动态遮蔽主要候选 token,强制模型进行迂回表达。Decoding-Level Taboo is introduced, a zero-prompt diagnostic stress test that intervenes directly in logit space at runtime, forcing models out of their nominal paths by dynamically masking primary candidate tokens at word boundaries, forcing machine circumlocution.

360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents
360CityArena:面向具身智能体的真实感虚拟城市导航基准
arXiv:2608.08814 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

360CityArena 为真实感城市区域导航与空间推理提供了必要且具有挑战性的测试平台;基于 SOTA LMM 智能体的评估显示,即使是最强的模型 Gemini 2.5 Flash,其表现仍远低于人类水平。360CityArena provides a necessary and challenging testbed for photorealistic urban-district navigation and spatial reasoning, and evaluation using state-of-the-art LMM-based agents shows that even the strongest model, Gemini 2.5 Flash, performs far below human level.

Solving Quantitative Reasoning Problems with Language Models
用语言模型解决定量推理问题
arXiv:2206.14858 评测基准 评测集 OA · 绿色 被引 2051 · S2
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives
LLM Agent 能否照本宣科?面向交互式叙事的长程一致性基准
arXiv:2608.08160 评测基准 评测集 OA · 绿色 被引 1 · S2

本文将该挑战形式化为叙事承诺保持 (Narrative Commitment Preservation, NCP),并提出 NCP-Bench:一个基于电影剧情梗概构建的包含 100 个叙事环境的基准,每个环境均提供可在玩家智能体与叙述者智能体交互过程中自动检查的结构化叙事规范。This work forms this challenge as Narrative Commitment Preservation (NCP), and introduces NCP-Bench, a benchmark of 100 narrative environments derived from movie synopses that each environment includes a structured narrative specification that can automatically check throughout the interaction between the player agent and the narrator agent.

OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution
OpenART:通过开放式环境演化扩展 Agent 红队测试
arXiv:2608.00677 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出进化式马尔可夫超图攻击(EMHA),这是一种黑盒策略,通过协调授权状态转移执行反馈驱动的环境演化,无需参数更新,并将 OpenART 确立为在复杂演化环境中研究 Agent 安全性的可扩展基础。This work proposes the Evolutionary Markov Hypergraph Attack (EMHA), a black-box policy that performs feedback-driven environment evolution by coordinating authorized state transitions without requiring parameter updates, and establishes OpenART as a scalable foundation for studying agent safety in complex, evolving environments.

H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models
H2R-Bench:世界模型中人到机器人操作视频生成的基准评测
arXiv:2608.13049 多模态 评测集 OA · 绿色 被引 1 · S2

H2R-Bench 提供了一个系统性诊断框架,用于评估视频世界模型能否跨越 human-to-robot 具身差距,并将人类操作观测转化为以机器人为中心的训练资源。H2R-Bench provides a systematic diagnostic framework for evaluating whether video world models can bridge the human-to-robot embodiment gap and convert human manipulation observations into robot-centric training resources.

CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing
CPI-Bench:面向真实图像编辑场景的综合、实用、智能基准
arXiv:2608.14546 评测基准 评测集 OA · 绿色 被引 2 · S2

基于 CPI-Bench 对主流图像编辑模型的评测结果显示,CPI-Bench 增强了模型间的性能区分度;排名分析表明 CPI-Bench 与 Arena Image Edit Leaderboard 的对齐度最高,与公开人类偏好排名具有更强的一致性。Evaluation results of mainstream image editing models based on CPI-Bench demonstrate that CPI-Bench enhances performance differentiation among models, and ranking analysis reveals that CPI-Bench achieves the highest alignment with the Arena Image Edit Leaderboard, indicating stronger consistency with public human preference rankings.

Forecast Collapse in Time-Series Foundation Models
时序基础模型中的预测坍缩现象
arXiv:2608.14106 评测基准 评测集 OA · 绿色 被引 1 · S2

在对 1,000 只美股进行小时级收益预测时,我们观察到一种意外现象:预测结果近乎平坦,且截面相关性衡量的股票排序能力很差,我们将其称为"预测坍缩"。令人惊讶的是,在同一设置下对交易量进行预测时,该现象基本消失。我们在多种时序基础模型(TSFMs)、12 个深度学习预测模型以及 97 个公开基准配置中系统考察了这一现象,发现其与目标可预测性密切相关,并识别出背后的两类成因:低 p……When forecasting hourly returns for 1,000 US equities, we observe an unexpected phenomenon: predictions become nearly flat and show poor stock ranking, as measured by cross-sectional correlation. We call this forecast collapse. Surprisingly, the phenomenon largely disappears when forecasting trading volume under the same setting. We investigate forecast collapse across time-series foundation models (TSFMs), twelve deep-learning forecasting models, and 97 public benchmark configurations, and find that it is closely tied to target predictability. We identify two distinct reasons behind it: low p

PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment
PRM-as-a-Judge 1.5:机器人过程评估工具包
arXiv:2608.14284 评测基准 评测集 OA · 绿色 被引 2 · S2

一个机器人过程评估工具包,将 rollout 视频转化为稠密进度曲线并衍生多项细粒度指标;引入 RoboPulse++ 用于评估过程奖励模型(PRM)的可靠性,为评测者提供更准确的测试平台。A toolkit for robot process assessment that turns rollout videos into dense progress curves and derives multiple fine metrics, and introduces RoboPulse++ to evaluate the reliability of process reward models (PRM), providing evaluators with a more accurate testing platform.

MobileMem: Learning from a Year of Mobile Experiences
MobileMem:从一年的移动端经验中学习
arXiv:2608.13606 评测基准 评测集 OA · 绿色 被引 1 · S2

提出 MobileMem,一个面向设备端长期记忆研究的 benchmark 和框架,基于长达一年的移动端经验集合,使 agent 能够记忆过去、理解当下并适应未来。This work introduces MobileMem, a benchmark and framework for studying on-device long-term memory, grounded in a year-scale collection of mobile experiences, and enables agents to remember the past, understand the present, and adapt to the future.

A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images
迈向通用科学 AI 的路径:科学图像的多模态理解
arXiv:2608.14075 多模态 评测集 被引 1 · S2

本工作提出将"图像中的科学概念理解"作为长期基准目标,未来方向涵盖更广泛的领域与图表类型、上下文与跨文档综合、假设评估、出处溯源、不确定性、反事实基础,以及开放式多模态研究。This work proposes “scientific conceptual understanding from images” as a long-term benchmark objective, with future directions including broader domains and figure types, contextual and cross-document synthesis, hypothesis evaluation, provenance, uncertainty, counterfactual grounding, and open-ended multimodal research.

UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers
UNMASK:文本分类器中虚假捷径的发现与因果验证
arXiv:2608.09209 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 UNMASK,一个全自动 pipeline,可在无需额外人工标注的情况下发现、因果验证并缓解文本分类器中的伪相关,并证明其发现与验证阶段可泛化至奖励模型的偏好数据。U N M ASK is presented, a fully automated pipeline that discovers, causally verifies, and mitigates spurious correlations in text classifiers without additional human annotation, and demonstrates that the discovery and validation stages generalize to reward model preference data.

Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models
增强并不意味着可预测:思维模型中的推理行为
arXiv:2608.13760 工程化 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

发现面向推理的训练并未优先放大具有最高 Lift 的行为,这促使研究者采用过程级目标,以奖励经过校准且有依据的推理,而不仅仅是表面形式。It is found that reasoning-oriented training does not preferentially amplify the highest-Lift behaviors, motivating process-level objectives that reward calibrated and grounded reasoning rather than surface form alone.

Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence
Apodex Discovery:用于评估与构建探索型 AI 的现实基准与环境
arXiv:2608.11341 评测基准 评测集 OA · 绿色 被引 1 · S2

提出 Apodex Discovery,一个通过 heavy-duty solver 构建和评估发现型 AI 的框架;该 solver 包含一个 foundation model、harness、工具和控制策略,用于执行长期的、有状态的、可验证的探索。This work introduces Apodex Discovery, a framework for building and evaluating discoverative AI through the heavy-duty solver, a system comprising a foundation model, harness, tools, and control policies that pursues extended, stateful, verifiable investigations.

Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems
Agent 抓 Agent:临床多 Agent 系统中的捷径级联与 benchmark 作弊
arXiv:2608.03744 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

探讨共享工作空间上 LLM Agent 委员会的审议过程是否能被捷径和线索(benchmark 所奖励但临床医生会忽略的)所博弈,以及委员会的社会可信度所构成的游戏。It is asked whether committees of language-model agents deliberating on a shared workspace can be gamed by shortcuts, cues a benchmark rewards but a clinician would ignore, and what games a committee is social plausibility.

GRIP: Grounded Reasoning via Information-Restricted Premises
GRIP:通过信息受限前提的扎根推理
arXiv:2608.16776 RAG 检索增强 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 GRIP(Grounded Reasoning via Information-Restricted Premises),引入容量不对称:decoder 对 query 保持全维度访问,而检索到的证据则通过一个严苛的随机瓶颈,迫使证据通道仅编码 query 中无法获得的残余信息。GRIP (Grounded Reasoning via Information-Restricted Premises), which imposes capacity asymmetry: the decoder keeps full-dimensional access to the query, while retrieved evidence passes through a severe stochastic bottleneck, which forces the evidence channel to encode only the residual information unavailable from the query.

HarnessEval-W: Agentifying the Evaluation of Visual Worlds
HarnessEval-W:将视觉世界模型的评估 Agent 化
arXiv:2608.16859 评测基准 评测集 OA · 绿色 被引 2 · S2

提出 HarnessEval-W,一个 agentified 的评估 pipeline,将 LLM 生态中的 harness 范式引入 world model 基准测试,并在 330 个评估用例上对 18 个代表性 world model 进行了评估。This work introduces HarnessEval-W, an agentified evaluation pipeline that brings the harness paradigm from the LLM ecosystem to world model benchmarking, and applies HarnessEval-W to 18 representative world models over 330 evaluation cases.

Gathered, Not Admitted: How Attention Brings a Latent Variable into Verbalizable Form
Gathered, Not Admitted:注意力如何将潜变量带入可言语化的形式
arXiv:2608.15022 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

语言模型以一种可被报告的形式持有潜在量,并且当任务需要灵活复用该量时,更多该量的信息会以这种形式存在。Language models hold latent quantities in a form they can report on, and more of a quantity is present in that form when the task requires reusing it flexibly when the task requires reusing it flexibly.

TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation
TRACE-Bench:多参考图像生成的分解与诊断
arXiv:2608.16765 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

认识到多样化的多参考任务共享一组共同的原子操作,本文形式化了四个算子:Anchor、Disentangle、Apply 和 Compose,并构建了 TRACE-Bench,包含约 1,600 个跨 slot 数量 1–8 的评估用例。Recognizing that diverse multi-reference tasks share a common set of atomic operations, this work formalizes four operators: Anchor, Disentangle, Disentangle, Apply, and Compose, and constructs TRACE-Bench, comprising approximately 1,600 evaluation cases across slot counts 1--8.