研究库 论文知识库
Papers · organized/paper_cards

论文

286 张论文卡片 · 评测集

开放获取 全部 绿色 · 1640
ClawProBench: Trace-Aware Evaluation of AI Agents with Runtime Coverage and Frozen Workplace-Style Holdouts
ClawProBench:基于 trace 感知、运行时覆盖与冻结式工作场景 holdout 的 AI Agent 评测
arXiv:2608.22510 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 ClawProBench:基于 OpenClaw(具备 workspace 工具及浏览、记忆、消息、调度、skill、subagent 等原生能力的实时 agent 运行时)实例化的 trace-aware、runtime-native agent 评估基准。ClawProBench is presented, a trace-aware benchmark for runtime-native agent evaluation instantiated on OpenClaw, a live agent runtime with workspace tools and native surfaces for browsing, memory, messaging, scheduling, skills, and subagents.

From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms
从看见到行动:智能眼镜作为第一人称智能平台
arXiv:2608.24877 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本综述首次以统一框架系统研究智能眼镜,形式化第一人称数据流与受限任务效用,并提出覆盖采集、反应式感知、上下文辅助、持续状态、受控行动与具身耦合的 L0–L5 框架。This survey is the first to systematically study smart glasses through a unified framework, formalizing first-person data flow and constrained task utility, and introducing an L0-L5 framework spanning capture, reactive perception, contextual assistance, persistent state, governed action, and embodied coupling.

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training
LAION-BVD:一个用于多模态预训练的千万小时级开放视频数据集
arXiv:2608.24845 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 LAION-BVD,面向多模态学习的大规模开放视频数据集,包含从 CommonCrawl 收集的 1.3B 条平台特定视频 URL,并通过抽取场景切换帧,将视频帧作为图文数据的替代来源加以探索。This work presents LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl, and explores video frames as an alternative source of image-text data by extracting scene-changing frames.

Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios
Video-IFBench:面向视频理解场景的多模态 LLM 指令遵循能力评估
arXiv:2608.25529 多模态 评测集 被引 0 · S2

本文对 20 余个近期 MLLM 进行大规模评估,结果表明视频指令跟随对当前模型仍具挑战,尤其是在涉及多约束、语义约束或需要根据视频内容选择正确分支/路径的复杂条件结构时。This work conducts a large-scale evaluation of more than 20 recent MLLMs and shows that video instruction following remains challenging for current models, especially for instructions with many constraints, semantic constraints, or complex conditional structures that require selecting the correct branch or path based on video content.

PlanSightRAG: A Visual-First Multimodal RAG for Automating Question Answering and Compliance Checking for Civil Standard Plans
PlanSightRAG:面向土木标准图自动化问答与合规审查的视觉优先多模态 RAG
arXiv:2608.26091 RAG 检索增强 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

PlanSightRAG 是一种 Visual-First 多模态 RAG 框架,直接对图纸图像建立索引并进行推理,集成了 ColNomic-3B 多向量检索、Agentic Planner-Retriever-Auditor-Synthesizer,并以 MaxSim 热力图作为证据链。A Visual-First Multimodal Retrieval-Augmented Generation (RAG) framework called PlanSightRAG, which indexes and reasons directly over plan imagery, integrates a ColNomic-3B multi-vector retrieval, an agentic Planner-Retriever-Auditor-Synthesizer, and MaxSim heatmaps as an evidence trail.

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?
SWE Refactor Bench:编码 Agent 能否完成长时序全仓库技术栈迁移?
arXiv:2608.23564 Agent 智能体 评测集 OA · 绿色 被引 3 · S2

本文发布 SWE Refactor Bench,一个包含 20 项全仓库迁移的基准,涵盖 4 类技术债务,作为开发面向可靠全仓库迁移的编码 Agent 的严格测试平台。SWE Refactor Bench is introduced, a benchmark comprising 20 whole-repository migrations, covering 4 kinds of technical debt, and SWE Refactor Bench is positioned as a rigorous testbed for developing coding agents for reliable whole-repository migrations.

Skill Issue: Are Skills Language-Invariant in LLMs?
Skill Issue:LLM 中的 Skill 是否具备语言无关性?
arXiv:2608.25832 评测基准 评测集 OA · 绿色 被引 1 · S2

本文通过多语言 self-play,正交于知识与综合基准性能对跨语言技能不一致性进行量化,表明技能差异是开发真正多语言模型过程中可衡量且主要的障碍。This work quantifies cross-lingual skill inconsistency orthogonally from knowledge and general benchmark performance via multilingual self-play, and shows that skill discrepancies are a measurable major roadblock in the development of truly multilingual models.

GUI-Primitives: Diagnosing Spatial Reasoning Failures in Vision-Language GUI Grounding
GUI-Primitives:诊断视觉语言 GUI grounding 中的空间推理失败
arXiv:2608.21832 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

计算机使用 Agent 将自然语言指令 grounding 到截图中以定位界面元素,但现有基准无法隔离模型是否将关系语言正确绑定到对应元素。我们提出 GUI-Primitives,一个包含 994 个条目的基准,由跨七种空间关系(左右、上下、包含、对齐、邻近、列表序数、遮挡)的对比指令对组成。每对保持截图和锚点不变,仅改变关系表达,使正确目标在两个指定候选之间切换。五位标注者……Computer-use agents ground natural-language instructions in screenshots to locate interface elements, yet existing benchmarks do not isolate whether models bind relational language to the correct element. We introduce GUI-Primitives, a 994-item benchmark of contrastive instruction pairs over seven spatial relations in graphical user interfaces (left/right, above/below, containment, alignment, proximity, list ordinal, occlusion). Each pair holds the screenshot and anchor fixed while changing the relation expression, so the correct target moves between two designated candidates. Five annotators

Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents
画出所见:多模态 Agent 中灵巧视觉工具使用基准评测
arXiv:2608.25417 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 EASEL benchmark,用于评估受控的灵巧视觉工具使用,以参考引导的视觉重建为主代理任务:agent 逐步在画布上绘制以匹配参考图像。EASEL is proposed, a benchmark evaluating a controlled instance of dexterous visual tool use that adopts reference-guided visual reconstruction as its primary proxy task: the agent incrementally paints a canvas to match a reference image.

LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering
LoopArena:将模型作为 Loop 工程运行时控制器的基准评测
arXiv:2608.28281 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

提出 LoopArena,一个用于评估一个模型在长时任务中引导另一个独立 coding agent 能力的 benchmark,并在执行范围与成本各不相同的三个互补设置下评估该能力。This work introduces LoopArena, a benchmark for evaluating how well one model can guide a separate coding agent through a long-running task, and evaluates this ability in three complementary settings that differ in execution scope and cost.

Acquire, Repair, Preserve: A Diagnosis-Guided Post-Training Recipe for Small-Model Dialogue Game Agents
获取、修复、保留:面向小模型对话游戏 Agent 的诊断驱动后训练方案
arXiv:2608.28458 工程化 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,广泛的 SFT 带来模型大部分能力提升;当失败检测精准时,turn-local 监督可发挥作用,且观察到的迁移主要集中在同族模型之间。The results suggest that broad SFT brings most of the model's capability improvement; turn-local supervision can be effective when failure detection is precise, with observed transfer concentrated primarily within-family.

MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents
MNIST-PRO:MNIST 作为部分可观察世界回归,用于 AI Agents
arXiv:2608.31022 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 MnIST-PRO 基准,将 Agent 感知能力隔离测试:把 MNIST 数字识别转换为带回看约束的序列式 glimpse 搜索任务,并评估了十个多模态模型,结果表明仅获取视觉证据是不够的,Agent 还必须能够构建并更新可靠的感知状态。MnIST-PRO is addressed, a benchmark that isolates agentic perception by converting MNIST digit recognition into a sequential, glimpse-based search task with lookback constraints, and ten multimodal models are evaluated, showing that simply acquiring visual evidence is not enough and agents must also be able to build and update a reliable perceptual state.

PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback
PaperBanana-Interact:基于多轮人类反馈的科学图表精修
arXiv:2608.30241 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 MTPaperBananaBench,一个面向多轮图表生成的基准,包含 292 张图像和 3,518 条用户需求标注,并引入 PaperBanana-Interact,一个通过内部 critique-and-refine 循环来优化图表的多智能体系统。MTPaperBananaBench, a benchmark for multi-turn diagram generation containing 292 images annotated with 3,518 user requirements, is presented and PaperBanana-Interact, a multi-agent system that refines diagrams via an internal critique-and-refine loop is introduced.

ContextBias: Controlled Evaluation of Bias Persistence Under Context Shift in Text-to-Image Models
ContextBias:受控评估文本到图像模型在上下文偏移下的偏见持续性
arXiv:2608.29847 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

评估四个 SOTA 模型发现,将角色置于语义无关的上下文中并不会抑制该角色的关联属性;相反,跨角色的属性集中度会上升(合并 BI $+0.047$)。Evaluating four state-of-the-art models finds that placing a role in a semantically unrelated context does not suppress role-linked attributes; instead, cross-role attribute concentration increases (pooled BI $+0.047$).

EvoGenUI-Bench: Evaluating LLMs as Multi-Turn Generative UI Assistants
EvoGenUI-Bench:评估 LLM 作为多轮生成式 UI 助手
arXiv:2608.29387 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 EvoGenUI-Bench,一个面向多轮界面维护的基准,包含 150 个五轮任务,共计 750 轮,覆盖三种场景:信息呈现、可执行交互和工具驱动的外部状态。EvoGenUI-Bench is introduced, a benchmark for multi-turn interface maintenance comprising 150 five-turn tasks and 750 turns across three scenarios: information presentation, executable interaction, and tool-grounded external state.

UI-Venus-2 Technical Report
UI-Venus-2 技术报告
arXiv:2609.00028 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

本文提出 UI-Venus-2,一个面向移动、Web 和桌面环境的通用 GUI 基础 Agent,采用统一的闭环推理-行动框架,并通过集成安全感知机制来确保关键操作的可控执行。UI-Venus-2 is presented, a general-purpose foundation GUI agent designed to operate across mobile, web, and desktop environments through a unified closed-loop reasoning-action framework that integrates safety-aware mechanisms to ensure controlled execution of consequential actions.

From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix
从生产流量到后训练:构建覆盖企业请求组合的自托管 LLM
arXiv:2609.01572 工程化 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

数据驻留约束迫使企业自托管 LLM,但不断引入新模型而不下线旧模型会扩张服务集群,分散有限的 GPU 池。我们通过沿指令遵循、函数调用和内部任务分布三个维度,针对生产错误分析所发现的质量差距,将 200 多个内部应用的流量整合到单一模型上。质量通过按生产流量分层的离线基准进行跟踪,并由确定性验证器或经过校准的 LLM 评判器打分。不同于针对Data-residency constraints force enterprises to self-host LLMs, but continuous adoption of newer models without decommissioning their predecessors expands the serving fleet, fragmenting a finite GPU pool. We consolidate traffic from over 200 internal applications onto a single model by closing quality gaps identified through production error analysis along three axes: instruction following, function-calling, and internal task distribution. Quality is tracked by offline benchmarks stratified to production traffic and scored by deterministic verifiers or calibrated LLM judges. Rather than optimi

DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation
DramaChain Bench:面向短剧生成的端到端基准
arXiv:2609.00646 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 DramaChain Bench,首个覆盖完整生产链各阶段的短剧基准,并验证最终剧集质量并非仅由视频生成决定。DramaChain Bench is presented, the first short-drama benchmark that evaluates every stage of the complete production chain, and confirms that final episode quality is not governed by video generation alone.

Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data
通过对非结构化数据的自适应结构化实现 token 高效的数据推理 Agent
arXiv:2608.31082 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 agentic data cracking,将非结构化数据自适应、推测性地结构化为推理自身的副产品,是面向非结构化数据上 Agentic reasoning 的下一代数据基础设施的第一步。This work proposes agentic data cracking, a method that structures unstructured data adaptively and speculatively as a byproduct of reasoning itself, a first step toward next-generation data infrastructure for agentic reasoning over unstructured data.

AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling
AgentJudgeBench:一个用于评估 LLM 法官在智能体工具调用方面表现的多难度基准
arXiv:2608.26623 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

首个系统性研究 LLM-as-a-judge 在工作流 DAG 上对 Agentic 工具调用评判可靠性的基准,区别于面向开放式文本或偏好的通用 LLM-as-a-judge 任务;揭示了当前 LLM judge 的根本局限,并给出面向 Agentic 系统可靠评估的实践指南。The first benchmark to systematically study LLM-as-a-judge reliability for agentic tool-calling over workflow DAGs, as distinct from the broader LLM-as-a-judge task of open-ended text or preference evaluation, exposes fundamental limitations of current LLM judges and yields practical guidelines for reliable evaluation in agentic systems.

HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
HarnessDev:LLM 能创建并演进自身的 Agent Harness 吗?
arXiv:2609.01437 评测基准 评测集 OA · 绿色 被引 5 · S2

研究发现,生成的 harness 在代码与搜索/研究任务上仍显著落后于成熟的人工参考方案,而在写作与机器学习实验任务上达到或超过所选参考方案,且执行成本差异巨大。It is found that generated harnesses remain substantially behind mature human-engineered references on code and on search and research, while matching or exceeding the selected references on writing and machine-learning experimentation, with large variation in execution cost.

Aspire: Can Models Self-Evolve from Vague Goals?
Aspire:模型能否从模糊目标中自我演进?
arXiv:2608.31111 评测基准 评测集 OA · 绿色 被引 1 · S2

本文提出 ASPIRE,面向模糊目标驱动自我演化的基准,表明模糊目标会将搜索资源导向目标解释阶段;并在涵盖 6 类目标的 520 题隐藏专家评测集上评估所得系统。This work introduces ASPIRE, a benchmark for vague-goal-driven self-evolution and shows that vague goals redirect search effort toward goal interpretation, and evaluates the resulting systems on a hidden, expert-authored set of 520 items spanning six goals.

5️⃣ arXiv · Is Agentic RAG Worth It? An Experimental Comparison of RAG Approaches(⭐⭐⭐⭐ 高优先级)
5️⃣ arXiv · Agentic RAG 是否值得?RAG 方法的实验对比(⭐⭐⭐⭐ 高优先级)
arXiv:2601.07711 RAG 检索增强 评测集 OA · 绿色 被引 6 · S2

基于实证对 "Enhanced" 与 "Agentic" RAG 范式进行评估,为真实场景中选取最有效的 RAG 设计(兼顾性能与成本)提供指导。An empirically driven evaluation of the "Enhanced" and "Agentic" RAG paradigms is conducted, offering guidance on selecting the most effective RAG design for real-world applications, considering both performance and costs.

S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?
S3Gym:LLM 能将自测试与自评判转化为自我提升吗?
arXiv:2608.31100 评测基准 评测集 OA · 绿色 被引 1 · S2

这些发现表明,仅识别成功动作并不足够;Agent 还需将反馈转化为可执行且可迁移的策略。本文给出统一框架以诊断该过程,并定位阻碍 Agent 将交互经验转化为可靠自我提升的瓶颈。These findings show that recognizing successful actions is insufficient; agents must also transform feedback into executable and transferable policies, and provide a unified framework for diagnosing this process and identifying the bottlenecks that prevent agents from translating interaction experience into reliable self-improvement.

SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions
SnapBench:面向移动端交互的即拍即问多模态检索 benchmark
arXiv:2608.29607 RAG 检索增强 评测集 OA · 绿色 被引 1 · S2

本文提出 SnapBench——首个面向鲁棒"拍照即问"多模态检索的配对基准,以及一种简单的自适应融合方法 MOOR(Modality-anchored, Outlier-aware, Optimal Reweighting),并指出在"拍照即问"检索中需要具备可靠性感知的模态校准。SnapBench is introduced, the first paired benchmark for robust snap-and-ask multimodal retrieval, and MOOR (Modality-anchored, Outlier-aware, Optimal Reweighting), a simple adaptive fusion approach, highlighting the need for reliability-aware modality calibration in snap-and-ask retrieval.

Improving Health Literacy through Lay Summarization of Radiological Reports: An Evaluation of BioNER and Retrieval-Augmented Generation
通过放射学报告的大众化摘要提升健康素养:BioNER 与检索增强生成的评估
arXiv:2609.02396 RAG 检索增强 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本研究探讨了相较于标准 LLM 生成方式,RAG 与命名实体识别在多大程度上能提升自动生成通俗摘要的质量、事实一致性及可读性。This study investigates the extent to which Retrieval-Augmented Generation and Named Entity Recognition improve the quality, factual consistency, and readability of automatically generated lay summaries compared with standard LLM-based generation.

Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents
无知还是无能?为 LLM 智能体构建知识门控的可验证任务
arXiv:2608.30322 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一种知识门控的任务构建协议,将任务指令与一个包含私有约定、参考表与效用算子的紧凑工件分离,并证明被保留的任务能够改善 post-training。A knowledge-gated task-construction protocol is introduced that separates a task instruction from a compact artefact containing private conventions, reference tables, and utility operators, and it is shown that the retained tasks improve post-training.

Last Translation Benchmark
终极翻译基准(Last Translation Benchmark)
arXiv:2609.04173 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 The Last Translation Benchmark,一组由人工编写并经同行评审的样本,可打破领先的机器翻译模型,并提出新评估方法:每个样本附带手工编写的验证规则,描述该样本上的具体失败案例,从而支持可靠且可操作的未来评估。The Last Translation Benchmark is introduced, a collection of human-authored and peer-reviewed examples that break leading machine translation models and a new evaluation approach: each example comes with handcrafted verification rules describing concrete failure cases on that example, therefore allowing reliable and actionable future evaluation.

Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization
优化之前先提问:交互式优化中的动态预形式化澄清。
arXiv:2609.05258 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出了 OR-Clarify,一个用于预表述澄清的 benchmark,并提出了 Interactive Optimization (InterOPT),这是一个两阶段框架,能识别未解决的、影响表述的关键缺口,并据此引导系统决定是提出下一个问题还是停止提问。This work introduces OR-Clarify, a benchmark for pre-formulation clarification and proposes Interactive Optimization (InterOPT), a two-stage framework that identifies unresolved formulation-critical gaps and uses them to guide whether to ask the next question or to stop.

τ^τ-Bench: An Environment for End-To-End, Realistic Agent Construction
τ^τ-Bench:面向端到端真实 Agent 构建的环境
arXiv:2609.04611 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出了 τ^τ-bench(读作 hyper-tau-bench),一个将 Agent 构建作为任务的 benchmark,将协作式 Agent 构建工作转化为面向 coding agent 的可度量目标。The $\tau^\tau$-bench (pronounced hyper-tau-bench), a benchmark that makes agent construction the task, is introduced to turn the work of cooperative agent building into a measurable target for coding agents.

Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models
知道不该回答什么:视觉语言模型中的选择性不遵从
arXiv:2609.04720 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 KoNA,一个评估 VLM 选择性不遵从的基准,覆盖五类情形:False Premise、Visual Inaccessibility、Universal Unknown、Task Feasibility 与 Safety。实验结果显示,微调后的模型能够区分可回答部分与需要不遵从的部分,并以符合任务要求的方式作答。KoNA is introduced, a benchmark for evaluating selective non-compliance in VLMs across five categories: False Premise, Visual Inaccessibility, Universal Unknown, Task Feasibility, and Safety, and results suggest that the fine-tuned models can distinguish between answerable components and those requiring non-compliance and respond in a task-appropriate manner.

HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals
HarvestBench:衡量 LLM Agent 是否愿意付出代价以避免杀死动物
arXiv:2609.04444 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

HarvestBench 是首个为避免 side effect 标定价格并将 side effect 命名为生物的 benchmark;在六个模型中有四个的每次作答遭遇 kill rate 对价格变化敏感。HarvestBench is the first benchmark to put a price on avoiding a side effect and name the side effect as a living creature and four out of six models'kill rate per answered encounter were sensitive to price changes.

RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?
RoboSPA:VLA 模型能否超越简单场景与短时序任务?
arXiv:2609.05324 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文发布大规模机器人操控数据集与基准,用于诊断 VLA 模型的具身推理能力,并将 RoboSPA 确立为面向更强、更可靠、更具泛化性具身 Agent 的挑战性诊断基准。A large-scale robotic manipulation dataset and benchmark for diagnosing embodied reasoning in VLA models, and establishes RoboSPA as a challenging diagnostic benchmark for developing more capable, reliable, and generalizable embodied agents.

Closing the Consistency Gap: Self-Evolving Agents That Learn to Stay on Course
弥合一致性差距:学会保持正轨的自进化 Agent
arXiv:2609.08832 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出一个自演化 Agent 框架,通过识别 Agent 轨迹中不稳定、低一致性的步骤并将其转化为情节记忆,以供后续运行调用,从而缩小一致性 gap。This work presents a self-evolving agent framework that reduces the consistency gap by identifying unstable, low-consistency steps in agent trajectories and converting them into episodic memory the agent can draw on in future runs.

VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification
VDiff-Bench:面向细粒度图像差异识别的挑战性 benchmark
arXiv:2609.06245 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

VDiff-Bench 提供一个针对性诊断基准,用于评估 MLLM 的比较视觉理解能力,揭示标准单图视觉-语言任务无法捕捉的失败模式。VDiff-Bench provides a targeted diagnostic for evaluating comparative visual understanding in MLLMs, exposing failures that are not captured by standard single-image vision-language tasks.

EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?
EVOHARNESSBENCH:你的 Agent 能否跟上不断演进的 Harness?
arXiv:2609.04280 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 EvoHarnessBench,一个在工具、技能与 Agent 三个维度上对可控 harness 演化条件下的 Agent 进行评测的 benchmark,并将 harness 演化确立为一项独立挑战:Agent 需要在持续演化的 harness 下保持原有有效行为EvoHarnessBench is introduced, a benchmark for evaluating agents under controlled harness evolution across three axes (tools, skills, and agents), and establishes harness evolution as a distinct challenge for building agents that can keep pace with an evolving harness while preserving previously effective behavior.