研究库 论文知识库
Papers · organized/paper_cards

论文

217 张论文卡片 · 评测基准 · OA 绿色

开放获取 全部 绿色 · 1640
OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories
OpenTumorBoard:多学科肿瘤委员会讨论轨迹的真实世界基准
arXiv:2609.32810 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出了 OpenTumorBoard,一个包含 611 个患者案例、19,157 轮讨论、涵盖十个专科角色的基准,源自 YouTube 上 12,534 分钟公开肿瘤委员会会议录音的转写,并发布了自动化整理流水线,以支持多学科、个性化癌症决策中 LLM 的开发与评估。OpenTumorBoard is introduced, a benchmark with 611 patient cases and 19,157 discussion turns across ten specialist roles, transcribed from 12,534 minutes of publicly available tumor board recordings on YouTube, and automated curation pipeline is released to support the development and evaluation of LLMs for multidisciplinary, personalized cancer decision-making.

"I think this is the most disruptive technology": Exploring Sentiments of ChatGPT Early Adopters using Twitter Data
"我认为这是最具颠覆性的技术":基于 Twitter 数据探索 ChatGPT 早期采用者的情感
arXiv:2212.05856 评测基准 应用落地 OA · 绿色 被引 284 · S2

基于 10,732 条早期 ChatGPT 用户推文的混合方法研究,对每个主题进行深入定性情感分析,结果显示大多数早期采用者在软件开发颠覆性、娱乐与创意发挥等主题上表达了压倒性的积极情感。A mixed-method study using 10,732 tweets from early ChatGPT users to conduct an in-depth qualitative sentiment analysis of each topic, showing that the majority of the early adopters have expressed overwhelmingly positive sentiments related to topics such as Disruptions to software development, Entertainment and exercising creativity.

LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models
LexReward:面向法律语言模型的分类驱动奖励框架
arXiv:2609.39071 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

实验表明,基于评分细则的奖励能可靠地区分不同质量的法律回答,且在该偏好数据上进行 DPO 训练在所有三个维度上均提升了性能;逐维度分析进一步支持了所提分类法与奖励构建的有效性。Experiments show that the rubric-based rewards reliably distinguish legal responses of different quality and that DPO training on the preference data improves performance across all three dimensions, and Dimension-wise analyses further support the effectiveness of the proposed taxonomy and reward construction.

Multilingual GSM-Symbolic: What determines capability transfer across languages?
多语言 GSM-Symbolic:跨语言能力迁移由什么决定?
arXiv:2610.03367 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Multilingual GSM-Symbolic,一个可扩展的多语言数学数据集,包含 30,000 对题项匹配的问答对,覆盖 15 种语言;并将能力差异的最大决定因素量化为模型规模、语言资源水平、推理能力与类型学距离。This work introduces Multilingual GSM-Symbolic, an extensible multilingual mathematical dataset covering 30,000 item-matched question-answer pairs and spanning 15 languages and quantifies the largest determinants of capability as model size, language resource level, reasoning, and typological distance.

4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes
4DCodeBench:动态场景逆向图形中的 Agent 基准测试
arXiv:2610.03715 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 4DCodeBench,一个面向 4D 逆向图形(通过代码生成)的基准,Agent 以可执行图形程序的形式从视频重建动态场景,结果表明强大的静态重建能力尚未转化为对复杂动态的可靠重建。4DCodeBench is introduced, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs, finding that strong static reconstruction capabilities do not yet translate into reliable reconstruction of complex dynamics.

Tail-Influence Sampling for CVaR Policy Evaluation
面向 CVaR 策略评估的尾部影响采样
arXiv:2609.38096 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

为每个可查询的条件分布推导出一个尾部影响度量,聚合其不确定性对每一次 Bellman 复用中 CVaR 的影响,并刻画了尾部最优与均值最优分配重合的精确网格(exact-grid)机制。A tail influence for each queryable conditional law that aggregates how its uncertainty affects CVaR across every Bellman reuse is derived, and an exact-grid regime in which tail- and mean-optimal allocations coincide is characterized.

The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models
缺失的原语:诊断与修复大语言模型中的数学推理
arXiv:2610.02191 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

引入 Mathematical Primitive 概念来探查结构性数学理解,并提出 \hlei{},一个沿发现(Discovery)、生成(Generation)、消化(Digestion)和执行(Execution)四个维度评估数学推理的新型基准。This paper introduces the notion of Mathematical Primitive to probe structural mathematical understanding and proposes \hlei{}, a novel benchmark that evaluates mathematical reasoning along four distinct dimensions: Discovery, Generation, Digestion, and Execution.

Technical Report on the Turba Fertilizer Machine Learning Stack in Morocco
摩洛哥 Turba 施肥机器学习技术报告
arXiv:2610.05949 评测基准 评测集 OA · 绿色 被引 0 · OpenAlex

场地特异性施肥推荐系统会根据地点、土壤属性、作物类型和生产目标调整养分建议,但当推荐功能主要通过交互式界面访问、输出未做版本管理、训练后的近似模型无法独立加载或基准测试时,其科学复用性受到限制。本技术报告介绍 Turba 施肥机器学习技术栈——面向摩洛哥场地特异性施肥推荐的可复现三层开源实现,其中 turba-client 提供对公开……Site-specific fertilizer recommendation systems adapt nutrient advice to location, soil properties, crop type, and production targets, but scientific reuse is constrained when recommendation functions remain accessible mainly through interactive interfaces, outputs are not versioned, and trained approximations cannot be independently loaded or benchmarked. This technical report presents the Turba fertilizer machine learning stack, a three-layer open-source implementation for reproducible site-specific fertilizer recommendation in Morocco. turba-client provides programmatic access to publicly a

RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies
RoboDojo:面向通用机器人操作策略综合评估的仿真-真机统一基准
arXiv:2607.04434 评测基准 评测集 OA · 绿色 被引 44 · S2

提出 RoboDojo,一个面向通用机器人操作策略综合评估的仿真-真机统一基准,将 30 种策略集成到 XPolicyLab 并在 RoboDojo 上进行评测,建立了公开的排行榜与系统性的策略性能分析。RoboDojo is introduced, a unified sim-and-real benchmark for comprehensive evaluation of generalist robot manipulation policies that integrates 30 policies into XPolicyLab and evaluates them on RoboDojo, establishing a public leaderboard and systematic analysis of current policy performance.

PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages
PluraMath:将数学推理评估拓展至丰富资源语言之外
arXiv:2607.05992 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

确认了丰富资源语言与代表性不足语言之间在数学推理性能上存在持续差距,性能更优主要与更强的指令遵循能力相关,并提出了完全开源的数据集、数据采集流程与评估框架。A persistent gap in mathematical reasoning performance between high-resource and underrepresented languages is confirmed, with stronger results largely associated with better instruction-following ability, and a fully open-source dataset, data acquisition pipeline, and evaluation framework is introduced.

HETERQA: Benchmarking Record Retrieval over Multiple Heterogeneous Sources
HETERQA:跨多个异构来源的记录检索基准
arXiv:2607.03028 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出 HETERQA,一个包含 857 个 QA 对的综合性基准,涵盖五个异构来源的记录检索,并表明 HETERQA 为异构来源下的记录检索提供了有效的测试平台,为未来检索方法留下了显著空间。This work introduces HETERQA, a comprehensive benchmark with 857 QA pairs for record retrieval over five heterogeneous sources and indicates that HETERQA provides an effective testbed for record retrieval over heterogeneous sources and leaves substantial room for future retrieval methods.

Measuring the Gap Between Human and LLM Research Ideas
衡量人类与 LLM 研究思路之间的差距
arXiv:2607.01233 评测基准 方法 OA · 绿色 被引 9 · S2

结果表明,强大的 LLM 能够产出一系列合理思路,但其范围仍比人类研究品味更窄,并存在系统性偏移。It is suggested that strong LLMs can produce a range of reasonable ideas, but that range remains narrower than, and systematically shifted relative to, human research taste.

AGVBench: A Reliability-Oriented Benchmark of Data Augmentation for Vein Recognition
AGVBench:面向可靠性的静脉识别数据增强基准
arXiv:2607.02271 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

静脉识别是一种安全生物特征技术,常受限于标注数据稀缺与成像差异;而面向自然图像设计的增强策略可能破坏其关键的细粒度拓扑与纹理。本文提出 AGVBench,在 5 个公开掌/指静脉数据集、7 种骨干网络(含经典 CNN、视觉 Transformer 及静脉专用模型)上评测 30 种代表性增强策略。结果显示,多图混合类方法(如 MixUp、PuzzleMix……Vein recognition is a secure biometric technology often constrained by limited annotated data and imaging variations. While data augmentation mitigates this, strategies designed for natural images may disrupt the fine-grained topology and textures essential for identity discrimination. We present AGVBench, which evaluates 30 representative augmentation strategies on five public palm- and finger-vein datasets with seven backbone architectures, covering classic CNNs, vision transformers, and vein-specific recognition models. Our results show that multi-image mixing methods (e.g., MixUp, PuzzleMi

EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments
EvoPolicyGym:在交互式环境中评估自主策略演化
arXiv:2607.02440 评测基准 评测集 OA · 绿色 被引 2 · S2

提出"自主策略演化"评估范式:在固定交互预算下,由 harness-model Agent 反复编辑可执行策略系统;并在 EvoPolicyGym 中实例化,该基准基于一组紧凑型交互式 RL 环境构建,用于评测 Agent 如何迭代改进已探索策略。This work introduces Autonomous Policy Evolution, a controlled evaluation setting in which a harness-model agent repeatedly edits an executable policy system under a fixed interaction budget, and instantiates this setting in EvoPolicyGym, a benchmark built from compact interactive RL environments that evaluates how agents iteratively improve explored policies.

Cross-Domain Generalization Failure in Lightweight Intrusion Detection Models for IIoT Networks
面向 IIoT 网络的轻量级入侵检测模型的跨域泛化失效
arXiv:2607.00553 评测基准 应用落地 OA · 绿色 被引 2 · S2

应在真实类别分布下使用跨网络评估来判断部署就绪度,而非仅依赖域内准确率。Deployment readiness should be assessed using cross-network evaluation under realistic class distributions, rather than within-domain accuracy alone, to suggest deployment readiness should be assessed using cross-network evaluation under realistic class distributions, rather than within-domain accuracy alone.

Beyond IID: How General Are Tabular Foundation Models, Really?
超越 IID:表格基础模型的泛化能力究竟如何?
arXiv:2606.30410 评测基准 评测集 OA · 绿色 被引 11 · S2

BeyondArena 是首个面向表格数据的统一整体基准,支持多种任务类型(IID、时序、分组),覆盖样本量与特征维度的不同尺度,并涵盖来自广泛学科的多样化特征类型。BeyondArena is the first unified holistic benchmark for tabular data that supports diverse task types (IID, temporal, grouped), across sample size and feature dimensionality scales, with diverse feature types from a broad range of disciplines.

PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception
PerceptionRubrics:将多模态评估校准至人类感知
arXiv:2606.28322 评测基准 评测集 OA · 绿色 被引 2 · S2

提出 PerceptionRubrics——一个基于评分量表的评估框架,旨在弥合饱和的基准分数与真实场景脆弱性之间的差距,并验证了严格的感知保真是可靠生成的前提。PerceptionRubrics is introduced, a rubric-based evaluation framework that addresses the gap between saturated benchmark scores and real-world brittleness, validating that strict perceptual fidelity is the prerequisite for reliable generation.

How Good Can Linear Models Be for Time-Series Forecasting?
线性模型在时间序列预测中能做到多好?
arXiv:2606.27282 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

所得到的模型在大多数数据集-预测步长组合上优于先前的线性预测器,并在八个基准中的六个上超越 Transformer、MLP 和 CNN 基线;同时它还可作为对数据本身的诊断工具,揭示那些被更大模型默默吸收进其学习参数中的结构。The resulting models beat prior linear forecasters on most dataset-horizon entries and exceed Transformer, MLP, and CNN baselines on six of eight benchmarks, and serve as a diagnostic on the data itself, revealing structures that larger models absorb silently into their learned parameters.

Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments
穿越 gauntlet:重新评估 Agent 在熟悉环境之外的能力
arXiv:2606.14397 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 GauntletBench,一个用于评估 Agent 在挑战性场景中泛化能力的 Web 基准,聚焦于三种被低估的能力(时间感知、图形理解与 3D 推理),揭示了当前 Agent 能力与复杂真实场景所需能力之间的巨大差距。GauntletBench, a web-based benchmark for evaluating agent generalisation in challenging scenarios, focusing on three underexplored capabilities (temporal perception, graphical understanding, and 3D reasoning), is introduced, revealing the substantial gap between current agent capabilities and those required for complex real-world scenarios.

The Galaxy's Guide to the Tokenizer: A Benchmark for Scientific Foundation Models
The Galaxy's Guide to the Tokenizer: A Benchmark for Scientific Foundation Models
arXiv:2606.25610 评测基准 综述 OA · 绿色 被引 1 · S2

研究发现重建质量与表示质量是解耦的;在所考察的任务中,没有任何单一方法能够在所有任务上稳定取得最优表现。It is found that reconstruction and representation quality are decoupled, and no single method consistently performs best across the tasks considered here and no single method consistently performs best across the tasks considered here.

Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation
管理 LLM Agent 中的程序性记忆:控制、适应与评估
arXiv:2606.23127 评测基准 评测集 OA · 绿色 被引 9 · S2

一个包含 382 个真实企业任务、覆盖 6 种专业角色和 22 项程序性技能的基准,用于评估技能在任务、角色和模型骨干间的迁移能力,发现部分技能可在任务和模型间广泛泛化,而另一些则专化为角色特定工作流,在迁移时失去效力。A benchmark of 382 realistic enterprise tasks spanning six professional roles and 22 procedural skills, designed to evaluate how skills transfer across tasks, roles, and model backbones finds that some skills generalize broadly across tasks and models, whereas others become specialized to role-specific workflows and lose effectiveness under transfer.

PowerAgentBench-SS: A Benchmark for Agentic AI in Power System Steady-State Studies
PowerAgentBench-SS:面向电力系统稳态研究的 Agentic AI 基准
arXiv:2606.18789 评测基准 评测集 OA · 绿色 被引 4 · S2

结果表明仅评估求解器或仅评估答案是不足的:Agent 的差异不仅体现在发现关键 contingency 上,还体现在验证预算的使用、显式提交、类型强制、重复验证、基于证据的报告以及缓解行为等方面。The results show why solver-only or answer-only evaluation is insufficient: agents are distinguished not only by top-contingency discovery, but also by validation-budget use, explicit submission, type coercions, duplicate validations, evidence-backed reporting, and mitigation behavior.

AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification
AdvancedMathBench:一个面向高等数学证明生成与验证的基准测试套件
arXiv:2607.11849 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本工作提出 AdvancedMathBench,用于评估 LLM 在高级数学证明上的推理能力;同时引入 VerifierBench,包含 888 条由模型生成的证明轨迹及专家真值,用于评估模型能否正确判断证明有效性并给出合理的验证依据。This work introduces AdvancedMathBench, a benchmark suite designed to evaluate the reasoning capabilities of LLMs on advanced mathematical proofs, and introduces VerifierBench, consisting of 888 model-generated proof trajectories paired with expert ground truth, to evaluate whether models can correctly judge proof validity and provide sound verification rationales.

MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning
MET:理论支撑且文化感知的多语言道德推理
arXiv:2607.11736 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 MCLASH,一个多语言道德决策 benchmark,用于捕捉跨语言的文化情境化道德直觉和社会规范;并提出 MET(Multilingual Ethics with Theory-grounded reasoning),一种基于心理学与哲学专家策划的理论依据的两步提示方法。This work introduces MCLASH, a multilingual moral decision-making benchmark to capture culturally situated moral intuitions and social norms across languages, and proposes MET (Multilingual Ethics with Theory-grounded reasoning), a two-step prompting method built on expert-curated, theory-based grounds drawn from psychology and philosophy.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models
Blind-Spots-Bench:评估多模态模型中的盲点
arXiv:2607.08317 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

开发了自动化评分流水线,用于评估多种模型,包括开源权重模型、闭源语言模型、视觉语言模型和图像生成模型,结果显示没有单一模型在所有任务类型上占优,且某些任务对所有评估模型仍具挑战性。An automated grading pipeline is developed to evaluate a wide range of models, including open-weight and closed-source language, vision-language, and image-generation models, and shows that no single model dominates across all task types, and that some tasks remain challenging for all evaluated models.

Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms
深度强化学习评估与设计范式的原理性分析
arXiv:2607.07769 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

介绍其潜在原因的理论基础,阐明强化学习算法的渐近性能在性能排名与数据规模之间不存在单调关系。The theoretical foundations of the underlying causes outlining that the asymptotic performance of reinforcement learning algorithms does not have a monotone relationship between performance rankings and data-regimes are introduced.

LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes
LakeQuest:面向跨数据湖根植式问答的三领域基准
arXiv:2607.12310 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 LakeQuest,一个经人工校验的基准,包含 9,846 个 QA 对,用于在真实数据湖场景下评估端到端检索与综合 pipeline,并揭示现代问答系统中的关键失效模式。LakeQuest is introduced, a human-validated benchmark of 9,846 QA pairs designed to evaluate the end-to-end retrieve-and-synthesize pipeline over realistic data lakes and exposes critical failure modes in modern QA systems.

Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists
LLMs 是否准备好进行科学发现?面向 AI 科学家的能力导向基准
arXiv:2607.11079 评测基准 评测集 OA · 绿色 被引 1 · S2

提出 SDABench,一个围绕六项能力(描述性、探索性、推断性、推断性、预测性、因果性、机制性)跨五大领域(生物、化学、环境、地理、物理)重新组织评估的基准。SDABench is introduced, a benchmark that reorganizes evaluation around six capabilities (descriptive, exploratory, inferential, inferential, predictive, causal, and mechanistic) across five domains (Biology, Chemistry, Environment, Geography, Physics).

SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding
SynthDocBench:面向长上下文视觉文档理解的受控基准
arXiv:2607.10400 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

视觉语言模型(VLMs)在 DocVQA、ChartQA、MMLongBench-Doc 等视觉文档理解基准上表现强劲。但真实文档融合长度、布局复杂度、模态、问题难度等多因素,难以将模型失败归因于具体原因。我们提出 SynthDocBench,一个全合成的长上下文视觉文档理解基准,系统性控制文档长度、布局结构、模态构成与问题类型等因子。该基准的构建……Vision language models (VLMs) have achieved strong performance on visual document understanding benchmarks such as DocVQA, ChartQA, and MMLongBench-Doc. However, real-world documents combine multiple factors such as length, layout complexity, modality, and question difficulty, which makes it difficult to attribute model failures to specific causes. We introduce SynthDocBench, a fully synthetic benchmark for long-context visual document understanding that systematically controls factors including document length, layout structure, modality composition, and question type. The benchmark is constr

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
arXiv:2607.13705 评测基准 评测集 OA · 绿色 被引 3 · S2

AgentCompass 被提出,它是一个开源、轻量且可扩展的面向 LLM-based Agent 的评估基础设施,将评估流程围绕三个独立组件组织,从而在不重新实现复杂执行逻辑的前提下支持灵活配置。AgentCompass is introduced, an open-source, lightweight, and extensible infrastructure for evaluating LLM-based agents that organizes the evaluation process around three independent components, thereby enabling flexible configurations without requiring the reimplementation of complex execution logic.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
从受控环境到真实世界:面向实际场景的渗透测试 Agent 评估
arXiv:2605.10834 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一种实用评估协议,将评估重点从任务完成转向经过验证的漏洞发现,可在涵盖多种攻击面与漏洞类型的足够复杂目标上开展评估,并结合结构化真值标注与基于 LLM 的语义匹配来识别漏洞。This paper presents a practical evaluation protocol that shifts assessment from task completion to validated vulnerability discovery, allowing evaluation in sufficiently complex targets spanning multiple attack surfaces and vulnerability classes, and combines structured ground-truth with LLM-based semantic matching to identify vulnerabilities.

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence
Self in Space:面向 UAV 具身智能的自我意识与空间认知基准
arXiv:2607.12477 评测基准 评测集 OA · 绿色 被引 5 · S2

本文提出 SIS-Bench,一个在统一 self-in-space 表述下评估 UAV 场景具身空间智能的基准,并探索了一种融合光流与视觉特征的运动感知表征,以纳入与自身相关的动态信息。SIS-Bench is introduced, a benchmark for evaluating embodied spatial intelligence in UAV scenarios under a unified self-in-space formulation, and a motion-aware representation that incorporates self-related dynamics through optical flow and visual feature fusion is explored.

Length Penalties Make Chain-of-Thought Less Monitorable
长度惩罚使 Chain-of-Thought 更难被监控
arXiv:2607.09786 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本工作通过对推理模型施加长度惩罚以抑制过度思考并降低推理成本,但表明这些惩罚会削弱思维链的可监控性,通过移除监控所依赖的证据,以可监控性换取推理成本。This work trains reasoning models with length penalties to curb overthinking and cut inference cost but shows that these penalties make the chain of thought less monitorable, and trades monitorability for inference cost by removing the evidence monitors depend on.

UniVR: Thinking in Visual Space for Unified Visual Reasoning
UniVR:在视觉空间中思考以实现统一视觉推理
arXiv:2607.12800 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 UniVR,这是首个从纯视觉演示中同时学习复杂推理、细粒度物理动力学和长期规划的研究,并配套首个在纯视觉协议下评估这些异质能力的综合评测套件。UniVR, the first investigation into simultaneously learning complex reasoning, fine-grained physical dynamics, and long-term planning from pure visual demonstrations, is introduced, the first comprehensive suite to assess these heterogeneous capabilities under a purely visual protocol.

Rethinking the Evaluation of Harness Evolution for Agents
Rethinking the Evaluation of Harness Evolution for Agents(重新审视 Agent 的 Harness 演化评估)
arXiv:2607.12227 评测基准 评测集 OA · 绿色 被引 45 · S2

重新审视自动 harness 演化流程的评估方法,强调需要采用匹配预算的基线(matched-budget baselines)和留出评估(held-out evaluation),以区分真正的 harness 改进与针对特定基准的搜索和过拟合。The evaluation of automatic harness evolution procedures is revisited and the need for matched-budget baselines and held-out evaluation to distinguish genuine harness improvements from benchmark-specific search and overfitting is highlighted.

Sparks of Artificial General Intelligence: Early experiments with GPT-4
Sparks of Artificial General Intelligence: Early experiments with GPT-4
arXiv:2303.12712 评测基准 评测集 OA · 绿色 被引 4491 · S2

认为(该早期版本的)GPT-4 属于新一代具备更通用智能的 LLM(与 ChatGPT、谷歌 PaLM 等并列),并讨论了这些模型不断增强的能力及其影响。It is argued that (this early version of) GPT-4 is part of a new cohort of LLMs (along with ChatGPT and Google's PaLM for example) that exhibit more general intelligence than previous AI models, and the rising capabilities and implications of these models are discussed.