提出 C4,一个受认知启发的成语跨概念创造力评估框架,揭示了当前 MLLM 在通过跨概念关系解码创造性编码语义方面存在的显著差距。C4, a cognition-inspired evaluation framework for Chengyu (Chinese idiom)-based Cross-Concept Creativity, is introduced, exposing a substantial gap in how current MLLMs decode creatively encoded meaning through cross-concept relations.
论文
139 张论文卡片 · 评测基准
本文提出一种结构化协议,通过对开源 LLM 评估与安全工具的分类驱动分析来自动化 AI 风险缓解,并给出一个可同时适用于开源与商用方案的分类驱动框架。This paper proposes a structured protocol to automate AI risk mitigation through a taxonomy-driven analysis of open-source LLM evaluation and security tools, and presents a taxonomy-driven framework applicable to open-source and proprietary solutions.
通过分别评估论文检索、证据 grounding 与答案准确性,LitTraceQA 为生成可验证答案、而非无依据摘要的科学 QA 系统提供了测试基准。By evaluating paper retrieval, evidence grounding, and answer accuracy separately, LitTraceQA provides a testbed for scientific QA systems that produce verifiable answers rather than unsupported summaries.
本文介绍了 CodeXGLUE,一个基准数据集,旨在推动面向程序理解与生成的机器学习研究,涵盖 14 个数据集上的 10 项任务,并提供模型评估与比较的平台。This paper introduces CodeXGLUE, a benchmark dataset to foster machine learning research for program understanding and generation that includes a collection of 10 tasks across 14 datasets and a platform for model evaluation and comparison.
于 2018 年 8 月 28 日中午 12:15 在 Pettit 微电子研究中心 102 A/B 室进行报告。Presented on August 28, 2018 at 12:15 p.m. in the Pettit Microelectronics Research Center, Room 102 A/B.
本文提出 MatrAIx,一个面向异构用户的群体规模模拟用户评估基础设施,为使用多样化模拟人类用户评估 AI 系统和数字产品提供端到端支撑。MatrAIx is introduced, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users and provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.
本文提出 Self-Modulating QKAN-based FWPs,用低秩生成的逐元素调制(作用于新提案分支、有界旧状态分支或两者)替代该广播门,并提出 Complementary Matrix Gating (CMG),在保持标量门控的有界凸更新和仿射前缀扫描结构的同时提供逐坐标的记忆控制。This work introduces Self-Modulating QKAN-based FWPs, which replace this broadcast gate with low-rank-generated element-wise modulation of the new-proposal branch, a bounded old-state branch, or both, and proposes Complementary Matrix Gating (CMG), which provides coordinate-wise memory control while preserving the bounded convex update and affine prefix-scan structure of scalar gating.
Evo-Bench 是首个跨 Search、Office 和 General Agent 领域评估模型内在 harness 演化能力的基准,并暴露了早期饱和等关键时序异常,同时证明所合成的 harness 是高度可迁移的推理结构,能持续提升多样化策略模型。Evo-Bench is the first benchmark designed to evaluate models'intrinsic harness-evolving capabilities across Search, Office, and General agent domains, and exposes critical temporal anomalies like early saturation, while demonstrating that the synthesized harnesses act as highly transferable reasoning structures, consistently boosting diverse policy models.
本文提出 BDH-CQ,一个结合上下文学习与循环潜在推理的推理模型,通过在高维潜在空间中的迭代计算来求解查询,且无需将中间推理外化为语言。BDH-CQ is introduced, a reasoning model that combines in-context learning with recurrent latent reasoning that solves a query through iterative computation in a high-dimensional latent space, without verbalizing its intermediate reasoning.
本文提出 MMOOC,一个用于评估 MLLMs 拒答与鲁棒回答能力的大规模 benchmark,并引入 LLM-as-a-Judge 指标来衡量模型推理的正确性。This work presents MMOOC, a large-scale benchmark for evaluating refusal and robust answering abilities of MLLMs, and introduces an LLM-as-a-Judge metric to assess the correctness of model reasoning.
本文针对 OOD 检测领域的近期技术发展空白,提出统一框架 generalized OOD detection(广义 OOD 检测),涵盖上述五类问题,即 AD、ND、OSR、OOD detection 与 OD。This paper addresses the gap in recent technical developments in recent technical developments in the field of OOD detection by presenting a unified framework called generalized OOD detection, which encompasses the five aforementioned problems, i.e.,AD, ND, OSR, OOD detection, and OD.
本文倡导源对比评估,并构建了 Cultivar——FLORES 的本地化子集,可用于特定 locale 的翻译评估;研究发现:MT 专用模型鲁棒性较差,少数模型可能对 FLORES 存在过拟合,且模型普遍更擅长翻译美国 locale 的内容,而非其他 locale,无论语种如何。This work advocates for source-contrastive evaluation and instantiates Cultivar, a localised subset of FLORES, which enables locale-specific translation evaluation and finds that MT-specialised models are less robust, a few models potentially overfit FLORES, and models tend to translate US content better than that of other locales, regardless of language.
面向策略性优化下的可测性设计指南:保留探针仅在不可枚举轴上保持有效性;门控必须衡量保留集性能,而非仅正确性;迁移率只有在附带每类失败机制评级时才可解读。Design guidance for measurement under strategic optimization is distill design guidance for measurement under strategic optimization: held-out probes retain validity only on non-enumerable axes; gates must measure held-out performance, not just correctness; and a transfer rate is interpretable only with per-failure mechanism grades.
本文实现了对时间序列数据集相似度方法的系统化、可复现比较,提供灵活的扩展性以添加自定义数据集、相似度方法及下游时间序列任务,并通过集成的时间序列数据集归约器对数据集级和序列级相似度方法进行一致评估。This work enables systematic and reproducible comparisons of time-series dataset similarity methods, flexible extensibility for users to add customized datasets, similarity methods, and downstream time-series tasks, and consistent evaluation of both dataset-level and series-level similarity methods through integrated time-series dataset reducers.
本文提出解码层禁忌 (Decoding-Level Taboo),一种零提示的诊断式压力测试,在运行时直接干预 logit 空间,于词边界处动态遮蔽主要候选 token,强制模型进行迂回表达。Decoding-Level Taboo is introduced, a zero-prompt diagnostic stress test that intervenes directly in logit space at runtime, forcing models out of their nominal paths by dynamically masking primary candidate tokens at word boundaries, forcing machine circumlocution.
360CityArena 为真实感城市区域导航与空间推理提供了必要且具有挑战性的测试平台;基于 SOTA LMM 智能体的评估显示,即使是最强的模型 Gemini 2.5 Flash,其表现仍远低于人类水平。360CityArena provides a necessary and challenging testbed for photorealistic urban-district navigation and spatial reasoning, and evaluation using state-of-the-art LMM-based agents shows that even the strongest model, Gemini 2.5 Flash, performs far below human level.
这篇立场论文定义了可解释性,阐述了何时需要(以及何时不需要)可解释性,并提出了一种用于严格评估的分类法,同时指出了迈向更严谨的可解释机器学习科学所面临的开放性问题This position paper defines interpretability and describes when interpretability is needed (and when it is not), and suggests a taxonomy for rigorous evaluation and exposes open questions towards a more rigorous science of interpretable machine learning.
引入了一个框架,通过提供简洁接口来跟踪实时能耗与碳排放、生成标准化的在线附录来简化核算,并为节能的强化学习算法建立排行榜以激励负责任的研究A framework is introduced that makes accounting easier by providing a simple interface for tracking realtime energy consumption and carbon emissions, as well as generating standardized online appendices, and creates a leaderboard for energy efficient reinforcement learning algorithms to incentivize responsible research.
本文将该挑战形式化为叙事承诺保持 (Narrative Commitment Preservation, NCP),并提出 NCP-Bench:一个基于电影剧情梗概构建的包含 100 个叙事环境的基准,每个环境均提供可在玩家智能体与叙述者智能体交互过程中自动检查的结构化叙事规范。This work forms this challenge as Narrative Commitment Preservation (NCP), and introduces NCP-Bench, a benchmark of 100 narrative environments derived from movie synopses that each environment includes a structured narrative specification that can automatically check throughout the interaction between the player agent and the narrator agent.
本文给出了 LLM routing 的统一形式化,将其刻画为由五个组件构成的序贯决策过程:context 编码器、模型编码器、评分函数、决策规则和学习信号,涵盖单轮、多轮和个性化 routing。This work presents a unified formulation of LLM routing as a sequential decision process characterized by five components: context encoders, model encoders, scoring functions, decision rules, and learning signals, covering single-turn, multi-turn, and personalized routing.
基于 CPI-Bench 对主流图像编辑模型的评测结果显示,CPI-Bench 增强了模型间的性能区分度;排名分析表明 CPI-Bench 与 Arena Image Edit Leaderboard 的对齐度最高,与公开人类偏好排名具有更强的一致性。Evaluation results of mainstream image editing models based on CPI-Bench demonstrate that CPI-Bench enhances performance differentiation among models, and ranking analysis reveals that CPI-Bench achieves the highest alignment with the Arena Image Edit Leaderboard, indicating stronger consistency with public human preference rankings.
在对 1,000 只美股进行小时级收益预测时,我们观察到一种意外现象:预测结果近乎平坦,且截面相关性衡量的股票排序能力很差,我们将其称为"预测坍缩"。令人惊讶的是,在同一设置下对交易量进行预测时,该现象基本消失。我们在多种时序基础模型(TSFMs)、12 个深度学习预测模型以及 97 个公开基准配置中系统考察了这一现象,发现其与目标可预测性密切相关,并识别出背后的两类成因:低 p……When forecasting hourly returns for 1,000 US equities, we observe an unexpected phenomenon: predictions become nearly flat and show poor stock ranking, as measured by cross-sectional correlation. We call this forecast collapse. Surprisingly, the phenomenon largely disappears when forecasting trading volume under the same setting. We investigate forecast collapse across time-series foundation models (TSFMs), twelve deep-learning forecasting models, and 97 public benchmark configurations, and find that it is closely tied to target predictability. We identify two distinct reasons behind it: low p
一个机器人过程评估工具包,将 rollout 视频转化为稠密进度曲线并衍生多项细粒度指标;引入 RoboPulse++ 用于评估过程奖励模型(PRM)的可靠性,为评测者提供更准确的测试平台。A toolkit for robot process assessment that turns rollout videos into dense progress curves and derives multiple fine metrics, and introduces RoboPulse++ to evaluate the reliability of process reward models (PRM), providing evaluators with a more accurate testing platform.
提出 MobileMem,一个面向设备端长期记忆研究的 benchmark 和框架,基于长达一年的移动端经验集合,使 agent 能够记忆过去、理解当下并适应未来。This work introduces MobileMem, a benchmark and framework for studying on-device long-term memory, grounded in a year-scale collection of mobile experiences, and enables agents to remember the past, understand the present, and adapt to the future.
提出 UNMASK,一个全自动 pipeline,可在无需额外人工标注的情况下发现、因果验证并缓解文本分类器中的伪相关,并证明其发现与验证阶段可泛化至奖励模型的偏好数据。U N M ASK is presented, a fully automated pipeline that discovers, causally verifies, and mitigates spurious correlations in text classifiers without additional human annotation, and demonstrates that the discovery and validation stages generalize to reward model preference data.
2023 年,一位纽约法官在 Mata v. Avianca 案中制裁了两位律师,因其提交的 brief 包含了由 ChatGPT 生成的虚构引用。此类失误大多能被数据库检索发现;但更棘手的问题在于检测那些指向真实案例、却不支持其所述命题的引用——这一失效模式是现有面向法律场景的 LLM 评测基本忽略的。本文通过对来自两个法律语料库的真实法律引用进行受控扰动(替换引用In 2023, a New York judge sanctioned two attorneys in Mata v. Avianca for filing a brief with hallucinated citations generated by ChatGPT. Such failures are largely caught by database lookups; the harder problem is detecting citations that point to real cases but do not support the propositions for which they are offered -- a failure mode that existing evaluations of LLMs for legal use cases largely overlook. In this paper, we study proposition-level citation support verification through controlled perturbations of real legal citations obtained from two legal corpora, either replacing the cite
提出 Apodex Discovery,一个通过 heavy-duty solver 构建和评估发现型 AI 的框架;该 solver 包含一个 foundation model、harness、工具和控制策略,用于执行长期的、有状态的、可验证的探索。This work introduces Apodex Discovery, a framework for building and evaluating discoverative AI through the heavy-duty solver, a system comprising a foundation model, harness, tools, and control policies that pursues extended, stateful, verifiable investigations.
提出 HarnessEval-W,一个 agentified 的评估 pipeline,将 LLM 生态中的 harness 范式引入 world model 基准测试,并在 330 个评估用例上对 18 个代表性 world model 进行了评估。This work introduces HarnessEval-W, an agentified evaluation pipeline that brings the harness paradigm from the LLM ecosystem to world model benchmarking, and applies HarnessEval-W to 18 representative world models over 330 evaluation cases.
语言模型以一种可被报告的形式持有潜在量,并且当任务需要灵活复用该量时,更多该量的信息会以这种形式存在。Language models hold latent quantities in a form they can report on, and more of a quantity is present in that form when the task requires reusing it flexibly when the task requires reusing it flexibly.
提出 Large Discovery Model (LDM),一种经验驱动的循环架构,将生成模型与贝叶斯非参数奖励代理模型耦合,产生一种感知不确定性的价值,用于引导候选的生成、精炼与选择。This work introduces the Large Discovery Model (LDM), an empirically grounded recurrent architecture that couples a generative model with a Bayesian non-parametric reward surrogate model, yielding an uncertainty-aware value that guides candidate generation, refinement, and selection.