基于 CPI-Bench 对主流图像编辑模型的评测结果显示,CPI-Bench 增强了模型间的性能区分度;排名分析表明 CPI-Bench 与 Arena Image Edit Leaderboard 的对齐度最高,与公开人类偏好排名具有更强的一致性。Evaluation results of mainstream image editing models based on CPI-Bench demonstrate that CPI-Bench enhances performance differentiation among models, and ranking analysis reveals that CPI-Bench achieves the highest alignment with the Arena Image Edit Leaderboard, indicating stronger consistency with public human preference rankings.
论文
226 张论文卡片 · 评测基准
在对 1,000 只美股进行小时级收益预测时,我们观察到一种意外现象:预测结果近乎平坦,且截面相关性衡量的股票排序能力很差,我们将其称为"预测坍缩"。令人惊讶的是,在同一设置下对交易量进行预测时,该现象基本消失。我们在多种时序基础模型(TSFMs)、12 个深度学习预测模型以及 97 个公开基准配置中系统考察了这一现象,发现其与目标可预测性密切相关,并识别出背后的两类成因:低 p……When forecasting hourly returns for 1,000 US equities, we observe an unexpected phenomenon: predictions become nearly flat and show poor stock ranking, as measured by cross-sectional correlation. We call this forecast collapse. Surprisingly, the phenomenon largely disappears when forecasting trading volume under the same setting. We investigate forecast collapse across time-series foundation models (TSFMs), twelve deep-learning forecasting models, and 97 public benchmark configurations, and find that it is closely tied to target predictability. We identify two distinct reasons behind it: low p
一个机器人过程评估工具包,将 rollout 视频转化为稠密进度曲线并衍生多项细粒度指标;引入 RoboPulse++ 用于评估过程奖励模型(PRM)的可靠性,为评测者提供更准确的测试平台。A toolkit for robot process assessment that turns rollout videos into dense progress curves and derives multiple fine metrics, and introduces RoboPulse++ to evaluate the reliability of process reward models (PRM), providing evaluators with a more accurate testing platform.
提出 MobileMem,一个面向设备端长期记忆研究的 benchmark 和框架,基于长达一年的移动端经验集合,使 agent 能够记忆过去、理解当下并适应未来。This work introduces MobileMem, a benchmark and framework for studying on-device long-term memory, grounded in a year-scale collection of mobile experiences, and enables agents to remember the past, understand the present, and adapt to the future.
提出 UNMASK,一个全自动 pipeline,可在无需额外人工标注的情况下发现、因果验证并缓解文本分类器中的伪相关,并证明其发现与验证阶段可泛化至奖励模型的偏好数据。U N M ASK is presented, a fully automated pipeline that discovers, causally verifies, and mitigates spurious correlations in text classifiers without additional human annotation, and demonstrates that the discovery and validation stages generalize to reward model preference data.
2023 年,一位纽约法官在 Mata v. Avianca 案中制裁了两位律师,因其提交的 brief 包含了由 ChatGPT 生成的虚构引用。此类失误大多能被数据库检索发现;但更棘手的问题在于检测那些指向真实案例、却不支持其所述命题的引用——这一失效模式是现有面向法律场景的 LLM 评测基本忽略的。本文通过对来自两个法律语料库的真实法律引用进行受控扰动(替换引用In 2023, a New York judge sanctioned two attorneys in Mata v. Avianca for filing a brief with hallucinated citations generated by ChatGPT. Such failures are largely caught by database lookups; the harder problem is detecting citations that point to real cases but do not support the propositions for which they are offered -- a failure mode that existing evaluations of LLMs for legal use cases largely overlook. In this paper, we study proposition-level citation support verification through controlled perturbations of real legal citations obtained from two legal corpora, either replacing the cite
提出 Apodex Discovery,一个通过 heavy-duty solver 构建和评估发现型 AI 的框架;该 solver 包含一个 foundation model、harness、工具和控制策略,用于执行长期的、有状态的、可验证的探索。This work introduces Apodex Discovery, a framework for building and evaluating discoverative AI through the heavy-duty solver, a system comprising a foundation model, harness, tools, and control policies that pursues extended, stateful, verifiable investigations.
提出 HarnessEval-W,一个 agentified 的评估 pipeline,将 LLM 生态中的 harness 范式引入 world model 基准测试,并在 330 个评估用例上对 18 个代表性 world model 进行了评估。This work introduces HarnessEval-W, an agentified evaluation pipeline that brings the harness paradigm from the LLM ecosystem to world model benchmarking, and applies HarnessEval-W to 18 representative world models over 330 evaluation cases.
语言模型以一种可被报告的形式持有潜在量,并且当任务需要灵活复用该量时,更多该量的信息会以这种形式存在。Language models hold latent quantities in a form they can report on, and more of a quantity is present in that form when the task requires reusing it flexibly when the task requires reusing it flexibly.
提出 Large Discovery Model (LDM),一种经验驱动的循环架构,将生成模型与贝叶斯非参数奖励代理模型耦合,产生一种感知不确定性的价值,用于引导候选的生成、精炼与选择。This work introduces the Large Discovery Model (LDM), an empirically grounded recurrent architecture that couples a generative model with a Bayesian non-parametric reward surrogate model, yielding an uncertainty-aware value that guides candidate generation, refinement, and selection.