研究库 论文知识库
Papers · organized/paper_cards

论文

224 张论文卡片 · 评测基准

开放获取 全部 绿色 · 1640
E2A-Bench: Benchmarking Evidence-to-Action Reliability in Financial Chart Reasoning
E2A-Bench:金融图表推理中证据到行动可靠性的基准测试
arXiv:2609.14302 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

论文提出了 E2A-Bench,一个面向金融图表推理的 969 查询基准,由 323 个 HS300 成分股在三种输入模态下构建,并附带由 OHLCV 确定性派生的证据锚点;结果表明金融 VLM 评估应追溯从证据到决策的完整链路,而非依赖单一幻觉分数。E2A-Bench is introduced, a 969-query benchmark for financial chart reasoning, constructed from 323 HS300 constituents under three input modalities with deterministic OHLCV-derived evidence anchors, and results show that financial VLM evaluation should trace the full evidence-to-action chain rather than rely on a single hallucination score.

Thought without systematicity? Evaluating reasoning models on rule induction tasks
没有系统性的思维?在规则归纳任务上评估推理模型
arXiv:2609.13948 评测基准 观点 被引 0 · S2

人类认知的一个核心原则是系统性,即理解一个概念本质上与理解该概念的相近变体相关联。推理模型能否稳健地展现这种系统性?如果可以,我们应能预期模型在其结构等价变体任务上表现一致。本文扩展了认知科学中既定的规则归纳任务,以评估当前推理模型思维的系统性。每个任务族都具有组合结构,我们借此通过任务同构创建结构等价的任务变体...A central tenet of human cognition is systematicity, the principle that understanding one concept is inherently tied to understanding close variations of that concept. Do reasoning models robustly exhibit such systematicity? If so, we would expect consistent performance on structurally equivalent variants of the same task. Here, we extend established rule induction tasks from cognitive science to assess the systematicity of thought in current reasoning models. Each task family has compositional structure that we use to create structurally equivalent task variations through task isomorphisms su

2. Systemic Measurement Bias in LLM Inference Benchmarking
LLM Inference Benchmarking 中的系统性测量偏差
arXiv:2605.24217 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一个无偏的多进程 evaluation 框架,能够有效分散 client 端负载,从而在每秒数千次 query 以上的生产规模下实现对 LLM 的精确、可复现 profiling。This work proposes an unbiased, multi-process evaluation framework that effectively distributes client-side load, enabling accurate, reproducible profiling of LLMs at production scales exceeding thousands of queries per second.

ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals
ImpossibleRubrics:对生成式评分标准作为奖励信号的应激测试
arXiv:2609.16816 评测基准 评测集 被引 0 · S2

论文提出了 ImpossibleRubrics,一个包含 169 个不可能任务的基准,覆盖六类不可能性类别,每项任务都配有可验证的 oracle 证书,规定诚实回答可以与不可以声明的内容,并附带 48 个可回答的对照样本。ImpossibleRubrics is introduced, a benchmark of 169 impossible tasks spanning six impossibility categories, each paired with a verifiable oracle certificate specifying what an honest answer may and may not claim, together with 48 answerable controls.

HarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness
HarnessVLN:通过智能体框架统一免训练的具身导航
arXiv:2609.15195 评测基准 方法 OA · 绿色 被引 4 · S2

本工作提出 HarnessVLN,一个零样本、无需训练的框架:通过共享的 Agent Harness 统一指令跟随与物体目标导航,并展示了其在真实世界中两类导航任务上的适用性This work introduces HarnessVLN, a zero-shot, training-free framework that unifies instruction-following and object-goal navigation through a shared Agent Harness and demonstrates its applicability to both navigation tasks in real-world environments.

ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement
ModularRSI:模块化且可泛化的递归框架自我改进
arXiv:2609.14857 评测基准 评测集 OA · 绿色 被引 6 · S2

论文提出了 ModularRSI,一个与基准解耦、对比式、模块化的可泛化 harness 进化框架,对同一任务下成功与失败的轨迹进行对比,并跨任务聚合证据以识别反复出现的行为缺陷。ModularRSI is proposed, a benchmark-disjoint, contrastive, and modular framework for generalizable harness evolution that contrasts successful and failed trajectories for the same task and aggregates evidence across tasks to identify recurring behavioral deficiencies.

Another Blueprint In The Wall: How to Ask Frontier AI Like a Kid?
墙上又一张蓝图:如何像孩子一样向前沿AI提问?
arXiv:2609.14803 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

论文报告了来自 OpenAI、Anthropic、xAI 与 Google DeepMind 的六类前沿模型实验,并使用 epistemic jailbreak 一词来指称随请求具体性增加而伴随出现的技术溯源严谨性丧失现象。This paper reports experiments across six frontier model types from OpenAI, Anthropic, xAI, and Google DeepMind, and uses the term epistemic jailbreak for the accompanying loss of discipline in technical provenance as requested specificity increases.

OmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning
OmniHarness:通过符号化策略学习实现可泛化的视觉生成
arXiv:2609.16057 评测基准 方法 OA · 绿色 被引 1 · S2

本文提出 OmniHarness,一个通过符号策略学习实现可泛化视觉生成的框架,将已验证的执行抽象为视觉生成任务族的符号策略,捕获共享流程和适用条件,同时去除实例特定的输入。OmniHarness is introduced, a framework for generalizable visual generation via symbolic policy learning that abstracts verified executions into symbolic policies for visual generation task families, capturing shared procedures and applicability conditions while removing instance-specific inputs.

1️⃣1️⃣ arXiv · RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic RAG Systems(⭐⭐⭐ 参考)
arXiv · RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic RAG Systems(⭐⭐⭐ 参考)
arXiv:2510.13910 评测基准 评测集 被引 5 · S2

本文提出 RAGCap-Bench,一个面向能力的 benchmark,用于对 agentic RAG workflow 中的中间任务进行细粒度评测,并构建了典型 LLM 错误的分类体系以设计针对性评测问题。This work proposes RAGCap-Bench, a capability-oriented benchmark for fine-grained evaluation of intermediate tasks in agentic RAG workflows, and constructs a taxonomy of typical LLM errors to design targeted evaluation questions.

VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering
VākQA:泰卢固语口语事实型问答基准与评估研究
arXiv:2609.19879 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 V\={a}kQA,一个覆盖六个领域、包含 2,001 对事实型问答的泰卢固语 SQA 基准;并观察到:泰卢固语措辞保留了翻译中会丢失的文化特异性,语音输入引入的音近混淆会改变问题含义,级联 ASR-MT 误差会逐步叠加放大。V\={a}kQA, a Telugu SQA benchmark of 2,001 factoid question-answer pairs across six domains, is introduced and it is observed that Telugu phrasing retains cultural specificity that is lost in translation, speech input introduces phonetic confusions that alter question meaning, and cascaded ASR-MT errors compound progressively.

1️⃣ arXiv · Learning Rate Matters: Vanilla LoRA May Suffice(⭐⭐⭐⭐⭐ 必读)
学习率至关重要:Vanilla LoRA 可能已足够
arXiv:2602.04998 评测基准 方法 Open MIND OA · 绿色 被引 11 · S2

本文通过大规模超参数搜索,系统地重新评估了 Vanilla LoRA 以及九个代表性 LoRA 变体,发现不同 LoRA 方法偏好的学习率区间各异,并将最优学习率区间的差异归因于最大 Hessian 特征值的变化,与经典学习理论相吻合。This work systematically re-evaluate nine representative LoRA variants alongside vanilla LoRA through extensive hyperparameter searches, finding that different LoRA methods favor distinct learning rate ranges and attributes the differing optimal learning rate ranges to variations in the largest Hessian eigenvalue, aligning with classical learning theories.

TeleAntiFraud 2.0: A Refreshable, Profile-Grounded, and Audio-Based Benchmark for Telecom Fraud Detection
TeleAntiFraud 2.0:面向电信诈骗检测的可刷新、基于画像的音频基准
arXiv:2609.18748 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 TeleAntiFraud 2.0,采用 Mixed-Tree Anti-Fraud Generation Pipeline 构建,并基于月度冻结评测协议进行评估,确立了近域构造与 collapse-aware 报告作为在现实易混淆条件下评测音频电信诈骗模型的核心要求。This work presents TeleAntiFraud 2.0, constructed with the Mixed-Tree Anti-Fraud Generation Pipeline and evaluated under a monthly frozen evaluation protocol, establishing near-domain construction and collapse-aware reporting as core requirements for evaluating audio-based telecom-fraud models under realistic confusable conditions.

DeformSmith: Physics Harness-Guided Hierarchical Generation of Deformable Assets for Robot Manipulation
DeformSmith:基于物理 Harness 引导的可形变资产分层生成,用于机器人操作
arXiv:2609.18620 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,DeformSmith 生成的资产在视觉质量与物理合理性上均优于 SOTA 基线,包括 PhysGen3D、PhysGM 和 PhysX-Omni,同时支持为可形变物体的机器人操作合成训练数据。Results show that DeformSmith generates assets with better visual quality and physical plausibility than state-of-the-art baselines, including PhysGen3D, PhysGM, and PhysX-Omni, while supporting the synthesis of data for robotic manipulation of deformable objects.

Gricea: An Open Science Platform for Conversational AI Research
Gricea:一个面向对话式 AI 研究的开放科学平台
arXiv:2609.22039 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Gricea,一个开放科学平台,将研究表示为可配置、可部署的研究 artifact,研究者可以运行、检查、共享和复用;该平台展示了如何通过共享的研究 artifact 构建、复现和扩展 CAI 研究,从而借助开放科学实现知识的累积构建。This work presents Gricea, an open-science platform representing studies as configurable, deployable research artifacts that researchers can run, inspect, share, and reuse, demonstrating Gricea's support for constructing, reproducing, and extending CAI studies through shared research artifacts, enabling cumulative knowledge building through open science.

CADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design
CADWorld:面向长视野计算机辅助设计的计算机使用基准
arXiv:2609.16251 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

研究发现,较弱的 Agent 常常在产出有效成果之前即告失败,而较强的 Agent 则越来越多地在结构、几何与构造过程要求上失败;CADWorld 揭示了通用 GUI 能力与可靠执行持久、可验证工程工作流之间的差距。It is found that weaker agents often fail before producing a valid artifact, whereas stronger agents increasingly fail on structural, geometric, and construction-process requirements, and CADWorld exposes a gap between general GUI competence and reliable execution of persistent, verifiable engineering workflows.

GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay
GameHorizon Suite:游戏中多时间历程的数据与评估。
arXiv:2609.25001 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

GameHorizon Suite 是一套统一的数据与评估套件,可在多个时间跨度下衡量不同模型族的游戏能力,为跨水平跨度与跨模型族的游戏能力评估提供标准化标尺。The GameHorizon Suite, a unified data and evaluation suite that measures gameplay capabilities at different horizons for diverse model families, can provide a standardized yardstick for evaluating gameplay capabilities across horizons and model families.

Harness-Zero: Harness Distillation via Agent-as-Harness
Harness-Zero:通过 Agent-as-Harness 进行 Harness 蒸馏。
arXiv:2609.24974 评测基准 应用落地 被引 3 · S2

本文提出 Harness-Zero,通过 agent-as-harness 实现 harness 蒸馏,并证明在使用相同演化 harness 的前沿 LLM 上,agent-as-harness 优于 code-as-harness。This work introduces Harness-Zero, which enables harness distillation through agent-as-harness, and shows that for frontier LLMs using the same evolved harness, agent-as-harness outperforms code-as-harness.

Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles
Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles
arXiv:2609.22220 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出将变异分析作为 kernel-benchmark 预言机的充分性度量:通过确定性规则向 188 个 KernelBench 问题的已验证 CUDA 实现中注入 10 个可编译故障,其中 7,384 个具备独立击杀见证;任何测试协议均按其检出比例评分。Mutation analysis as an adequacy metric for kernel-benchmark oracles is introduced: deterministic rules inject 10 compilable faults into verified CUDA implementations of 188 KernelBench problems, 7{,}384 of them with an independent kill witness; any test protocol is scored by the fraction it detects.

A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal
大语言模型的测谎测试:读出模型不愿透露的知识
arXiv:2609.21996 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

内部识别探针(Probe of Internal Recognition, PIR)可区分"不愿回答"与"无法回答"的模型,支持装傻审计与遗忘验证,并从选择题扩展至自由生成。Probe of Internal Recognition (PIR) separates a model that will not answer from one that cannot, which supports sandbagging audits and unlearning verification, and extends from multiple-choice questions to free-form generation.

ALPINE: Adaptive Localization for Parameter- and Sample-Efficient Few-Shot Learning
ALPINE:面向参数与样本高效少样本学习的自适应定位
arXiv:2609.22323 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一种面向 few-shot 图像分类的超轻量级空间-关系架构,结合固定的 Gabor 边缘能量引导与窗口化的、内容自适应的 patch locator。实验表明,该架构中虽然存在 pairwise relational computation,但它并非性能的主要驱动因素;真正起决定作用的是 content-adaptive patch locator。An ultra-lightweight spatial-relational architecture for few-shot image classification that combines fixed Gabor edge-energy guidance with a windowed, content-adaptive patch locator is presented, showing that the architecture's pairwise relational computation, while present, is not the primary driver of its performance; the content-adaptive patch locator is.

JEV-as-a-Judge: Accept When Confident, Escalate When Unsure
JEV-as-a-Judge:自信则接受,不确定则升级
arXiv:2609.26550 评测基准 方法 OA · 绿色 被引 14 · S2

LLM-as-a-judge 支持跨任务评估,但在大规模场景下推理成本与置信度可靠性成为关键问题。本文研究仅做判断的 judge 能否提供经济高效的首轮判别,并识别何时需要更强的评估。在与十六种生成式与奖励模型 judge 的对比中(采用盲法人类裁定作为参照),我们发现 jev-as-a-judge 在普通偏好与有证据支撑的事实性任务上,与作为最强对照的 SOTA LLM judge 仅相差 3 个百分点,成本仅为后者的 0.36%。在需要核查LLM-as-a-judge enables evaluation across diverse tasks, but inference cost and confidence reliability become critical at scale. We study whether a decision-only judge can provide an economical first pass and identify when stronger evaluation is needed. Comparing jev-as-a-judge with sixteen generative and reward-model judges, with blinded human adjudication, we find it within three percentage points of a state-of-the-art LLM judge, our strongest comparator, on ordinary preference and evidence-grounded factuality at 0.36% of the comparator's fee. Larger gaps arise when judgments require checking

StudentBench: AI and human tutoring yield equivalent GRE learning gains
StudentBench:AI 辅导与人类辅导在 GRE 学习成效上相当
arXiv:2609.28470 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

论文证实,AI 辅导在 GRE 学习收益上与专家人类辅导具有统计等效性(p = .015),且在 GRE 的七个领域中,有五个领域的最佳 AI 导师平均超越了人类导师。It is established that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains (p = .015), and in five of the seven GRE domains, the best performing AI tutor surpassed the human tutor, on average.

HappyWorld-Bench
HappyWorld-Bench
arXiv:2609.24308 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 HappyWorld-Bench,一个用于评估生成世界在 Agent 交互过程中是否保持可靠的综合基准,并强调对世界模型的评估不仅应看视觉质量,还应考察状态一致性以及其对动作和干预响应的正确性。HappyWorld-Bench is introduced, a comprehensive benchmark that evaluates whether generated worlds remain reliable as agents interact with them, and highlights the need to evaluate world models not only by visual quality, but also by state consistency and the correctness of their responses to actions and interventions.

FLEET: From Logits Entropy to Enhanced Trajectories in Text Generation
FLEET:从 logits 熵到文本生成中的增强轨迹
arXiv:2609.27657 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

FLEET 将每次生成表示为穿越状态的稀疏轨迹(状态的熵超过预设阈值),并基于这些轨迹推断逐 token 的效用分数以调整 logits,是一种将 memory 机制融入生成过程的新方法。FLEET represents each generation as a sparse trajectory through states whose entropy exceeds a predefined threshold and uses these trajectories to infer per-token utility scores that adjust the logits, a novel method that integrates a memory mechanism into the generation process.

Calibration as a First-Class Criterion in LLM Evaluation
将校准作为 LLM 评估中的一等标准
arXiv:2609.26489 评测基准 评测集 OA · 绿色 被引 1 · S2

本文主张每个 NLP 子领域应将主要性能指标与一个校准分数配对,呼吁将校准视为每个模型的基本属性而非边缘话题。It is argued that each NLP subfield should pair its main performance metric with a calibration score and call for treating calibration as an essential property of every model rather than a niche topic.

Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models
能力出众却简洁高效:提取并刻画前沿模型中隐藏的思维链
arXiv:2609.26637 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果发现 Astra 表现出 token 高效的有向推理,能更早选择正确轨迹,在内部解决基础步骤,仅外化关键推理,为前沿模型推理提供了超越基准分数的行为视角。It is found that Astra exhibits token-efficient directed reasoning, selecting a correct trajectory earlier, while resolving elementary steps internally and externalizing only crucial reasoning, which provides a behavioral lens on frontier-model reasoning beyond benchmark scores.

RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation
RGBD20K:面向 RGB-D 语义分割的大规模基准
arXiv:2609.29028 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一种新颖的 score-purified fusion (SPF) 方法,在所有评测 benchmark 上均达到 SOTA 性能,验证了该方法在利用高质量多模态信息进行 RGB-D 语义分割任务中的有效性。A novel score-purified fusion (SPF) method is proposed, which achieves state-of-the-art performance across all evaluated benchmarks, demonstrating the effectiveness of the approach in leveraging high-quality multimodal information for RGB-D semantic segmentation.

Learning to Discover Interesting Mathematics
学习发现有趣的数学
arXiv:2609.28603 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

近年来,大语言模型(LLMs)解决高级数学问题的能力日益增强,包括许多悬而未决数十年的难题。这为以空前规模扩展数学知识打开了大门。然而,尽管 LLM 能够猜想并证明越来越多的定理,这些新数学知识是否有趣或有用仍属未知。我们将定理的内在有趣度定义为其证明长度与陈述长度之比。证明该指标与下载量的外在度量高度相关。Recently, Large Language Models (LLMs) have been increasingly able to solve advanced mathematical problems, including many that have been open for decades. This opens the door to expansion of mathematical knowledge at unprecedented scale. Yet, while LLMs may be able to conjecture and prove more and more theorems, it remains open whether this new mathematical knowledge is interesting or useful. We define intrinsic interestingness of a theorem as the ratio between the length of its proof and the length of its statement. We show that this correlates strongly with an extrinsic measure of the downs

1. AlphaEval: Evaluating Agents in Production
AlphaEval: 在生产环境中评估 Agent
arXiv:2604.12162 评测基准 评测集 OA · 绿色 被引 1 · S2

本工作提出 AlphaEval,一个基于真实生产环境的基准,包含来自七家在其核心业务中部署 AI Agent 的公司的 94 个任务,覆盖六个 O*NET (Occupational Information Network) 领域;并贡献了一套从需求到基准的构建框架,将从需求到评估的完整流程标准化。This work presents AlphaEval, a production-grounded benchmark of 94 tasks sourced from seven companies deploying AI agents in their core business, spanning six O*NET (Occupational Information Network) domains, and contributes a requirement-to-benchmark construction framework that standardizes the entire pipeline from requirement to evaluation.

Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations
[标题中文] 气候政策因果评估中识别假设的循证审计
arXiv:2609.30867 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

[TLDR中文] 双重差分(DID)研究被广泛用于评估气候政策,但评估支持其识别假设的证据仍具挑战。我们提出 ARGUS,一个结构化的语言模型流水线,针对十一维的假设—含义—证据评估标准对所报告的证据进行审计,并在无法检索到相关证据时选择弃答。我们通过注入缺陷、经济学论文以及一个使用经协调标签的小规模试点评估 ARGUS。在 11 类缺陷基准上,ARGUS 检测出 73% 的植入缺陷,而基于关键词的流水线仅能检测出 18%。Difference-in-differences (DID) studies are widely used to evaluate climate policy, but assessing the evidence supporting their identification assumptions remains challenging. We introduce ARGUS, a structured language-model pipeline that audits reported evidence against an eleven-dimension assumption-implication-evidence rubric and abstains when relevant evidence cannot be retrieved. We evaluate ARGUS using injected flaws, economics papers, and a small pilot with reconciled labels. On the 11-flaw benchmark, ARGUS detects 73% of planted flaws, compared with 18% for a keyword-based pipeline. Acr

IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking
[标题中文] IndicBankBench:评估语言模型助手在印度零售银行中的安全性与可靠性
arXiv:2609.29167 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 IndicBankBench,一个涵盖五个业务领域、一个能力/拒答领域以及二十个主轴、共 799 个案例的零售银行 benchmark,并提供 case 级诊断分析,以区分那些提出不必要追问的系统与那些采取行动却未能调和客户上下文或完整解决诉求的系统。IndicBankBench is introduced, a 799-case benchmark for Indian retail banking spanning five operational domains, a capability/refusal domain, and twenty primary axes, and a case-level diagnostics that distinguish systems that ask unnecessary questions from those that act but fail to reconcile customer context or fully resolve the request.

SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance
SAGE:通过拓扑引导缓解长程推理偏置
arXiv:2609.30192 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 SAGE(Structural Admissibility-Guided Exploration),一个通过注入结构引导来缓解长程推理中探索偏差与累积偏差的统一框架,在 Andrews-Curtis 问题上取得了最高 8 倍的提升。This work proposes SAGE (Structural Admissibility-Guided Exploration), a unified framework that injects structural guidance to alleviate exploration bias and compounding bias in long-horizon reasoning and achieves up to an 8-fold improvement on the Andrews-Curtis problem.

Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents
Mid-Harness:在 Model 与 Harness 之间扩展终端 Agent 的动作规模
arXiv:2609.39982 评测基准 方法 OA · 绿色 被引 1 · S2

本文提出 Mid-Harness,在执行前对候选动作进行采样和验证,同时保持 generator 和 harness 不变,并将 action scaling 识别为终端 Agent 中 test-time compute scaling 的一个有前景的目标。Mid-Harness is introduced, which samples and verifies candidate actions before forwarding one for execution, while keeping the generator and harness unchanged, and identifies action scaling as a promising target for test-time compute scaling in terminal agents.

OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software
OSWorld-Science:面向科学与科研软件学习与使用的计算机操作 Agent 基准
arXiv:2609.39903 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 OSWorld-Science,这是一个结合科学意义任务、基于 artifact 的评估以及高效 agent harness 的基准与评估环境,用于研究科学领域的计算机使用,从而在科学工作流中系统评估 agent 能力与 harness 设计。OSWorld-Science is introduced, a benchmark and evaluation environment that combines scientifically meaningful tasks, artifact-based evaluation, and an efficient agent harness for studying computer use in the scientific domain, enabling systematic evaluation of both agent capabilities and harness design in scientific workflows.

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
DyRAD:面向动态驾驶场景的雷达新视角合成
arXiv:2609.39841 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

DyRAD 使用静态背景反射器和运动追踪的动态点反射器建模动态驾驶场景,渲染完整的距离-方位-多普勒 (RAD) 张量,并通过从雷达信号处理链推导出的固定解析点扩散函数渲染反射器,避免传感器引起的扩散被烘焙到场景表示中。DyRAD is presented, which models dynamic driving scenes using static background reflectors and motion-tracked dynamic point reflectors to render complete range-azimuth-Doppler (RAD) tensors, and renders reflectors through a fixed analytic point-spread function derived from the radar's signal-processing chain, preventing sensor-induced spread from being baked into the scene representation.

The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence
无尽之试:当代模型迈向超智能的数学构造
arXiv:2609.24555 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

我们提出 Endless Exam,一个涵盖十四个参数化数学构造问题族的 benchmark,具有可验证的分数,能够区分在已发表数学前沿之前与之后的进展。每个提交的对象会被自动检验有效性,并依据已发表前沿或构造基线获得相对质量分数,不将改进上限设为 1。该 benchmark 从开放性数学问题中汲取长期挑战,并通过改变参数生成更大规模的实例。紧凑证书使得大型构造能被快速验证We introduce the Endless Exam, a benchmark spanning fourteen parameterised families of mathematical construction problems, with verifiable scores that distinguish progress before and beyond published mathematical frontiers. Each submitted object is checked automatically for validity and assigned a relative quality score against a published frontier or construction baseline, without capping improvements at 1. The benchmark draws long-term challenges from open mathematical problems and generates larger instances by varying their parameters. Compact certificates allow large constructions to be ve