研究库 论文知识库
Papers · organized/paper_cards

论文

138 张论文卡片 · 评测基准 · 评测集

开放获取 全部 绿色 · 1640
MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes
MetroLLM-Bench:将语言模型作为交通售票机运行时进行评估
arXiv:2609.10016 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

引入 MetroLLM-Bench,一个包含 955 个用例的基准,用于测试语言模型作为交通信息亭策略层的能力,并评估了来自六个厂商的 26 个模型,其中 23 个被排名。This work introduces MetroLLM-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk, and evaluates twenty-six models from six vendors, of which twenty-three are ranked.

2️⃣ arXiv · Generating Leakage-Free Benchmarks for Robust RAG Evaluation(⭐⭐⭐⭐⭐ 必读评测方法论)
arXiv · Generating Leakage-Free Benchmarks for Robust RAG Evaluation(⭐⭐⭐⭐⭐ 必读评测方法论)
arXiv:2605.08838 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出了 SeedRG,一个用于缓解 knowledge leakage 并应对 benchmark aging 问题的半合成 benchmark 生成 pipeline。SeedRG is introduced, a semi-synthetic benchmark generation pipeline that mitigates knowledge leakage and addresses the issue of benchmark aging.

IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications
IdeaAMBIG:研究构想规范中面向实现的关键缺口基准
arXiv:2609.10539 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

研究面向实现的研究方法规范的可编码化就绪度,定义为其是否为胜任的实现者或编码 Agent 提供足够的方法学信息,以在不引入未支持假设的情况下构建预期方法。This work studies the codification readiness of implementation-facing research-method specifications, defined by whether they provide sufficient methodological information for a competent implementer or coding agent to construct the intended method without unsupported assumptions.

Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation
Benchmark Radar:面向 AI 基准与评测的活体数据库与搜索引擎
arXiv:2609.11115 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Benchmark Radar,一个用于 AI 基准检索与发现的活体数据库和搜索引擎,覆盖 LLM 评测、Agent 与工具使用基准、代码、推理、安全及领域评测。Benchmark Radar is presented, a living database and search engine for retrieval and discovery of AI benchmarks, covering LLM evaluation, agentic and tool-use benchmarks, coding, reasoning, safety, and domain-specific evaluations.

DataFlex-RL: An Evaluation Platform for RLVR Data Policies
DataFlex-RL:面向 RLVR 数据策略的评测平台
arXiv:2609.06107 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出了 DataFlex-RL,一个在统一 GRPO 方案下比较数据策略的评测平台;研究发现,改变数据策略会显著影响训练过程,但相对于均匀训练并未带来可复现的提升。DataFlex-RL, an evaluation platform for comparing choices under a common GRPO recipe, is introduced, finding that changing the data policy measurably changes the training process but does not produce a reproducible improvement over uniform training.

Feature Recovery for Object Understanding After Irreversible Fire Damage
Feature Recovery for Object Understanding After Irreversible Fire Damage
arXiv:2609.12078 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出即插即用的 Feature Recovery Module (FRM),可在保持宿主网络冻结的前提下,将退化编码器特征映射到与干净图像对齐的表征;该模块提升了场景级检测、CLIP/SigLIP2 特征恢复以及全部四项物体级 VLM 任务,且退化越严重增益越大。The Feature Recovery Module (FRM), a plug-and-play module that maps degraded encoder features to pristine-aligned representations while keeping the host frozen, is proposed, which improves scene-level detection, CLIP/SigLIP2 feature recovery, and all four object-level VLM tasks, with larger gains under more severe degradation.

E2A-Bench: Benchmarking Evidence-to-Action Reliability in Financial Chart Reasoning
E2A-Bench:金融图表推理中证据到行动可靠性的基准测试
arXiv:2609.14302 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

论文提出了 E2A-Bench,一个面向金融图表推理的 969 查询基准,由 323 个 HS300 成分股在三种输入模态下构建,并附带由 OHLCV 确定性派生的证据锚点;结果表明金融 VLM 评估应追溯从证据到决策的完整链路,而非依赖单一幻觉分数。E2A-Bench is introduced, a 969-query benchmark for financial chart reasoning, constructed from 323 HS300 constituents under three input modalities with deterministic OHLCV-derived evidence anchors, and results show that financial VLM evaluation should trace the full evidence-to-action chain rather than rely on a single hallucination score.

2. Systemic Measurement Bias in LLM Inference Benchmarking
LLM Inference Benchmarking 中的系统性测量偏差
arXiv:2605.24217 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一个无偏的多进程 evaluation 框架,能够有效分散 client 端负载,从而在每秒数千次 query 以上的生产规模下实现对 LLM 的精确、可复现 profiling。This work proposes an unbiased, multi-process evaluation framework that effectively distributes client-side load, enabling accurate, reproducible profiling of LLMs at production scales exceeding thousands of queries per second.

ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals
ImpossibleRubrics:对生成式评分标准作为奖励信号的应激测试
arXiv:2609.16816 评测基准 评测集 被引 0 · S2

论文提出了 ImpossibleRubrics,一个包含 169 个不可能任务的基准,覆盖六类不可能性类别,每项任务都配有可验证的 oracle 证书,规定诚实回答可以与不可以声明的内容,并附带 48 个可回答的对照样本。ImpossibleRubrics is introduced, a benchmark of 169 impossible tasks spanning six impossibility categories, each paired with a verifiable oracle certificate specifying what an honest answer may and may not claim, together with 48 answerable controls.

ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement
ModularRSI:模块化且可泛化的递归框架自我改进
arXiv:2609.14857 评测基准 评测集 OA · 绿色 被引 6 · S2

论文提出了 ModularRSI,一个与基准解耦、对比式、模块化的可泛化 harness 进化框架,对同一任务下成功与失败的轨迹进行对比,并跨任务聚合证据以识别反复出现的行为缺陷。ModularRSI is proposed, a benchmark-disjoint, contrastive, and modular framework for generalizable harness evolution that contrasts successful and failed trajectories for the same task and aggregates evidence across tasks to identify recurring behavioral deficiencies.

Another Blueprint In The Wall: How to Ask Frontier AI Like a Kid?
墙上又一张蓝图:如何像孩子一样向前沿AI提问?
arXiv:2609.14803 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

论文报告了来自 OpenAI、Anthropic、xAI 与 Google DeepMind 的六类前沿模型实验,并使用 epistemic jailbreak 一词来指称随请求具体性增加而伴随出现的技术溯源严谨性丧失现象。This paper reports experiments across six frontier model types from OpenAI, Anthropic, xAI, and Google DeepMind, and uses the term epistemic jailbreak for the accompanying loss of discipline in technical provenance as requested specificity increases.

1️⃣1️⃣ arXiv · RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic RAG Systems(⭐⭐⭐ 参考)
arXiv · RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic RAG Systems(⭐⭐⭐ 参考)
arXiv:2510.13910 评测基准 评测集 被引 5 · S2

本文提出 RAGCap-Bench,一个面向能力的 benchmark,用于对 agentic RAG workflow 中的中间任务进行细粒度评测,并构建了典型 LLM 错误的分类体系以设计针对性评测问题。This work proposes RAGCap-Bench, a capability-oriented benchmark for fine-grained evaluation of intermediate tasks in agentic RAG workflows, and constructs a taxonomy of typical LLM errors to design targeted evaluation questions.

VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering
VākQA:泰卢固语口语事实型问答基准与评估研究
arXiv:2609.19879 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 V\={a}kQA,一个覆盖六个领域、包含 2,001 对事实型问答的泰卢固语 SQA 基准;并观察到:泰卢固语措辞保留了翻译中会丢失的文化特异性,语音输入引入的音近混淆会改变问题含义,级联 ASR-MT 误差会逐步叠加放大。V\={a}kQA, a Telugu SQA benchmark of 2,001 factoid question-answer pairs across six domains, is introduced and it is observed that Telugu phrasing retains cultural specificity that is lost in translation, speech input introduces phonetic confusions that alter question meaning, and cascaded ASR-MT errors compound progressively.

TeleAntiFraud 2.0: A Refreshable, Profile-Grounded, and Audio-Based Benchmark for Telecom Fraud Detection
TeleAntiFraud 2.0:面向电信诈骗检测的可刷新、基于画像的音频基准
arXiv:2609.18748 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 TeleAntiFraud 2.0,采用 Mixed-Tree Anti-Fraud Generation Pipeline 构建,并基于月度冻结评测协议进行评估,确立了近域构造与 collapse-aware 报告作为在现实易混淆条件下评测音频电信诈骗模型的核心要求。This work presents TeleAntiFraud 2.0, constructed with the Mixed-Tree Anti-Fraud Generation Pipeline and evaluated under a monthly frozen evaluation protocol, establishing near-domain construction and collapse-aware reporting as core requirements for evaluating audio-based telecom-fraud models under realistic confusable conditions.

Gricea: An Open Science Platform for Conversational AI Research
Gricea:一个面向对话式 AI 研究的开放科学平台
arXiv:2609.22039 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Gricea,一个开放科学平台,将研究表示为可配置、可部署的研究 artifact,研究者可以运行、检查、共享和复用;该平台展示了如何通过共享的研究 artifact 构建、复现和扩展 CAI 研究,从而借助开放科学实现知识的累积构建。This work presents Gricea, an open-science platform representing studies as configurable, deployable research artifacts that researchers can run, inspect, share, and reuse, demonstrating Gricea's support for constructing, reproducing, and extending CAI studies through shared research artifacts, enabling cumulative knowledge building through open science.

CADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design
CADWorld:面向长视野计算机辅助设计的计算机使用基准
arXiv:2609.16251 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

研究发现,较弱的 Agent 常常在产出有效成果之前即告失败,而较强的 Agent 则越来越多地在结构、几何与构造过程要求上失败;CADWorld 揭示了通用 GUI 能力与可靠执行持久、可验证工程工作流之间的差距。It is found that weaker agents often fail before producing a valid artifact, whereas stronger agents increasingly fail on structural, geometric, and construction-process requirements, and CADWorld exposes a gap between general GUI competence and reliable execution of persistent, verifiable engineering workflows.

GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay
GameHorizon Suite:游戏中多时间历程的数据与评估。
arXiv:2609.25001 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

GameHorizon Suite 是一套统一的数据与评估套件,可在多个时间跨度下衡量不同模型族的游戏能力,为跨水平跨度与跨模型族的游戏能力评估提供标准化标尺。The GameHorizon Suite, a unified data and evaluation suite that measures gameplay capabilities at different horizons for diverse model families, can provide a standardized yardstick for evaluating gameplay capabilities across horizons and model families.

Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles
Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles
arXiv:2609.22220 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出将变异分析作为 kernel-benchmark 预言机的充分性度量:通过确定性规则向 188 个 KernelBench 问题的已验证 CUDA 实现中注入 10 个可编译故障,其中 7,384 个具备独立击杀见证;任何测试协议均按其检出比例评分。Mutation analysis as an adequacy metric for kernel-benchmark oracles is introduced: deterministic rules inject 10 compilable faults into verified CUDA implementations of 188 KernelBench problems, 7{,}384 of them with an independent kill witness; any test protocol is scored by the fraction it detects.

HappyWorld-Bench
HappyWorld-Bench
arXiv:2609.24308 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 HappyWorld-Bench,一个用于评估生成世界在 Agent 交互过程中是否保持可靠的综合基准,并强调对世界模型的评估不仅应看视觉质量,还应考察状态一致性以及其对动作和干预响应的正确性。HappyWorld-Bench is introduced, a comprehensive benchmark that evaluates whether generated worlds remain reliable as agents interact with them, and highlights the need to evaluate world models not only by visual quality, but also by state consistency and the correctness of their responses to actions and interventions.

Calibration as a First-Class Criterion in LLM Evaluation
将校准作为 LLM 评估中的一等标准
arXiv:2609.26489 评测基准 评测集 OA · 绿色 被引 1 · S2

本文主张每个 NLP 子领域应将主要性能指标与一个校准分数配对,呼吁将校准视为每个模型的基本属性而非边缘话题。It is argued that each NLP subfield should pair its main performance metric with a calibration score and call for treating calibration as an essential property of every model rather than a niche topic.

RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation
RGBD20K:面向 RGB-D 语义分割的大规模基准
arXiv:2609.29028 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一种新颖的 score-purified fusion (SPF) 方法,在所有评测 benchmark 上均达到 SOTA 性能,验证了该方法在利用高质量多模态信息进行 RGB-D 语义分割任务中的有效性。A novel score-purified fusion (SPF) method is proposed, which achieves state-of-the-art performance across all evaluated benchmarks, demonstrating the effectiveness of the approach in leveraging high-quality multimodal information for RGB-D semantic segmentation.

1. AlphaEval: Evaluating Agents in Production
AlphaEval: 在生产环境中评估 Agent
arXiv:2604.12162 评测基准 评测集 OA · 绿色 被引 1 · S2

本工作提出 AlphaEval,一个基于真实生产环境的基准,包含来自七家在其核心业务中部署 AI Agent 的公司的 94 个任务,覆盖六个 O*NET (Occupational Information Network) 领域;并贡献了一套从需求到基准的构建框架,将从需求到评估的完整流程标准化。This work presents AlphaEval, a production-grounded benchmark of 94 tasks sourced from seven companies deploying AI agents in their core business, spanning six O*NET (Occupational Information Network) domains, and contributes a requirement-to-benchmark construction framework that standardizes the entire pipeline from requirement to evaluation.

Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations
[标题中文] 气候政策因果评估中识别假设的循证审计
arXiv:2609.30867 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

[TLDR中文] 双重差分(DID)研究被广泛用于评估气候政策,但评估支持其识别假设的证据仍具挑战。我们提出 ARGUS,一个结构化的语言模型流水线,针对十一维的假设—含义—证据评估标准对所报告的证据进行审计,并在无法检索到相关证据时选择弃答。我们通过注入缺陷、经济学论文以及一个使用经协调标签的小规模试点评估 ARGUS。在 11 类缺陷基准上,ARGUS 检测出 73% 的植入缺陷,而基于关键词的流水线仅能检测出 18%。Difference-in-differences (DID) studies are widely used to evaluate climate policy, but assessing the evidence supporting their identification assumptions remains challenging. We introduce ARGUS, a structured language-model pipeline that audits reported evidence against an eleven-dimension assumption-implication-evidence rubric and abstains when relevant evidence cannot be retrieved. We evaluate ARGUS using injected flaws, economics papers, and a small pilot with reconciled labels. On the 11-flaw benchmark, ARGUS detects 73% of planted flaws, compared with 18% for a keyword-based pipeline. Acr

IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking
[标题中文] IndicBankBench:评估语言模型助手在印度零售银行中的安全性与可靠性
arXiv:2609.29167 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 IndicBankBench,一个涵盖五个业务领域、一个能力/拒答领域以及二十个主轴、共 799 个案例的零售银行 benchmark,并提供 case 级诊断分析,以区分那些提出不必要追问的系统与那些采取行动却未能调和客户上下文或完整解决诉求的系统。IndicBankBench is introduced, a 799-case benchmark for Indian retail banking spanning five operational domains, a capability/refusal domain, and twenty primary axes, and a case-level diagnostics that distinguish systems that ask unnecessary questions from those that act but fail to reconcile customer context or fully resolve the request.

OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software
OSWorld-Science:面向科学与科研软件学习与使用的计算机操作 Agent 基准
arXiv:2609.39903 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 OSWorld-Science,这是一个结合科学意义任务、基于 artifact 的评估以及高效 agent harness 的基准与评估环境,用于研究科学领域的计算机使用,从而在科学工作流中系统评估 agent 能力与 harness 设计。OSWorld-Science is introduced, a benchmark and evaluation environment that combines scientifically meaningful tasks, artifact-based evaluation, and an efficient agent harness for studying computer use in the scientific domain, enabling systematic evaluation of both agent capabilities and harness design in scientific workflows.

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
DyRAD:面向动态驾驶场景的雷达新视角合成
arXiv:2609.39841 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

DyRAD 使用静态背景反射器和运动追踪的动态点反射器建模动态驾驶场景,渲染完整的距离-方位-多普勒 (RAD) 张量,并通过从雷达信号处理链推导出的固定解析点扩散函数渲染反射器,避免传感器引起的扩散被烘焙到场景表示中。DyRAD is presented, which models dynamic driving scenes using static background reflectors and motion-tracked dynamic point reflectors to render complete range-azimuth-Doppler (RAD) tensors, and renders reflectors through a fixed analytic point-spread function derived from the radar's signal-processing chain, preventing sensor-induced spread from being baked into the scene representation.

The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence
无尽之试:当代模型迈向超智能的数学构造
arXiv:2609.24555 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

我们提出 Endless Exam,一个涵盖十四个参数化数学构造问题族的 benchmark,具有可验证的分数,能够区分在已发表数学前沿之前与之后的进展。每个提交的对象会被自动检验有效性,并依据已发表前沿或构造基线获得相对质量分数,不将改进上限设为 1。该 benchmark 从开放性数学问题中汲取长期挑战,并通过改变参数生成更大规模的实例。紧凑证书使得大型构造能被快速验证We introduce the Endless Exam, a benchmark spanning fourteen parameterised families of mathematical construction problems, with verifiable scores that distinguish progress before and beyond published mathematical frontiers. Each submitted object is checked automatically for validity and assigned a relative quality score against a published frontier or construction baseline, without capping improvements at 1. The benchmark draws long-term challenges from open mathematical problems and generates larger instances by varying their parameters. Compact certificates allow large constructions to be ve

Benchmarking and Enhancing Skill-Level Memory for Partially Observable Robotic Manipulation
在部分可观测机器人操作任务上对技能级记忆的基准评测与增强
arXiv:2609.38886 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 HIDE,一个用于在部分可观测条件下评估操作记忆的基准,并提出 SEEK 框架,结合三种互补的记忆机制来保留历史证据并追踪执行状态。This work introduces HIDE, a benchmark for evaluating manipulation memory under partial observability, and proposes $SEEK, a framework combining three complementary memory mechanisms to retain historical evidence and track execution state, and proposes a framework combining three complementary memory mechanisms to retain historical evidence and track execution state.

Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs
泛化即稳定性,而非准确率:LLM 的多轴评估
arXiv:2610.01428 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出稳定性感知的泛化目标(SAGO),一个用于衡量模型在同一输入上面对不同扰动和基准时行为变化程度的评估框架,涵盖生成一致性、内部激活、置信度以及响应镜像等多个维度的变异性。The Stability-Aware Generalization Objective (SAGO) is introduced, a framework that measures how much model behavior changes for the same input under different variations and benchmarks, capturing variability across several dimensions including generation consistency, internal activations, confidence, and response mirroring.

PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop
PhysVista:通过感知-推理-评估闭环评测 VLM 的物理智能
arXiv:2610.00559 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

PhysVista 是一个旨在通过借鉴人类“感知-推理-评估”过程的认知闭环框架来评测 VLM 物理智能的基准,揭示了视觉识别与真实物理理解之间持续存在的差距,并为面向物理基础的多模态智能设计提供了更具原则性的方向。PhysVista is a benchmark designed to evaluate physical intelligence in VLMs through a closed cognitive loop framework inspired by the human seeing-reasoning-assessment process, highlighting a persistent gap between visual recognition and genuine physical understanding, and pointing toward more principled designs for physically grounded multimodal intelligence.

OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories
OpenTumorBoard:多学科肿瘤委员会讨论轨迹的真实世界基准
arXiv:2609.32810 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出了 OpenTumorBoard,一个包含 611 个患者案例、19,157 轮讨论、涵盖十个专科角色的基准,源自 YouTube 上 12,534 分钟公开肿瘤委员会会议录音的转写,并发布了自动化整理流水线,以支持多学科、个性化癌症决策中 LLM 的开发与评估。OpenTumorBoard is introduced, a benchmark with 611 patient cases and 19,157 discussion turns across ten specialist roles, transcribed from 12,534 minutes of publicly available tumor board recordings on YouTube, and automated curation pipeline is released to support the development and evaluation of LLMs for multidisciplinary, personalized cancer decision-making.

SimuVerity: Benchmarking Agents for Engineering-Grade Simulink Model Generation
SimuVerity:面向工程级 Simulink 模型生成的智能体基准
arXiv:2610.02304 评测基准 评测集 被引 0 · S2

SimuVerity 为评估智能体在可执行 Simulink 模型生成中的工程能力与诊断失败提供了系统性基础,并表明结构相似性是衡量工程性能的一个糟糕代理指标。SimuVerity provides a systematic basis for assessing agents' engineering capabilities and diagnosing failures in executable Simulink model generation, and shows that structural similarity is a poor proxy for engineering performance.

4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes
4DCodeBench:动态场景逆向图形中的 Agent 基准测试
arXiv:2610.03715 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 4DCodeBench,一个面向 4D 逆向图形(通过代码生成)的基准,Agent 以可执行图形程序的形式从视频重建动态场景,结果表明强大的静态重建能力尚未转化为对复杂动态的可靠重建。4DCodeBench is introduced, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs, finding that strong static reconstruction capabilities do not yet translate into reliable reconstruction of complex dynamics.

The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models
缺失的原语:诊断与修复大语言模型中的数学推理
arXiv:2610.02191 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

引入 Mathematical Primitive 概念来探查结构性数学理解,并提出 \hlei{},一个沿发现(Discovery)、生成(Generation)、消化(Digestion)和执行(Execution)四个维度评估数学推理的新型基准。This paper introduces the notion of Mathematical Primitive to probe structural mathematical understanding and proposes \hlei{}, a novel benchmark that evaluates mathematical reasoning along four distinct dimensions: Discovery, Generation, Digestion, and Execution.

Technical Report on the Turba Fertilizer Machine Learning Stack in Morocco
摩洛哥 Turba 施肥机器学习技术报告
arXiv:2610.05949 评测基准 评测集 OA · 绿色 被引 0 · OpenAlex

场地特异性施肥推荐系统会根据地点、土壤属性、作物类型和生产目标调整养分建议,但当推荐功能主要通过交互式界面访问、输出未做版本管理、训练后的近似模型无法独立加载或基准测试时,其科学复用性受到限制。本技术报告介绍 Turba 施肥机器学习技术栈——面向摩洛哥场地特异性施肥推荐的可复现三层开源实现,其中 turba-client 提供对公开……Site-specific fertilizer recommendation systems adapt nutrient advice to location, soil properties, crop type, and production targets, but scientific reuse is constrained when recommendation functions remain accessible mainly through interactive interfaces, outputs are not versioned, and trained approximations cannot be independently loaded or benchmarked. This technical report presents the Turba fertilizer machine learning stack, a three-layer open-source implementation for reproducible site-specific fertilizer recommendation in Morocco. turba-client provides programmatic access to publicly a

RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies
RoboDojo:面向通用机器人操作策略综合评估的仿真-真机统一基准
arXiv:2607.04434 评测基准 评测集 OA · 绿色 被引 44 · S2

提出 RoboDojo,一个面向通用机器人操作策略综合评估的仿真-真机统一基准,将 30 种策略集成到 XPolicyLab 并在 RoboDojo 上进行评测,建立了公开的排行榜与系统性的策略性能分析。RoboDojo is introduced, a unified sim-and-real benchmark for comprehensive evaluation of generalist robot manipulation policies that integrates 30 policies into XPolicyLab and evaluates them on RoboDojo, establishing a public leaderboard and systematic analysis of current policy performance.