研究库 论文知识库
Papers · organized/paper_cards

论文

224 张论文卡片 · 评测基准

开放获取 全部 绿色 · 1640
9️⃣ arXiv · Benchmarking Multimodal Memory for Realistic User-Agent Interactions(M3Exam)(⭐⭐⭐ 参考)
9️⃣ arXiv · 面向真实用户-Agent 交互的多模态记忆基准测试(M3Exam)(⭐⭐⭐ 参考)
arXiv:2606.07402 评测基准 评测集 OA · 绿色 被引 2 · S2

本文提出 M$^3$Exam,一个以查询为中心、基于真实用户-Agent 交互构建的多模态对话记忆基准,涵盖跨模态定位与隐式信息推断等多维度评估。M$^3$Exam is introduced, a query-centric multimodal conversational memory benchmark built on realistic user-agent interaction, with multi-dimensional evaluation spanning cross-modal grounding and implicit information inference.

1. Recursive Agent Harnesses (RAH)
1. Recursive Agent Harnesses(RAH)
arXiv:2606.13643 评测基准 方法 OA · 绿色 被引 5 · S2

本文命名并研究这两条研究脉络之间的模式:其递归单元是配备文件系统工具、代码执行与规划的完整 Agent harness,而非无工具的模型调用,并给出针对长上下文推理的受控评估。This work names and studies the pattern between these two lines of work, where the recursive unit is a full agent harness with filesystem tools, code execution, and planning rather than a model call with no tools, and provides a controlled evaluation on long-context reasoning.

5.2 ForeSci:研究判断型 agent 评测
arXiv:2606.00644 评测基准 评测集 OA · 绿色 被引 1 · S2

本文提出 ForeSci,一个时间受控的基准,用于评估 LLM Agent 是否能从历史证据中做出前瞻性研究判断,并在四种骨干模型上评测原生 LLM、Hybrid RAG 以及三种 research-agent 适配方案。ForeSci is introduced, a temporally controlled benchmark for evaluating whether LLM agents can make such forward-looking research judgements from historical evidence, and agent-based methods improve traceability over Hybrid RAG, while their gains in future-target alignment over native LLMs are modest and task dependent.

MMLongEmbed: 多模态嵌入模型长上下文基准测试
arXiv:2606.14747 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 MMLongEmbed,首个面向长上下文场景评估 MEM 的综合基准,并发现现有架构严重依赖浅层特征匹配,难以捕捉深层语义与结构依赖。This work introduces MMLongEmbed, the first comprehensive benchmark for evaluating MEMs in long-context scenarios, and finds that current architectures rely heavily on superficial feature matching and struggle to capture deep semantic and structural dependencies.

🟡 保留 4:"The Last Harness" — Meta-Evolution 双层循环
arXiv:2604.21003 评测基准 方法 OA · 绿色 被引 5 · S2

一个两级框架将手动 harness 工程转变为自动化 harness 工程,并更进一步——将"自动化本身的设计"也自动化。A two-level framework shifts manual harness engineering into automated harness engineering, and takes one step further --automating the design of the automation itself.

🔴 保留 3:Agentic Harness Engineering (AHE) — arXiv 实证论文
🔴 保留 3:Agentic Harness Engineering(AHE)— arXiv 实证论文
arXiv:2604.25850 评测基准 方法 OA · 绿色 被引 111 · S2

提出 Agentic Harness Engineering(AHE),一个通过三个相互匹配的 observability 支柱应对 harness 工程挑战的闭环,将每一次编辑转化为可证伪的契约,使 harness 演进能够自主进行而不退化为试错。Agentic Harness Engineering (AHE) is introduced, a closed loop that addresses harness engineering challenges through three matched observability pillars that turn every edit into a falsifiable contract, so harness evolution proceeds autonomously without collapsing into trial-and-error.

Agent runtime / security / harness 补充候选
Agent runtime / security / harness 补充候选
arXiv:2603.25723 评测基准 方法 OA · 绿色 被引 53 · S2

本文提出 Natural-Language Agent Harnesses,即可编辑的、描述运行级 harness 策略的文档,以及 Intelligent Harness Runtime(IHR),一个将上述文档解释为 agent 调用、交接、状态更新、验证门控与 artifact 契约的共享运行时。This paper introduces Natural-Language Agent Harnesses, editable documents that describe run-level harness policy, and Intelligent Harness Runtime (IHR), a shared runtime that interprets these documents into agent calls, handoffs, state updates, validation gates, and artifact contracts.

6. Evaluation and Benchmarking of LLM Agents: A Survey
LLM Agent 的评估与基准测试:综述
arXiv:2507.21504 评测基准 综述 KDD 2025 被引 217 · S2

本文对 LLM agent 评估这一新兴领域进行了深入综述,提出一个二维分类体系,沿评估目标维度组织已有工作,为系统性评估提供框架,使研究者与从业者能够面向真实场景部署评估 LLM agent。An in-depth overview of the emerging field of LLM agent evaluation is provided, introducing a two-dimensional taxonomy that organizes existing work along evaluation objectives and provides a framework for systematic assessment, enabling researchers and practitioners to evaluate LLM agents for real-world deployment.

🔴 保留 · `Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Benchmarking`
🔴 保留 · `Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Benchmarking`
arXiv:2606.10749 评测基准 评测集 OA · 绿色 被引 5 · S2

文中指出,安全的 LLM Agent 需要显式的信任边界、原则化的权限控制、具备溯源能力的 state 管理,以及与真实运行场景对齐的评估实践;现有 benchmark 仍未能充分覆盖长程、具状态、对部署敏感的风险。It is argued that secure LLM agents require explicit trust boundaries, principled privilege control, provenance-aware state management, and evaluation practices aligned with realistic operational settings, as well as existing benchmarks still underrepresent long-horizon, stateful, and deployment-sensitive risks.

🔴 保留 · `DeNovoSWE: Scaling Long-Horizon Environments for Generating Entire Repositories from Scratch`
🔴 保留 · `DeNovoSWE: Scaling Long-Horizon Environments for Generating Entire Repositories from Scratch`
arXiv:2606.10728 评测基准 评测集 OA · 绿色 被引 6 · S2

在 DeNovoSWE 上对 Qwen3-30B-A3B 进行微调可显著提升长程 SWE 性能,在具有挑战性的 BeyondSWE-Doc2Repo benchmark 上将其得分从 5.8% 提升至 47.2%。Fine-tuning Qwen3-30B-A3B on DeNovoSWE substantially improves long-horizon SWE performance, raising its score on the challenging BeyondSWE-Doc2Repo benchmark from 5.8% to 47.2%.

🔴 保留 · `Agent Skill Evaluation and Evolution: Frameworks and Benchmarks`
🔴 保留 · `Agent Skill Evaluation and Evolution: Frameworks and Benchmarks`
arXiv:2606.11435 评测基准 综述 OA · 绿色 被引 7 · S2

本综述系统梳理了超越基础 Skill 创建的 Skill 演化与评估图景,将其归纳为四种范式:执行反馈、轨迹蒸馏、压缩与强化学习,并指出了构建可泛化、高效且可验证安全的 Skill 生态的开放方向。This survey systematically examines the landscape of skill evolution and evaluation beyond foundational skill creation into four distinct paradigms, spanning execution feedback, trajectory distillation, compression, and reinforcement learning, and identifies open directions for building skill ecosystems that are generalizable, efficient, and verifiably safe.

条目D3:UnWeaving GraphRAG — GraphRAG vs VectorRAG 理论分析(arXiv 2603.29875v3)
条目D3:UnWeaving GraphRAG — GraphRAG vs VectorRAG 理论分析(arXiv 2603.29875v3)
arXiv:2603.29875 评测基准 观点 OA · 绿色 被引 0 · S2 + OpenAlex

文章认为基于实体的分解能形成对原始信息更精炼的表示,并有助于降低索引与生成过程中的噪声;在端到端 QA 评测中,VectorRAG 表现优于标准 GraphRAG,且接近当前 SOTA 图方法的效果。It is argued that entity-based decomposition yields a more distilled representation of original information, and additionally serves to reduce noise in the indexing, and generation process, and on end to end QA evaluation VectorRAG performs better than standard GraphRAG and almost as good as current SOTA graph-based solutions.

条目A1:EvoArena + EvoMem — 动态环境下的LLM Agent记忆演进基准(arXiv:2606.13681)
arXiv:2606.13681 评测基准 评测集 OA · 绿色 被引 4 · S2

介绍 EvoArena 基准套件,将环境变化建模为跨终端、软件与社会领域的渐进式更新序列;并提出 EvoMem,一种基于 patch 的记忆范式,将记忆演化记录为结构化的更新历史,使 Agent 能通过记忆的变化推理环境的演化。EvoArena is introduced, a benchmark suite that models environment changes as sequences of progressive updates across terminal, software, and social domains, and EvoMem is proposed, a patch-based memory paradigm that records memory evolution as structured update histories, enabling agents to reason about environmental evolution through changes in their memory.

Accuracy and Order Sensitivity Diverge Under Label-Free Strategies
Accuracy and Order Sensitivity Diverge Under Label-Free Strategies
arXiv:2608.11947 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文测试在模型作答时阻止其看到选项标签能否消除位置影响并进而提升性能,并评估了两种不同的偏置缓解策略。This paper test whether preventing a model from seeing option labels while committing to an answer removes positional influence and, in turn, improves performance, and evaluates two different strategies for mitigating bias.

Do Large Language Models Play Six Degrees of Separation? Measuring Topological Compression in Long-Context Manifolds
大语言模型是否在玩"六度分隔"?长上下文流形中的拓扑压缩度量
arXiv:2608.17950 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作数学形式化了 Transformer 如何执行抽象推理,并提出一种新颖的严格几何签名用于评估事实可靠性,证明了深度 LLM 潜空间天然组织为小世界网络。This work mathematically formalizes how transformers execute abstract reasoning and provides a novel, strictly geometric signature for evaluating factual reliability, proving that deep LLM latent spaces natively organize into Small-World networks.

HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety
HarnessRisk:面向 Agent Harness 全生命周期安全的基准
arXiv:2608.17597 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

HarnessRisk 是一个面向生命周期的 benchmark,将 agent harness 安全组织为六个运行阶段,包括 Harness Configuration、Capability Extension、Runtime Operation、State Persistence、Action Control 和 Incident Recovery,发现显式的风险识别并不能可靠地带来安全的行动——某些配置在超过 90% 的运行中检测到风险,同时仍保留显著的攻击成功率。HarnessRisk, a lifecycle oriented benchmark that organizes agent harness safety into six operational phases including Harness Configuration, Capability Extension, Runtime Operation, State Persistence, Action Control, and Incident Recovery, finds that explicit risk recognition does not reliably lead to safe action as some configurations detect risks in more than 90% of runs while retaining substantial attack success.

PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX
PTXBench:面向 GPU kernel 优化与架构特定 PTX 的 LLM 基准与适配
arXiv:2608.17379 评测基准 评测集 OA · 绿色 被引 1 · S2

PTXBench 提供了一个可审计的测试平台,用于衡量并提升 LLM 利用持续演进 GPU 架构的能力,并表明各 LLM 在架构特定 PTX 能力上仍参差不齐。PTXBench provides an auditable testbed for measuring and improving LLMs'ability to exploit evolving GPU architectures, and shows that architecture-specific PTX capability remains uneven.

SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation
SoftVTBench:面向可形变物体操作的形变感知视触觉数据集与基准
arXiv:2608.18701 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

实验表明,仅提供触觉本身并不能确保有效的多模态融合,SoftVTBench 为研究策略不仅能否成功,还在于其如何与可形变物体物理交互,以及触觉在何时改善这种交互,提供了统一的视触觉资源Results show that making touch available does not by itself ensure effective multimodal fusion, and SoftVTBench provides a common visuo-tactile resource for studying not only whether a policy succeeds, but how it physically interacts with deformable objects and when touch improves that interaction.

Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence
Zetta ζ:面向自进化物理智能的高效闭环具身 Harness
arXiv:2608.16590 评测基准 方法 OA · 绿色 被引 9 · S2

本文提出 Zetta,一种闭环具身 harness,在保持基础策略冻结的同时在线演化基于代码的运行时评判器与恢复技能,表明闭环 harness 的自进化为可靠的物理智能开辟了一条可扩展的路径Zetta is presented, a closed-loop embodied harness that evolves code-based runtime critics and recovery skills online while keeping the base policy frozen, and shows that closed-loop harness self-evolution opens a scaling path for reliable physical intelligence.

Evaluating Music Context Preservation: A Multi-facet Framework for Music Editing Systems
音乐上下文保留评估:面向音乐编辑系统的多面框架
arXiv:2512.14629 评测基准 应用落地 OA · 绿色 被引 1 · S2

提出首个 MuseCP 评估框架,涵盖四类音乐 facet,使用细粒度且量身定制的指标来捕捉音乐属性的细微变化,并希望为开发更有效、更可靠、具备强大 MuseCP 能力的音乐编辑策略提供实践指导。The first MuseCP evaluation framework is introduced that covers four categories of music facets with fine-grained and well-tailored metrics to capture nuanced changes in music attributes and hopes it can offer practical guidance for developing more effective and reliable music editing strategies with strong MuseCP capability.

Towards Quantifying Benchmark Optimization in ASR Models
迈向 ASR 模型中基准过拟合的量化
arXiv:2608.19936 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一种量化基准优化的方法论,聚焦于音频对参考转写不充分确定的情形,指出高性能模型会表现出基准条件化行为,从而虚高基准得分,却未必反映通用转写能力的真正提升。This work presents a methodology for quantifying benchmark optimization, focusing on cases where the audio underdetermines the reference transcript, and indicates that high-performing models exhibit benchmark-conditioned behaviors that can inflate benchmark performance without reflecting improved general-purpose transcription ability.

Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See
低资源语言下的思考:SFT 构建什么,RL 修复什么,准确率看不到什么
arXiv:2608.17744 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

选取三个前沿混合专家模型在低资源语言上进行推理微调,提出六个可度量的行为维度,且每维度均设门拒绝任何与输出长度相关的指标,并报告其自家评测工具为何失效。Take three frontier mixture-of-experts models and fine-tune them to reason in a low-resource language and propose six behavioural dimensions that make changes measurable, each gated to reject any metric that correlates with output length, and report how their own instruments lied.

NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video
NARU:面向日语超长视频中叙事演化与文化细微理解 benchmark
arXiv:2608.13210 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 NARU,一个用于评估日语长视频中叙事演进和文化理解推理能力的基准;该工作提出一种基于分层记忆的标注流水线,可将原始视频转换为结构化的事件、叙事和文化标注,并通过任务导向合成与迭代式捷径去除来生成问题。NARU, a benchmark designed to evaluate Narrative evolution and Reasoning on cultural Understanding in Japanese long-form video, is introduced, a hierarchical memory-based annotation pipeline that transforms raw video into structured event, narrative, and cultural annotations, then generates questions via task-oriented synthesis and iterative shortcut removal.

TinyCast: Probabilistic Zero-Shot Forecasting with Computed Periodicity
TinyCast:基于计算周期性的概率性零样本预测
arXiv:2608.15767 评测基准 方法 OA · 绿色 被引 1 · S2

我们提出 TinyCast,一个注意力无关的零样本预测器,仅用 146,505 个参数输出预测分布,其前提是在该规模下,上下文中的周期结构值得通过计算而非学习方式得到。一个零参数谱检测器给出主导周期,上下文按其相位进行折叠,再由一个膨胀卷积编码器和一个分块自回归分位数解码器建模其余部分。它在 GIFT-Eval 榜单上所有可确认参数量的零样本条目中体积最小;在概率准确性方面,它划定了 size-accuracy 前沿。We introduce TinyCast, an attention-free zero-shot forecaster that emits a predictive distribution from 146,505 parameters, on the premise that at this size the periodic structure of a context is worth computing rather than learning. A zero-parameter spectral detector supplies the dominant periods, the context is folded on their phase, and a dilated convolutional encoder and a block-autoregressive quantile decoder model the rest. It is smaller than every zero-shot entry on the GIFT-Eval board whose parameter count can be established. On probabilistic accuracy it defines the size-accuracy front

Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses
分层自改进:面向任务特定可进化 Agent Harness 的框架
arXiv:2608.08466 评测基准 应用落地 OA · 绿色 被引 5 · S2

结果表明,任务特定的 harness 进化是改进冻结 LLM Agent 的一条可行路径,但其效果存在明确的经验性边界。Results demonstrate task-specific harness evolution as a viable axis for improving frozen LLM agents under clear empirical limits.

9️⃣ arXiv · 下一代云原生内存数据库:从 Redis 到 Valkey ⭐⭐⭐⭐⭐ 必读评测
arXiv:2510.19805 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本研究对新兴的内存键值存储进行了全面的性能与可行性评估,突出了性能、兼容性与长期可行性(包括项目成熟度、社区支持与持续开发)之间的权衡。This study presents a comprehensive performance and viability assessment of the emerging in-memory key-value stores and highlights trade-offs between performance, compatibility, and long-term viability, including project maturity, community support, and sustained development.

Benchmarking Patent Drafting from Inventor-Style Disclosures
基于发明人风格 disclosure 的专利起草 benchmark
arXiv:2608.21249 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 Dis2Pat,一个从披露到专利(disclosure-to-patent)的数据集,其设计贴合真实专利工作流,要求直接从发明人风格的、去法律化的披露生成完整专利申请;并提出强基线 Patent-MAF,一个可在本地部署、面向专利撰写的多 Agent 框架。Dis2Pat is introduced, a disclosure-to-patent dataset that reflects realistic patenting workflows by requiring the generation of complete patent applications directly from inventor-style, de-legalized disclosures and a strong baseline named Patent-MAF is proposed, a multi-agent framework for locally deployable patent drafting.

FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth
FlavourBench:基于可执行烹饪真值的 frontier 语言模型排名
arXiv:2608.20574 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 FlavorBench:一个基于版本化烹饪嵌入模型编译稠密确定性答案映射的基准,并报告了一项 3 seed 的后训练研究——在 Epicure 的 270 条最优答案上对 Qwen3-0.6B checkpoint 进行 LoRA SFT 后,在该任务集上获得 13.3 分的提升。This work introduces FlavorBench: a benchmark for Compiling Dense Deterministic Answer Maps from a Versioned Culinary Embeddings Model and presents a 3-seed post-training study where LoRA SFT of a Qwen3-0.6B checkpoint on 270 optimal answers for Epicure to score on this task-set resulted in a 13.3 point gain.

Human-Centric Intelligence in the Era of Foundation Models: A Survey
基础模型时代下以人为本的智能:综述
arXiv:2608.18184 评测基准 综述 OA · 绿色 被引 0 · S2 + OpenAlex

提出一个全谱系的人类上下文分类法,将六个相互关联的层级整合在一起:将人类视为通过视觉外观与空间几何可观测的主体、视为通过运动学动态与交互建模的动态行动者,以及视为通过世界仿真与具身智能定位的处境化 Agent。A full-spectrum human context taxonomy is introduced that integrates six interconnected levels by viewing humans as observable subjects through visual appearance and spatial geometry, as dynamic actors through kinematic dynamics and interaction modeling, and as situated agents through world simulation and embodied agency.

MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks
MobilePA-Bench:在复杂真实任务上对移动端 Planner Agent 的基准评测
arXiv:2608.23035 评测基准 评测集 OA · 绿色 被引 1 · S2

通过将交互式函数调用沙箱与基于证据的验证相结合,MobilePA-Bench 既可作为实用的诊断基准,也是 Agent 强化学习的交互式基础,加速可靠移动 Agent 的开发。By pairing an interactive function-calling sandbox with evidence-based verification, MobilePA-Bench serves as both a practical diagnostic benchmark and an interactive foundation for agentic reinforcement learning---accelerating the development of dependable mobile agents.

Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection
Task-CoEvolve:通过自适应验证任务选择实现 Harness 高效优化
arXiv:2608.20169 评测基准 方法 OA · 绿色 被引 4 · S2

Task-CoEvolve 基于以下观察:相比被一致解决或一致失败的候选任务,候选 harness 之间存在分歧的任务更能提供区分信息;它利用基于历史结果的方差加权采样,将评估聚焦在能力前沿附近的任务上。Task-CoEvolve builds on the observation that tasks on which candidate harnesses disagree are more informative for distinguishing among them than tasks that are consistently solved or failed, and uses variance-weighted sampling based on past outcomes to focus evaluation on tasks near the capability frontier.

LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks
LongWoF-Bench:评估 EvoMap Gene 的可验证长工作流任务基准
arXiv:2608.23200 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

EvoMap 的结果表明,经过验证的执行经验可以被保留并共享为可复用的外部资源,使模型能够提升长工作流完成度,而无需反复承担经验探索的全部成本。EvoMap results show that verified execution experience can be retained and shared as a reusable external resource, enabling models to improve long-workflow completion without repeatedly paying the full cost of experience discovery.

ClawProBench: Trace-Aware Evaluation of AI Agents with Runtime Coverage and Frozen Workplace-Style Holdouts
ClawProBench:基于 trace 感知、运行时覆盖与冻结式工作场景 holdout 的 AI Agent 评测
arXiv:2608.22510 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 ClawProBench:基于 OpenClaw(具备 workspace 工具及浏览、记忆、消息、调度、skill、subagent 等原生能力的实时 agent 运行时)实例化的 trace-aware、runtime-native agent 评估基准。ClawProBench is presented, a trace-aware benchmark for runtime-native agent evaluation instantiated on OpenClaw, a live agent runtime with workspace tools and native surfaces for browsing, memory, messaging, scheduling, skills, and subagents.

Skill Issue: Are Skills Language-Invariant in LLMs?
Skill Issue:LLM 中的 Skill 是否具备语言无关性?
arXiv:2608.25832 评测基准 评测集 OA · 绿色 被引 1 · S2

本文通过多语言 self-play,正交于知识与综合基准性能对跨语言技能不一致性进行量化,表明技能差异是开发真正多语言模型过程中可衡量且主要的障碍。This work quantifies cross-lingual skill inconsistency orthogonally from knowledge and general benchmark performance via multilingual self-play, and shows that skill discrepancies are a measurable major roadblock in the development of truly multilingual models.

StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments
StarHarness:通过分层搜索为企业环境进化 Harness
arXiv:2608.24804 评测基准 方法 OA · 绿色 被引 1 · S2

StarHarness 提供了一种实用方法,通过根据基线失败行为对任务进行分层、将 proposer 可见的搜索任务与 proposer 隐藏的选择任务分离,并为评估泛化能力保留 held-out 任务,从而缓解工具密集型企业任务中持续的 model-environment mismatch。StarHarness offers a practical way to reduce persistent model-environment mismatch in tool-rich enterprise tasks by stratifying tasks according to baseline failure behavior, separating proposer-visible search tasks from proposer-hidden selection tasks, and reserves held-out tasks for evaluating generalization.

Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered
思维链忠实性随偏好线索的传递位置与方式而变化
arXiv:2608.29464 评测基准 方法 OA · 绿色 被引 1 · S2

结果表明,在所测试的单调用、预填工具场景下,当偏好信息通过工具传入或需从原始产物中推断时,CoT 监控的可靠性可能下降。The results suggest that CoT monitoring may be less reliable when preference information arrives through tools or must be inferred from raw artifacts, within the single-call, prefilled-tool setting tested here.