理论分析表明 RWTD 的不动点分布在参考策略的 off-policy 奖励倾斜与当前模型的 on-policy 倾斜之间插值,提供了一种在奖励适配与保留先验知识之间取得平衡的原则性方法。Theoretical analysis shows that the fixed-point distributions of RWTD interpolate between off-policy reward tilting of the reference and on-policy tilting of the current model, providing a principled approach to balancing reward adaptation with retention of prior knowledge.
论文
63 张论文卡片 · 评测基准 · 方法
实验表明,基于评分细则的奖励能可靠地区分不同质量的法律回答,且在该偏好数据上进行 DPO 训练在所有三个维度上均提升了性能;逐维度分析进一步支持了所提分类法与奖励构建的有效性。Experiments show that the rubric-based rewards reliably distinguish legal responses of different quality and that DPO training on the preference data improves performance across all three dimensions, and Dimension-wise analyses further support the effectiveness of the proposed taxonomy and reward construction.
本文提出 Multilingual GSM-Symbolic,一个可扩展的多语言数学数据集,包含 30,000 对题项匹配的问答对,覆盖 15 种语言;并将能力差异的最大决定因素量化为模型规模、语言资源水平、推理能力与类型学距离。This work introduces Multilingual GSM-Symbolic, an extensible multilingual mathematical dataset covering 30,000 item-matched question-answer pairs and spanning 15 languages and quantifies the largest determinants of capability as model size, language resource level, reasoning, and typological distance.
为每个可查询的条件分布推导出一个尾部影响度量,聚合其不确定性对每一次 Bellman 复用中 CVaR 的影响,并刻画了尾部最优与均值最优分配重合的精确网格(exact-grid)机制。A tail influence for each queryable conditional law that aggregates how its uncertainty affects CVaR across every Bellman reuse is derived, and an exact-grid regime in which tail- and mean-optimal allocations coincide is characterized.
定性分析表明,EMHO 超越了从失败与无效动作中恢复的范畴,重塑了具身 agent 对环境的解释与交互方式,并在 EmbodiedBench 的导航和操作任务上进行了评估。Qualitative analysis shows that EMHO goes beyond recovering from failures and unproductive actions to reshape how the embodied agent interprets and interacts with its environment, and is evaluated on EmbodiedBench across navigation and manipulation tasks.
结果表明,强大的 LLM 能够产出一系列合理思路,但其范围仍比人类研究品味更窄,并存在系统性偏移。It is suggested that strong LLMs can produce a range of reasonable ideas, but that range remains narrower than, and systematically shifted relative to, human research taste.
介绍其潜在原因的理论基础,阐明强化学习算法的渐近性能在性能排名与数据规模之间不存在单调关系。The theoretical foundations of the underlying causes outlining that the asymptotic performance of reinforcement learning algorithms does not have a monotone relationship between performance rankings and data-regimes are introduced.
本工作通过对推理模型施加长度惩罚以抑制过度思考并降低推理成本,但表明这些惩罚会削弱思维链的可监控性,通过移除监控所依赖的证据,以可监控性换取推理成本。This work trains reasoning models with length penalties to curb overthinking and cut inference cost but shows that these penalties make the chain of thought less monitorable, and trades monitorability for inference cost by removing the evidence monitors depend on.
本文提出 UniVR,这是首个从纯视觉演示中同时学习复杂推理、细粒度物理动力学和长期规划的研究,并配套首个在纯视觉协议下评估这些异质能力的综合评测套件。UniVR, the first investigation into simultaneously learning complex reasoning, fine-grained physical dynamics, and long-term planning from pure visual demonstrations, is introduced, the first comprehensive suite to assess these heterogeneous capabilities under a purely visual protocol.
发现对 GPT 语言模型进行重复采样,是为困难 prompt 生成可用解的一种出乎意料的有效策略,并讨论了部署强大代码生成技术在安全性、安全保障与经济等方面的潜在更广泛影响。It is found that repeated sampling from the GPT language model is a surprisingly effective strategy for producing working solutions to difficult prompts, and the potential broader impacts of deploying powerful code generation technologies, covering safety, security, and economics are discussed.
在该研究的评估协议下,部分 AI 生成回答获得了高于可用参考答案的评分,尤其当其包含准确且相关的支撑细节时;但这不应被解读为模型总体上优于合格法律专业人士的证据。Under the study’s evaluation protocol, some AI-generated responses received higher ratings than the available reference answers, particularly when they contained accurate and relevant supporting detail, should not be interpreted as evidence that the models generally outperform qualified legal professionals.
本文提出 Verbatimize 这一新任务,能够以高质量规范化逐字转录对语音语料库进行可扩展的创建与扩充,并通过有监督的 cross-attention 微调,将不流畅语音上的词级时间戳表现提升至超过强制对齐 baseline。Verbatimize is proposed, a new task enabling scalable creation and enrichment of speech corpora with high-quality canonical verbatim transcriptions and supervised cross-attention fine-tuning that improves word-level timestamps on disfluent speech beyond forced-alignment baselines.
FinMMEval 2026 Task 2 围绕多语言证据评估金融领域的短答问答,采用文档 RAG、跨语言证据处理、结构化提示、答案压缩与验证策略。FinMMEval 2026 Task 2 evaluates short-answer financial question answering over multilingual evidence over multilingual evidence using document retrieval-augmented generation, cross-lingual evidence handling, structured prompting, answer compression, and validation strategies.
提出 OpenForgeRL,一个用于在多样化环境中端到端训练基于 harness 的 Agent 的开源框架,并在多种复杂的 harness 和环境中得到验证,涵盖工具/爪型 Agent 以及多模态 GUI 浏览器和计算机使用 Agent。OpenForgeRL is presented, an open-source framework for training harness-based agents end-to-end in diverse environments, and validated across diverse, complex harnesses and environments, spanning tool/claw-based agents and multimodal GUI browser- and computer-use agents.
本文提出 PAJAMA,该系统将程序合成为 judge,将其决策聚合为联合裁决(joint verdict),并通过 fallback 机制选择性地将低置信度用例升级交由 LLM 处理。PAJAMA is introduced, a system that synthesizes programs as judges, aggregates their decisions into a joint verdict, and incorporates a fallback mechanism to selectively escalate low-confidence cases to an LLM.
介绍 DecoEvo(Decoupled Co-Evolution),在解耦目标下协同进化一个求解器 skill 和一个 rubric 生成器 skill,优化过程中不使用 gold rubric。DecoEvo (Decoupled Co-Evolution), which co-evolves a solver skill and a rubric-generator skill under decoupled objectives without using gold rubrics during optimization, is introduced.
本文提出 CoRT,一种用于 rubric 条件化 GRPO 的 token 级 credit 加权方法,通过反事实回放在原始 rubric 条件化 prompt 和匹配的免准则 prompt 下对同一样本响应重新打分,并表明策略内部反事实似然对比为响应内 credit 分配提供了有效的训练信号。CoRT is proposed, a token-level credit weighting method for rubric-conditioned GRPO that uses counterfactual replay to rescore the same sampled response under the original rubric-conditioned prompt and a matched criteria-free prompt, and suggests that policy-internal counterfactual likelihood contrasts provide an effective training signal for within-response credit allocation.
对 Fairness Pruning 的实证评估表明,群体偏置处理与模型能力运行在可分离的电路上,奠定了从盲目零化向定向行为调制过渡的方法论基础。Empirical evaluation of Fairness Pruning empirically confirm that demographic bias processing and model capabilities operate on dissociable circuits, establishing the methodological foundations for transitioning from blind zeroing toward directional behavior modulation.
多维评估-验证奖励(EVR)将评估分解为独立视觉准则;针对每个准则,MLLM Evaluator生成多个候选假设,Verifier在具体视觉证据中grounding每个claim以接受或拒绝,产生可靠且细粒度的奖励信号。A Multi-dimensional Evaluation-Verification Reward (EVR) decomposes evaluation into distinct visual criteria; for each criterion, an MLLM Evaluator generates multiple candidate hypotheses, and a Verifier grounds each claim in concrete visual evidence to accept or reject it, producing reliable and fine-grained reward signals.
RecHarness 将优化过程分为两步:bandit 路由器根据历史验证反馈选择下一步修改方向,LLM 在选定方向内生成具体优化假设与可执行代码编辑RecHarness separates the optimization process into two steps: a bandit router selects the next modification direction according to historical validation feedback, while the LLM generates a concrete optimization hypothesis and executable code edit within the selected direction.
在后训练中教授删除操作可减少删除回避行为并提升更广泛的代码编辑性能,表明该行为是训练不足而非不可达成Teaching deletion during post-training reduces deletion avoidance and improves broader code-editing performance, suggesting the behavior is undertrained rather than beyond reach.
提出 OSReward,一个面向真实场景的高质量基准,用于评估 VLM 评判器在 CUA 轨迹上的表现;同时发布一个面向 CUA 社区的、带推理标注的开放轨迹判断语料库,以弥补大规模可靠 CUA 奖励的缺口。OSReward is introduced, a realistic, high-quality benchmark that evaluates VLM judges on CUA trajectories, and an open corpus of reasoning-annotated trajectory judgments for the CUA community, to close the gap in reliable CUA reward at scale.
本工作提出 BERTScore——一种文本生成自动评估指标,与人类判断的相关性更强,并在模型选择性能上优于现有指标。This work proposes BERTScore, an automatic evaluation metric for text generation that correlates better with human judgments and provides stronger model selection performance than existing metrics.
提出 C4,一个受认知启发的成语跨概念创造力评估框架,揭示了当前 MLLM 在通过跨概念关系解码创造性编码语义方面存在的显著差距。C4, a cognition-inspired evaluation framework for Chengyu (Chinese idiom)-based Cross-Concept Creativity, is introduced, exposing a substantial gap in how current MLLMs decode creatively encoded meaning through cross-concept relations.
本文提出 Self-Modulating QKAN-based FWPs,用低秩生成的逐元素调制(作用于新提案分支、有界旧状态分支或两者)替代该广播门,并提出 Complementary Matrix Gating (CMG),在保持标量门控的有界凸更新和仿射前缀扫描结构的同时提供逐坐标的记忆控制。This work introduces Self-Modulating QKAN-based FWPs, which replace this broadcast gate with low-rank-generated element-wise modulation of the new-proposal branch, a bounded old-state branch, or both, and proposes Complementary Matrix Gating (CMG), which provides coordinate-wise memory control while preserving the bounded convex update and affine prefix-scan structure of scalar gating.
本文提出 BDH-CQ,一个结合上下文学习与循环潜在推理的推理模型,通过在高维潜在空间中的迭代计算来求解查询,且无需将中间推理外化为语言。BDH-CQ is introduced, a reasoning model that combines in-context learning with recurrent latent reasoning that solves a query through iterative computation in a high-dimensional latent space, without verbalizing its intermediate reasoning.
提出 Large Discovery Model (LDM),一种经验驱动的循环架构,将生成模型与贝叶斯非参数奖励代理模型耦合,产生一种感知不确定性的价值,用于引导候选的生成、精炼与选择。This work introduces the Large Discovery Model (LDM), an empirically grounded recurrent architecture that couples a generative model with a Bayesian non-parametric reward surrogate model, yielding an uncertainty-aware value that guides candidate generation, refinement, and selection.