本文提出 SHAPER,一种免训练具身自适应的自演化框架,保持模型参数冻结,通过目标环境 rollout 演化可复用的 skills 和 context-code harness 来改进非参数化 Agent 系统。This work proposes SHAPER, a self-evolving framework for train-free embodied adaptation that keeps model parameters frozen and improves the non-parametric agent system by evolving reusable skills and a context-code harness through target-environment rollouts.
论文
1686 张论文卡片 · OA 绿色
提出 Simplax,一种精确的 Dirichlet–categorical 增强方法,将每个被损坏的 categorical 状态与一个辅助的单纯形值变量耦合,同时保持均匀扩散过程作为其 categorical 边际。Simplax is introduced, an exact Dirichlet--categorical augmentation that couples each corrupted categorical state with an auxiliary simplex-valued variable while preserving the uniform diffusion process as its categorical marginal.
本文提供了参数空间探索能够改进 LLM 强化学习的证据,并提出称为扰动参数策略优化(3PO)的方法族,使用不同的采样策略和不同的 rollout 分组进行 reward 估计。Evidence that parameter-space exploration can improve reinforcement learning for LLMs is presented, and a family of methods called Perturbed Parameter Policy Optimization (3PO) is introduced which use different sampling strategies and different rollout grouping for reward estimation.
本文提出 SkillZip,一种执行感知的程序化抽象框架,对 section 级图执行保持契约的压缩,加载一个紧凑、依赖闭合的 context,并仅在需要时展开宏。SkillZip is proposed, an execution-aware procedural abstraction framework that performs contract-preserving compression over section-level graphs that hydrates a compact, dependency-closed context and expands macros only when required.
一个仅依赖字符级输入的简单神经语言模型,仅从字符即可编码语义和正字法信息,表明在许多语言中,字符输入足以完成语言建模。A simple neural language model that relies only on character-level inputs that is able to encode, from characters only, both semantic and orthographic information and suggests that on many languages, character inputs are sufficient for language modeling.
描述了一种无监督学习通用分布式句子编码器的方法,利用书籍文本的连续性,训练编码器-解码器模型以重建编码段落的周围句子。The approach for unsupervised learning of a generic, distributed sentence encoder is described, using the continuity of text from books to train an encoder-decoder model that tries to reconstruct the surrounding sentences of an encoded passage.
本工作训练了一个预测的计算最优模型 Chinchilla,使用与 Gopher 相同的计算预算,但参数量为 70B、数据量为 4 倍,达到 SOTA 平均准确率,比 Gopher 提升超过 7%。This work trains a predicted compute-optimal model, Chinchilla, that uses the same compute budget as Gopher but with 70B parameters and 4$\times$ more more data, and reaches a state-of-the-art average accuracy, greater than a 7% improvement over Gopher.
本工作综述了可用于推断用户需求的贝叶斯用户模型研究,这些模型综合考虑用户的背景、操作和查询,并提出了一种智能用户界面的整体架构。This work reviews work on Bayesian user models that can be employed to infer a user's needs by considering a users' background, actions, and queries and proposes an overall architecture for an intelligent user interface.
探索以交错方式使用 LLM 同时生成推理轨迹和任务特定动作,使两者产生更大协同:推理轨迹帮助模型归纳、跟踪和更新动作计划以及处理异常,而动作使其与外部源交互以获取额外信息。The use of LLMs are explored to generate both reasoning traces and task-specific actions in an interleaved manner, allowing for greater synergy between the two: reasoning traces help the model induce, track, and update action plans as well as handle exceptions, while actions allow it to interface with external sources to gather additional information.
本文提出提示式注视目标估计(PGE)任务,一种用于注视分析的端到端、概念驱动新范式,并推出首个为 PGE 设计的模型 GazeAnywhere,它使用基于 Transformer 的检测器融合冻结编码器特征,同时解决主体定位、画内/画外存在性以及注视目标热图估计。The Promptable Gaze Target Estimation (PGE) task is introduced, a new end-to-end, concept-driven paradigm for gaze analysis and GazeAnywhere, the first model designed for PGE, uses a transformer-based detector to fuse features from frozen encoders and simultaneously solves subject localization, in/out-of-frame presence, and gaze target heatmap estimation.
ParliamentRAG 是一种主题相关的权威模型,根据当前 query 估计每位发言者的权威性,结合职业、教育和此前发言等可解释组件,以应对政治敏感文本中最高频发言者主导、无法按主题专长加权发言者以及引用归属错误的风险。ParliamentRAG is a topic-dependent authority model that estimates each speaker's authority as a function of the current query, combining interpretable components such as profession, education, and previous interventions that addresses risks of dominance of the most frequent speakers, inability to weight speakers according to topical expertise, and citation misattribution in politically sensitive text.
本文提出一种用于机器人操作的动作条件视频世界模型,给定观测帧、语言指令以及由末端执行器位姿和夹爪状态组成的预定动作序列,预测对应的未来观测结果。An action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action sequence comprising end-effector poses and gripper states, predicts the resulting future observations is presented.
本文提出 UniSwap,这是首个用于说话视频中流式联合音视频身份替换的框架,并引入 swap-and-reconstruct 流程,从真实片段中移除视觉和声音身份,同时使用原始片段作为重建目标。This work presents UniSwap, the first framework for streaming joint audio-visual identity replacement in talking videos, and introduces a swap-and-reconstruct pipeline that removes visual and vocal identity from real clips and uses the original clips as reconstruction targets.
本文提出 LiveAnimate,据作者所知是首个将实时流式生成与十亿参数规模下的稳定长视频生成相结合的系统,基于 140 亿参数的视频 Diffusion Transformer(DiT)。This work presents LiveAnimate, to their knowledge the first animation system to combine real-time streaming with stable long-form generation at billion scale, built on a 14B-parameter video Diffusion Transformer (DiT).
H2R-Bench 提供了一个系统性诊断框架,用于评估视频世界模型能否跨越 human-to-robot 具身差距,并将人类操作观测转化为以机器人为中心的训练资源。H2R-Bench provides a systematic diagnostic framework for evaluating whether video world models can bridge the human-to-robot embodiment gap and convert human manipulation observations into robot-centric training resources.
SKILLER 是一个由自然语言驱动的强化学习框架,旨在为小模型自动生成执行器特定的 skills,使用强模型作为 actor 和 critic,将小模型 Agent 系统视为环境,并通过自然语言完全传递所有强化学习信号。SKILLER is a natural-language-driven reinforcement learning framework designed to automatically generate executor-specific skills for small models, which employs a strong model as the actor and critic, treats the small-model agent system as the environment, and propagates all reinforcement learning signals entirely via natural language.
本文给出了 LLM routing 的统一形式化,将其刻画为由五个组件构成的序贯决策过程:context 编码器、模型编码器、评分函数、决策规则和学习信号,涵盖单轮、多轮和个性化 routing。This work presents a unified formulation of LLM routing as a sequential decision process characterized by five components: context encoders, model encoders, scoring functions, decision rules, and learning signals, covering single-turn, multi-turn, and personalized routing.
本文认为关键在于设计合适的 search-state 表示,将 LLM 的内部知识与 visual-token 剪枝的结构要求和约束相连接,并提出 AutoPrune,一种用于 LLM 驱动的 visual-token 剪枝策略设计的免训练框架。This paper argues that the key lies in designing an appropriate search-state representation that connects the internal knowledge of LLMs with the structural requirements and constraints of visual-token pruning, and proposes AutoPrune, a training-free framework for LLM-driven visual-token pruning policy design.
本文将 S2G-RAG 的结构化充分性-缺口判断适配到冻结的 Search-R1 流程中,并在来自 900 个不相交 HotpotQA 问题的 3009 个状态上训练了一个 Qwen3.5-2B judge,以减少检索次数同时广泛保持答案准确性。This work adapts S2G-RAG's structured sufficiency-and-gap judgment to a frozen Search-R1 pipeline and trains a Qwen3.5-2B judge on 3,009 states from 900 disjoint HotpotQA questions to reduce retrieval while broadly preserving answer accuracy.
本文提出 Syfer,一种用于多语言多跳问答的 synthesizer-folding 框架,默认推迟翻译而非直接应用翻译,在保持具有竞争力准确性的同时,在性能与计算成本之间取得良好平衡。The method Syfer is introduced, a synthesizer-folding framework for multilingual multi-hop question answering that defers translation rather than applying it by default and attains competitive accuracy while striking a favourable balance between performance and computational cost.
本文提出 RAGSieve,为每个检测范围构建匹配的参考集合,其构建需要投毒标签、可信语料或训练过程。This work presents RAGSieve, which constructs a reference matched to each detection scope, which requires poison labels, a trusted corpus, or training to be constructed.
本文提出 Reduced Matrix Multiplication,一种免训练的输入自适应推理方法,通过沿收缩维度选取信息性切片来减少 Transformer 矩阵乘积,且不修改模型权重,并表明同一原理可扩展到多模态视觉-语言推理。Reduced Matrix Multiplication is proposed, a training-free, input-adaptive inference method that reduces Transformer matrix products by selecting informative slices along their contraction dimensions, without modifying model weights, and it is shown that the same principle extends to multimodal vision-language inference.
提出 Agentic Video Auto-Encoder(AVA-Encoder),一种由 agentic 自我进化驱动的新型自编码框架,用于学习 agent-native 视频表示,在 shot-level 和 keyframe-level system-prompt token 使用量减少 74.3% 的同时,性能优于精心人工调优的策略。The Agentic Video Auto-Encoder (AVA-Encoder), a novel auto-encoding framework driven by agentic self-evolution to learn agent-native video representations that outperforms a carefully human-tuned policy while using 74.3% fewer shot-level and keyframe-level system-prompt tokens.
论文表明,使用 Stanford Natural Language Inference 数据集有监督训练的通用句子表示,在广泛的迁移任务上能持续优于 SkipThought vectors 等无监督方法。It is shown how universal sentence representations trained using the supervised data of the Stanford Natural Language Inference datasets can consistently outperform unsupervised methods like SkipThought vectors on a wide range of transfer tasks.
提出 Context-Matched Distillation(CMD),一种因果 DMD 框架,将 teacher 监督信号与每个 target 生成时可用信息对齐,并可自然扩展到逐帧和逐 chunk 生成、长视频蒸馏以及相机条件蒸馏。Context-Matched Distillation (CMD) is introduced, a causal DMD framework that aligns teacher supervision with the information available when each target is generated, and naturally extends to frame-wise and chunk-wise generation, long video distillation, and camera-conditioned distillation.
提出 HPSE,构建一种混合 rollout,能够在 student 自身轨迹中 coverage 缺失的精确位置补齐缺失事实,同时在其他位置保持 on-policy。HPSE is proposed, which builds a hybrid rollout that steps in to place missing facts onto the student's own trajectory precisely where its coverage fails, while staying on-policy elsewhere.
本文报告了一项完整的、有完整记录的案例研究:在规范优先协议下,由 AI 编码 Agent 对大规模架构进行重构,期间无人工代码审查、无预先存在的预言机来验证目标行为。该任务是在一个大型相互依赖的代码库中拆除核心不变量,作者评估认为通过增量重构基本上不可行,这类变更通常需要重写。本文所述协议下,Agent 成功完成了任务。该系统包含 717,725 行This paper reports a single, fully instrumented case study of a large-scale architectural refactoring by an AI coding agent under a specification-first protocol, with no human review of the generated code and no pre-existing oracle to validate the target behaviour. The task, dismantling a central invariant across a large interdependent codebase, was assessed by the author as effectively infeasible through incremental refactoring, the kind of change that conventionally calls for a rewrite instead. Under the protocol described here, the agent completed it successfully. The system is a 717,725-li
Gambit 通过周期性剪除低潜力轨迹并立即从高质量前缀分支,利用轻量 scorer 探测 hidden state,将算力动态集中到最有潜力的推理路径上,同时保持持续的高硬件利用率。By periodically pruning unpromising trajectories and immediately branching from high-quality prefixes, Gambit dynamically concentrates compute onto the most promising reasoning traces via a light-weight scorer probing hidden states while maintaining continuous high hardware utilization.
一种具有固定大小记忆的循环 Transformer 架构,可泛化 sliding-window attention,同时在训练期间保持并行性,并在验证 loss 和下游预训练 benchmark 上优于 sliding-window 和 latent recurrent Transformer 基线。A recurrent Transformer architecture with fixed-size memory that generalizes sliding-window attention while remaining parallelizable during training and improves validation loss and downstream pretraining benchmarks over sliding-window and latent recurrent transformer baselines.
本文提出一种权重裁剪的替代方案:对 critic 相对于其输入的梯度范数施加惩罚。其性能优于标准 WGAN,能以几乎无需调参的方式稳定训练多种 GAN 架构。This work proposes an alternative to clipping weights: penalize the norm of gradient of the critic with respect to its input, which performs better than standard WGAN and enables stable training of a wide variety of GAN architectures with almost no hyperparameter tuning.
当将监督规模扩展到 680,000 小时的多语言、多任务数据时,所得到的模型在标准 benchmark 上泛化良好,在 zero-shot transfer 设置下常可与此前全监督方法的结果相当,且无需任何微调。When scaled to 680,000 hours of multilingual and multitask supervision, the resulting models generalize well to standard benchmarks and are often competitive with prior fully supervised results but in a zero-shot transfer setting without the need for any fine-tuning.
建立一个可复现的选择性 3D 定位框架,并指出跨视角对应(cross-view correspondence)是主要操作瓶颈。The study establishes a reproducible framework for selective 3D localization and identifies cross-view correspondence as the dominant operational bottleneck.
首个同时使用 LLM 推理和 tag-aware 翻译来明确处理并评估英罗 MT 中性别偏见的方法。This is the first method to explicitly address and evaluate gender bias in English-Romanian MT using both LLM inference and tag-aware translation.
对八个法律 RAG 系统在 GDPR 和某国国内民法两个法律语料上的幻觉行为进行细粒度分析,发现含有错误假设的 false-premise 问题在人工编写问题上产生高幻觉率。A fine-grained analysis of hallucination behavior in eight legal RAG systems across two legal corpora, the GDPR and a national civil law, finds that false-premise questions, containing incorrect assumptions that must be rejected, produce high hallucination rates on the manually-drafted questions.
基于 CPI-Bench 对主流图像编辑模型的评测结果显示,CPI-Bench 增强了模型间的性能区分度;排名分析表明 CPI-Bench 与 Arena Image Edit Leaderboard 的对齐度最高,与公开人类偏好排名具有更强的一致性。Evaluation results of mainstream image editing models based on CPI-Bench demonstrate that CPI-Bench enhances performance differentiation among models, and ranking analysis reveals that CPI-Bench achieves the highest alignment with the Arena Image Edit Leaderboard, indicating stronger consistency with public human preference rankings.
在对 1,000 只美股进行小时级收益预测时,我们观察到一种意外现象:预测结果近乎平坦,且截面相关性衡量的股票排序能力很差,我们将其称为"预测坍缩"。令人惊讶的是,在同一设置下对交易量进行预测时,该现象基本消失。我们在多种时序基础模型(TSFMs)、12 个深度学习预测模型以及 97 个公开基准配置中系统考察了这一现象,发现其与目标可预测性密切相关,并识别出背后的两类成因:低 p……When forecasting hourly returns for 1,000 US equities, we observe an unexpected phenomenon: predictions become nearly flat and show poor stock ranking, as measured by cross-sectional correlation. We call this forecast collapse. Surprisingly, the phenomenon largely disappears when forecasting trading volume under the same setting. We investigate forecast collapse across time-series foundation models (TSFMs), twelve deep-learning forecasting models, and 97 public benchmark configurations, and find that it is closely tied to target predictability. We identify two distinct reasons behind it: low p