理论分析表明 RWTD 的不动点分布在参考策略的 off-policy 奖励倾斜与当前模型的 on-policy 倾斜之间插值,提供了一种在奖励适配与保留先验知识之间取得平衡的原则性方法。Theoretical analysis shows that the fixed-point distributions of RWTD interpolate between off-policy reward tilting of the reference and on-policy tilting of the current model, providing a principled approach to balancing reward adaptation with retention of prior knowledge.
论文
1094 张论文卡片 · 方法 · OA 绿色
提出一种自适应、自动化的数据提取攻击流程,在黑盒设置下针对 MRAG(其中检索到的视觉产物本身就是答案)发起攻击,表明亟需专门面向多模态数据设计的安全防护。An adaptive and automatic data extraction attack procedure operating in a black box setting against MRAG, a configuration in which the retrieved visual artifact is itself the response, and shows the urgent need for safeguards specifically designed for multimodal data.
InterEvolve 提出了一个物体感知的前向-后向行为基础模型,其在冻结身体先验上的物体残差可在测试时将关于身体或物体的奖励转化为 loco-manipulation 行为,并将任务以奖励程序的形式指定:带完成条件与可调常数的分阶段奖励。InterEvolve develops an object-aware forward-backward behavioral foundation model, whose object residuals on a frozen body prior turn a new reward about the body or objects into loco-manipulation behavior at test time, and specifies tasks as reward programs: staged rewards with completion conditions and tunable constants.
结果表明,有效的拒绝能够保留任务结构,同时限制与参考策略的耦合,并且较小的冻结模型可以低成本地提供此类拒绝能力。The results suggest that effective rejects preserve task structure while limiting coupling to the reference policy, and that smaller frozen models can provide them at low cost.
结果表明,纯视觉设置会降低准确率并增加 token 成本,因为 Agent 缺乏足够的符号化细节,需通过重复的视觉查询进行补偿;研究指向一种面向下一代编码 Agent 的实用文本与视觉混合设计。The results show that a strictly vision-only setup degrades accuracy and increases token cost, because agents lack sufficient symbolic detail and compensate with repeated visual queries, and point to a practical hybrid text-and-vision design for next-generation coding agents.
结果表明,HeteroFold 能够在接收端无需预填充的前提下实现高效的跨系列 KV 复用,并在全部四个长上下文基准和大多数短上下文设置上取得最佳的 cache 迁移性能。Results show that HeteroFold enables efficient cross-family KV reuse without receiver prefill, and achieves the best cache-transfer performance on all four long-context benchmarks and most short-context settings.
DataMagic 通过声明式多 Agent 编排,从原始表格数据自动生成数据可视化视频,在完全自动化与细粒度人工控制之间架起桥梁,并提升了创作效率、降低了感知认知负荷。DataMagic, which authors data videos from raw tabular data through declarative multi-agent orchestration, bridging full automation with fine-grained human control, and improves creation efficiency and reduces perceived cognitive load.
本文提出 LOCI,一种混合的空间记忆架构,同时维护键值缓存与循环记忆两种表示,在重访内容的复现上比代表性世界模型以及同配置的 full-softmax 模型都更为忠实。LOCI is introduced, a hybrid spatial-memory architecture that keeps both representations of key-value caches and recurrent memory that reproduces revisited content more faithfully than representative world models and a same-recipe full-softmax model.
本文提出 AntPlan,一个包含 505 张真实专业建筑平面图、覆盖 92 类物体和十类住宅房间且具有密集家具标注的精选数据集,以及 Architect-Ant,一个用于生成家具布局的框架,可在不依赖高成本迭代式 Agent 推理的情况下直接进行约束感知的布局生成。AntPlan, a curated dataset of 505 real professional architectural floor plans with dense furniture annotations spanning 92 object classes and ten residential room categories, and Architect-Ant, a framework for generating furniture layouts are introduced, enabling direct constraint-aware layout generation without relying on costly iterative agentic inference.
SAKIKO 是一个审计框架,通过定向错误发现、路由器条件干预、目标解析验证以及前瞻性冻结统计许可来形式化表示修复,并确立了在声明内部修复之前必须进行结果解析裁定的必要性。SAKIKO is presented, an auditing framework that formalizes representation repair via directional error discovery, router-conditioned intervention, destination-resolved verification, and prospectively frozen statistical licensing and establishes the necessity of outcome-resolved adjudication before claiming internal repair.
提出一种新颖的深度架构与 GAN 形式化方法,有效衔接文本与图像建模领域的进展,将视觉概念从字符转换为像素。A novel deep architecture and GAN formulation is developed to effectively bridge advances in text and image modeling, translating visual concepts from characters to pixels.
提出 MMAgent-R$^2$,一种将视觉重排序与主动拒答作为内部验证机制的 Agentic mRAG 框架,并通过 GRPO 训练实现外部检索、内部验证与答案生成的联合优化。MMAgent-R$^2$, an agentic mRAG framework that integrates visual reranking and active rejection as its internal verification mechanism, is proposed and achieves joint optimization of external retrieval, internal verification, and answer generation via GRPO training.
MatRAG 在检索质量上优于其最强的竞争者;同时,通过避免 KG 构建和基于 LLM 的摘要降低了索引成本,并通过维度感知的相似度降低了查询成本。MatRAG outperforms its strongest competitors in terms of retrieval quality; furthermore, it reduces indexing costs by avoiding KG construction and LLM-based summarization, and lowers query-time costs through dimension-aware similarity.
提出了 Honeycomb,一种基于 HexMemory 构建的视频世界模型;HexMemory 是作者提出的低秩表示,用于在固定大小的内存中存储场景特征,共包含六个空间与时空平面。Honeycomb is introduced, a video world model built on HexMemory, the authors' proposed low-rank representation for storing scene features in a fixed-size memory with a total of six spatial and spatiotemporal planes.
在数据和预算匹配的条件下,X-Tree 相较标准方案在 WebArena 上 SR 提升最高 4.5%,在 ScienceWorld 上提升 5.8%,在 WebShop 上成功率提升 4.1%;匹配分析表明增益源自 X-Tree 结构及其三项集成设计。X-Tree improves over standard recipes at matched data and budget by up to 4.5% SR on WebArena, 5.8% SR on ScienceWorld and 4.1% success on WebShop, and matched analyses attribute the gains to the X-Tree structure and the three integrations.
本文提出 MemFold,根据所支持的行为来优化固定预算的软记忆;在 PersonaMem-32K 和 PersonaMem-128K 上取得了作者所测得的最高准确率,且在更长历史长度下优势进一步扩大。This work presents MemFold, which optimizes a fixed-budget soft memory by the behavior it supports, and attains the highest accuracy the authors measure on PersonaMem-32K and PersonaMem-128K, with margins that widen at the longer history length.
在统计物理中,多元硬核模型描述一个粒子系统,每个粒子拥有各自的逸度。用图论语言表述,该模型的配分函数对应多元独立多项式,即独立多项式的多重仿射推广,定义为 $Z_G(λ_1,\dots,λ_n) := \sum_{I\in\mathcal{I}(G)} \prod_{v\in I}λ_v$,其中 $\mathcal{I}(G)$ 表示 $[n]:=\{1,2,\dots,n\}$ 上图 $G$ 的所有独立集。我们证明对于 $[n]$ 上的每个简单图 $G$ 以及 $λ_1,\dots,λ_n\geq 0$,\[ Z_G(λ_1,\dots,In statistical physics, the multivariate hard-core model describes a system of particles, each of which receives its own fugacity. In graph-theoretic language, the partition function of the model translates to the multivariate independence polynomial, i.e., the multiaffine generalisation of the independence polynomial, defined by $Z_G(λ_1,\dots,λ_n) := \sum_{I\in\mathcal{I}(G)} \prod_{v\in I}λ_v$, where $\mathcal{I}(G)$ denotes the set of all independent sets in a graph $G$ on $[n]:=\{1,2,\dots,n\}$. We prove that for every simple graph $G$ on $[n]$ and $λ_1,\dots,λ_n\geq 0$, \[ Z_G(λ_1,\dots,
提出一种基于不确定性感知框架的自适应问答方法,通过 LLM 内部表征中区分知识不足与知识歧义/冲突的显式信号,在单次前向传播中即可由隐状态高效估计。This work proposes an uncertainty-aware framework for adaptive QA based on explicit signals derived from LLM internal representations that distinguish between knowledge insufficiency and knowledge ambiguity or conflict, and efficiently estimate these from hidden states in a single forward pass.
预训练 Transformer 仅利用其深度的一小部分来跟踪上下文中的引用。十三个基础模型仅能可靠地跟随 1.4–3.6 行,额外的预训练循环收益甚微。在一个早期层上训练的 rank-8 LoRA 在所有模型权重冻结的情况下扩展了这一计算能力。Qwen3-8B 在 24 行链上的精确准确率从 15.5% 提升至 99%;更长训练的 LoRA 可达 50 行。Ouro-1.4B 经过四轮循环达到 60 行,八轮后至少达到 160 行。该 LoRA 启动了一场接力:程序行通过中间层的一段短距离传递其链身份。冻结的 head 逐层读取渐进式进展信号……Pretrained transformers use little of their depth to follow references in context. Thirteen base models reliably follow only 1.4-3.6 lines, and extra pretrained loops add little. A task-trained rank-8 LoRA at one early layer extends this computation with all model weights frozen. Qwen3-8B improves from 15.5% to 99% exact accuracy on 24-line chains; a longer-trained LoRA reaches 50 lines. Ouro-1.4B reaches 60 lines after four loops and at least 160 after eight. The LoRA starts a relay: program lines pass on their chain identity through a short range of middle layers. Frozen heads read progressi
结果将控制器所学到的行为范围与该范围内请求的准确性分离开来,尽管连贯性下限并非对每种特性都成立。The results separate the behavioral range learned by a controller from the accuracy of requests within that range, although the coherence floor does not hold for every trait there.
在 Encyclopedic-VQA 和 InfoSeek 上的实验表明,CLIMB 持续优于基于检索增强的多模态基线;消融实验显示互补池化、基于评论家的打分以及迭代置信度控制的精炼各自对最终性能均有贡献。Experiments on Encyclopedic-VQA and InfoSeek show that CLIMB consistently improves over retrieval-augmented multimodal baselines, andlations indicate that complementary pooling, critic-based scoring, and iterative confidence-controlled refinement each contribute to the final performance.
ProAR 引入两个关键组件:一是将生成锚定到长远结果,二是通过非对称注意力掩码将目标帧预测集成到自回归循环中,使预测的目标帧可指导中间状态的生成而不被其干扰。ProAR introduces two key components: to anchor generation to the long-range outcome, and to integrate goal-frame prediction into the autoregressive loop via an asymmetric attention mask, enabling the predicted goal frame to guide the generation of intermediate states without being disrupted by them.
提出 LingBot-Video,一种专为具身智能设计的基于 DiT 的视频预训练范式,并将其作为首个大规模开源 MoE 视频基础模型贡献给社区,致力于在数字创意与物理执行之间搭建桥梁。LingBot-Video is presented, a DiT-based video pretraining paradigm specifically tailored for embodied intelligence, and is contributed as the inaugural large-scale, open-source MoE video foundation model to the community, in a pioneering effort to bridge digital creativity and physical actuation.
实验表明,基于评分细则的奖励能可靠地区分不同质量的法律回答,且在该偏好数据上进行 DPO 训练在所有三个维度上均提升了性能;逐维度分析进一步支持了所提分类法与奖励构建的有效性。Experiments show that the rubric-based rewards reliably distinguish legal responses of different quality and that DPO training on the preference data improves performance across all three dimensions, and Dimension-wise analyses further support the effectiveness of the proposed taxonomy and reward construction.
提出了 MotorMind,一个机器人操作框架,将 VLM 提出的中级动作与确定性机器人控制及反馈相连接,支持异步监控与后台记忆更新;研究表明,通用 VLM 在配备合适的中级动作表示和异步执行框架后,可执行有效的零样本机器人操作。MotorMind is introduced, a robot manipulation harness that connects VLM-proposed mid-level actions to deterministic robot control and feedback, with asynchronous monitoring and background memory updates, and shows that a general-purpose VLM, when equipped with an appropriate mid-level action representation and asynchronous execution harness, can perform effective zero-shot robotic manipulation.
本文提出 Multilingual GSM-Symbolic,一个可扩展的多语言数学数据集,包含 30,000 对题项匹配的问答对,覆盖 15 种语言;并将能力差异的最大决定因素量化为模型规模、语言资源水平、推理能力与类型学距离。This work introduces Multilingual GSM-Symbolic, an extensible multilingual mathematical dataset covering 30,000 item-matched question-answer pairs and spanning 15 languages and quantifies the largest determinants of capability as model size, language resource level, reasoning, and typological distance.
本文用三种以不同方式施加长度压力的方法(固定生成预算、逐样本长度目标、组相对长度奖励)对多种模型进行微调,发现它们对忠实性和可监控性具有不同的影响。This work fine-tune a variety of models with three methods that apply length pressure differently, namely a fixed generation budget, a per-example length target, and a group-relative length reward, and finds that it affects faithfulness and monitorability differently.
提出 ProWAM,一种渐进式世界动作模型,联合预测动作和稀疏视觉子目标的有序序列,为任务执行过程中的动作生成提供显式视觉引导,展示了进度索引视觉前瞻对闭环控制的价值。ProWAM is presented, a progressive world action model that jointly predicts actions and an ordered sequence of sparse visual sub-goals, providing explicit visual guidance to anchor action generation throughout task execution, demonstrating the value of progress-indexed visual foresight for closed-loop control.
提出 TQP,一种训练策略,将视频按帧分块处理,块间传递被追踪对象信息,但仅在单个块内进行反向传播,从而在无显存溢出、推理开销或梯度消失问题的情况下实现更长时间范围的监督。TQP is introduced, a training strategy in which the model processes a video in chunks of frames, where information about tracked objects is propagated between chunks but backpropagation is only conducted in individual chunks, enabling longer temporal supervision without out-of-memory issues, inference overhead, or vanishing gradients.
提出 Rollout-Marginal Distillation,在 AR 预测中保留生成历史,但对每个 chunk 独立对照 chunk teacher 打分,确保其质量修正不受不完美的时序上下文影响。Rollout-Marginal Distillation is introduced, which retains the generated history for AR prediction but scores each chunk independently against a chunk teacher, ensuring its quality correction is not compromised by an imperfect temporal context.
一种 3D 世界动作模型,将世界分解为场景与手,并在共享时空坐标系内将二者联合预测为 3D 点轨迹,无需任务特定物体或关键点选择即可在大量人类演示视频上进行有效预训练。A 3D world action model that decomposes the world into a scene and hands and hands, and jointly forecasts both as 3D point trajectories within a shared space-time coordinate frame, which enables effective pre-training on large-scale human demonstration videos without requiring any task-specific object or keypoint selection.
EyeRobot 2.0 通过旋转两个眼动视角将其注视中心对准场景中的 3D 注视点实现物理上的视觉聚焦,并将夹爪信息规范化到以注视点为参考的 SE(3) 坐标系,从而压缩待学习的动作分布规模。EyeRobot 2.0 physically attends to a 3D fixation point in the scene by swiveling two eye viewpoints to center their gaze on it, and takes advantage of fixation by canonicalizing gripper information into a fixation-relative SE(3) frame, which compacts the size of the action distribution to learn.
首次在 World Modeling 领域引入 Agentic Harness 框架:由 pilot agent 负责规划并执行角色行为,director agent 负责随着场景推进合成新的环境要素。The integration of an agentic harness within the domain of world modeling is pioneered, wherein a pilot agent is tasked with planning and executing character behaviors, while a director agent is responsible for synthesizing novel environmental elements as the scene progresses.
提出 FOVEATED,一种即插即用框架,通过随机偏移分配给前文上下文 key 的 Rotary Position Embedding 位置来构建每个句子的聚焦视图,并从理论上分析了 FOVEATED 如何抵消上下文引发的难度低估。FOVEATED, a plug-and-play framework that constructs focused views of each sentence by randomly shifting the Rotary Position Embedding positions assigned to the keys of its preceding context, is proposed and theoretically analyzed how FOVEATED counteracts context-induced difficulty underestimation is analyzed.
数字环境的可编程性远超其 GUI 界面所呈现的程度,且一个极简的、以终端为中心的 harness 是获得更优性能与效率的关键。Digital environments are far more programmable than their GUI interfaces imply, and a minimal terminal-focused harness is the key for better performance and efficiency, according to these results.
JEPA-TTT 在测试时间内持续适配预训练动作条件 Joint-Embedding Predictive Architecture 世界模型的潜在动力学预测器,平均将自回归潜在预测误差降低 83%,并将规划性能相较冻结的 JEPA 世界模型提升 153%。JEPA-TTT, which adapts the latent dynamics predictor of a pretrained action-conditioned Joint-Embedding Predictive Architecture world model throughout test time, reduces autoregressive latent prediction error by 83% on average and improves planning performance by 153% over the frozen JEPA world model.