提出了一个能力门控(competence gate),根据已解决的结果估计领域级的源权重,将不确定的估计向全局权重收缩,并重新校准合并后的预测,为基于已测边际价值的选择性模型使用提供了实用方案。A competence gate is introduced that estimates domain-level source weights from resolved outcomes, shrinks uncertain estimates toward a global weight, and recalibrates the pooled forecast and provides a practical approach for selective model use based on measured marginal value.
论文
1173 张论文卡片 · 方法
本文提出 CATS,一种 self-speculative decoding 框架,在内存受限设备上结合 memory budget 与 parameter offloading pattern 进行级联式 verify 与 correction,在设备峰值显存与单独运行 target model 相当的前提下,最大化 token acceptance rate 与端到端加速比。CATS, a self-speculative decoding framework that conducts cascaded verification and correction based on the memory budget and parameter offloading patterns on memory-limited devices, is proposed, which maximizes token acceptance rate and end-to-end speedup while keeping the peak memory footprint on the device equal to that of the target model alone.
Pelican-Sim 在轨迹、场景、物体、具身和视角变化上的定性泛化能力,凸显其作为通用世界模型模拟器的潜力。Qualitative generalization across trajectory, scene, object, embodiment, and viewpoint shifts highlights Pelican-Sim's potential as a general-purpose world model simulator.
该工作提交于 ECCV 2026 Wearable AI Challenge 的 EgoProactive 赛道,在大模型组排名第一、≤2B 组排名第二,表明对当前任务而言视觉定位比标注量更为重要。This submission to the EgoProactive track of the ECCV 2026 Wearable AI Challenge is presented, which ranked first in the large-model division and second in the<=2B division, suggesting that visual grounding is more important than annotation volume for this task.
该系统仅用一个 2B 视觉-语言模型,通过一次贪心前向传播即可回答关于十分钟第一人称视频的多选题,并仅用大型 Agentic pipeline 1.1% 的参数即达到其 89% 的准确率。The system is a single 2B vision-language model that answers multiple-choice questions about ten-minute egocentric videos in one greedy forward pass, and reaches 89% of the accuracy of the large agentic pipeline using 1.1% of its parameters.
即便在公式、变量与笔记访问条件一致时,向 Program-Solve 接口加入 executor 对不同开源权重模型的帮助程度各异,且无论哪种情况都无法替代经过验证的公式或可靠的变量抽取。Adding an executor to the Program-Solve interface helps some open-weight models more than others even under matched formula, variable and note access, and is not a substitute for verified formulas or reliable variable extraction either way.
提出结构感知的 RAG 框架 ReMoMask-2:耦合 Hierarchical Bidirectional Momentum 对比学习以对齐全局与部件级特征与文本;采用 Semantic Spatial-Temporal Attention (SSTA) 实现拓扑感知的融合;通过 Topology Structured Masking (TSM) 借助自适应掩码强化鲁棒的部件级 grounding。ReMoMask-2, a structure-aware RAG framework coupling Hierarchical Bidirectional Momentum contrastive learning to align global and part-level features with text; Semantic Spatial-Temporal Attention (SSTA) for topology-aware fusion; and Topology Structured Masking (TSM) to force robust part-level grounding via adaptive masking are introduced.
提出 MobileVLA-R1 2.0,一个 RL 增强的 VLA 框架,显式地将结构化具身推理与可执行的移动机器人控制耦合,并引入推理条件化的动作解码器,将多模态推理表征映射到任务级动作目标,再由机器人控制器翻译为具身特定的指令。This work proposes MobileVLA-R1 2.0, an RL-enhanced VLA framework that explicitly couples structured embodied reasoning with executable mobile robot control, and introduces a reasoning-conditioned action decoder that maps multimodal reasoning representations to task-level action targets, which are subsequently translated into embodiment-specific commands by robot controllers.
提出成本高效的协同模型 Occamy-1.0,由后训练 checkpoint Qwen3.6-35B-A3B 继续训练得到,位于所观测成本-性能帕累托前沿的低成本拐点处。Occamy-1.0, a cost-efficient co-work model obtained by further training the post-trained Qwen3.6-35B-A3B checkpoint, is presented and placed at the low-cost knee of the observed cost--performance Pareto frontier.
提出 RSIAgent,一种无需训练、通过自主构建记忆实现递归自我改进的多 Agent 框架,显著增强了强开源模型,使 Kimi-K3 与 GLM-5.3 超越包括 GPT-6.3 在内的前沿闭源模型。RSIAgent is introduced, a training-free multi-agent framework for recursive self-improvement through autonomous memory construction that substantially improves strong open-source models, enabling Kimi-K3 and GLM-5.3 to outperform frontier closed-source models including GPT-6.3.
提出 HazardAuditor,一个执行驱动的框架,在受控环境中运行异构 Agent 并将其交互归一化为规范事件表示以实现跨框架监督;观察到 token 级后训练目标与生成式守卫存在结构性失配,导致更长的推理链主导梯度更新。HazardAuditor is introduced, an execution-grounded framework that runs heterogeneous agents in controlled environments and normalizes their interactions into a canonical event representation for cross-framework supervision, and observes that token-level post-training objectives create a structural mismatch for generative guards, causing longer rationales to dominate gradient updates.
将规模化拐点定义为边际 Elo 增益等于独立采样参考时的单次会话预算;提出 Elo-per-token 分析,跟踪每个 token 预算下找到的最优解,并使用 Bradley-Terry 模型将各任务内的排序聚合为跨不同评分尺度任务的 Elo 评分。This work defines the scaling inflection point as the per-session budget where marginal Elo gains match the independent-sampling reference, and proposes Elo-per-token analysis, which tracks the best solution found at each token budget and uses a Bradley-Terry model to aggregate within-task orderings into Elo ratings across tasks with different score scales.
将投毒选择形式化为 oracle 预算下的集合优化问题,提出 SAILS (Set-level Audit-Informed Iterative Learned Selection),通过数百次微调-评估运行学习集合打分器,对百万级候选集合排序,仅审计少量入围集合。This work formalizes poison selection as oracle-budgeted set optimization and introduces SAILS (Set-level Audit-Informed Iterative Learned Selection), which learns a set scorer from a few hundred finetune-and-evaluate runs, ranks millions of candidate sets, and audits only a small shortlist.
提出结构化设计规范作为探索替代 UI 概念的实用控制点,同时通过将设计方向显式化为中间决策的推理架构,保持下游生成设置固定。This work identifies structured design specifications as a practical control point for exploring alternative UI concepts while keeping downstream generation settings fixed through an inference architecture that makes design direction an explicit intermediate decision.
Realtime-Venus 是一个主动全双工交互系统,由两个独立训练的 9B 模型组成:Realtime-Venus-Omni 用于音视频交互,Realtime-Venus-Audio 用于语音交互;在 MMAU、Llama Questions 与 Speech CMMLU 上均领先于对比模型Realtime-Venus, a proactive full-duplex interaction system with two separately trained 9B models: Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for spoken interaction, which leads the compared models on MMAU, Llama Questions, and Speech CMMLU.
本工作证明通用 agent 可在整个任务执行过程中直接驱动物理机器人,无需任何任务或环境专属训练;并提出 Agent as Policy(AGP),将任务规划与执行置于 agent 的控制之下This work demonstrates that a general-purpose agent can directly drive a physical robot throughout task execution without any task-specific or environment-specific training and introduces Agent as Policy (AGP), which places task planning and execution under the agent's control.
视频扩散模型具有随机性且难以控制:精确内容往往需要反复采样且无法保证成功,长时场景在外观、交互和时间一致性上会发生漂移。Agent式视觉创作可提供显式参考、可编辑的3D场景或可执行的游戏状态以实现稳定控制,但本身并不能保证高对象或角色保真度。二者结合可实现稳定且高质量的生成。为实现该结合,我们提出 LynnReal-Omni,一个基于32B共享多模态Video diffusion models are stochastic and hard to control: precise content often requires repeated sampling without guaranteed success, and long-horizon scenes drift in appearance, interactions, and temporal coherence. Agentic visual creation provides explicit references, editable 3D scenes, or executable game states for stable control, but does not by itself guarantee high object or character fidelity. Combining the two can enable stable, high-quality generation. To realize this combination, we present LynnReal-Omni, a native multimodal video generation framework built on a 32B shared multimo
展示并论证了算法无损性与其在有限精度算术下实现之间存在差距,主张应从精确生成轨迹与下游任务性能两个层面评估无损投机解码A gap between algorithmic losslessness and its implementation under finite-precision arithmetic is demonstrated and motivated, to motivate evaluating lossless speculative decoding at the level of exact generation trajectories as well as downstream task performance.
ModaLens 是一种配对图像交换审计,用于衡量报告可用性如何改变图像敏感性,并限制对视觉正确性的结论;该方向在另外两个模型系列中得到复现。ModaLens, a paired image-swap audit, measures how report availability changes image sensitivity and limits conclusions about visual correctness; the direction replicates in two further model lineages.
本文在 Argonne Leadership Computing Facility 的 Polaris 超级计算机上对分布式向量数据库性能进行了实证研究,选取 Qdrant 评估在最多 32 个 worker 下的插入、索引构建与 query latency。This work presents an empirical study of distributed vector database performance on the Polaris supercomputer in the Argonne Leadership Computing Facility, and selects Qdrant to evaluate insertion, index construction, and query latency with up to 32 workers.
提出 ESRL,一种架构感知的框架,显式探索 MoE 模型的专家路由空间,将高置信度专家保留为锚点,并把随机路由限制在合理候选池内,从而保留可靠的计算路径。ESRL is introduced, an architecture-aware framework that explicitly explores the expert-routing space of MoE models, and preserves high-confidence experts as anchors, and restricts stochastic routing to a plausible candidate pool, thereby retaining reliable computation paths.
该工作提出了最小干预强化学习(Minimal Intervention Reinforcement Learning, MInTRL),通过在原本的 on-policy rollout 中引入稀疏的局部干预来扩展探索边界,确立了最小干预作为增强 on-policy RL 的有效范式。This work introduces Minimal Intervention Reinforcement Learning (MInTRL), which expands the exploration frontier through sparse, local interventions in otherwise on-policy rollouts, and establishes minimal intervention as an effective paradigm for enhancing on-policy RL.
Dynin-Robotics 在 LIBERO 和零样本 LIBERO-Plus 上取得了具有竞争力的性能,并在 Franka Research 3 机器人的四种操作条件下达到了 78.4% 的平均成功率。Dynin-Robotics achieves competitive performance on LIBERO and zero-shot LIBERO-Plus, together with a 78.4% average success rate across four manipulation conditions on a Franka Research 3 robot.
外部评估显示,尽管引用有效性保持稳健,但在领域偏移下证据利用、片段对齐与拒答校准变得更加困难,表明可信的 RAG 系统需要在检索与最终答案交付之间进行显式验证。External evaluation shows that while citation validity remains strong, evidence utilization, span alignment, and refusal calibration become harder under domain shift, indicating that trustworthy RAG systems require explicit validation between retrieval and final answer delivery.
该工作提出了 SCoRE(Selection and Consolidation for Robust Evidence),一个用于显式证据选择与整合的统一 agent 循环,将最终推理与探索式试错解耦,并通过索引化的声明-图像关联确保严格的视觉锚定。This work proposes SCoRE (Selection and Consolidation for Robust Evidence), a unified agent loop for explicit evidence selection and consolidation, which decouples final reasoning from exploratory trial-and-error while ensuring strict visual grounding via indexed claim-to-image linkages.
检索增强生成(RAG)管线通常依赖在预处理阶段确定的固定索引与检索配置。这种一刀切的设计难以适配领域专家场景,因为异构查询需要不同的分块粒度、元数据约束与来源选择策略。因此,针对某一类查询有效的配置,往往在其他类查询上表现欠佳。本文提出 ORDER(Optimal Routing for Dynamic Evidence Retrieval),一种查询条件化的 RAG 框架,可联合自适应地调整索引与……Retrieval-Augmented Generation (RAG) pipelines typically rely on a fixed indexing and retrieval configuration determined at preprocessing time. This one-size-fits-all design is ill-suited to domain-expert settings, where heterogeneous queries require different chunking granularities, metadata constraints, and source-selection strategies. As a result, configurations that are effective for one family of queries often perform poorly for others. In this paper, we introduce ORDER (Optimal Routing for Dynamic Evidence Retrieval), a query-conditioned RAG framework that jointly adapts indexing and ret
论文提出了 InceptionRAG,一种颠覆 RAG 标准攻击范式的隐蔽攻击机制,展现出更优的规避能力,能够有效绕过针对传统单文档注入的既有防御。This paper introduces InceptionRAG, a stealthy attack mechanism that subverts the standard attack paradigm of RAG, and shows superior evasion capabilities, effectively bypassing established defenses that mitigate traditional single-document injections.
论文提出了 ModAR,这是首个在预测动作前以自回归方式对多种未来模态进行去噪的 WAM;研究发现 ModAR 的序列化生成优于现有 WAM 形式,并在所有评估数据规模下取得最高的平均成功率。ModAR is introduced, the first WAM to autoregressively denoise multiple future modalities before predicting actions, and it is found that ModAR's sequential generation outperforms existing WAM formulations, with the highest average success rate at all evaluated data scales.
Mind2Dialogue 框架提出了一个心理学引导的模拟器,在交互过程中保留个人特征并更新心智状态以生成连贯对话;通过对 Oracle 信息充分的回复进行训练,使模型在部署时无需直接访问用户心智状态即可提供帮助。The Mind2Dialogue framework proposes a psychology-guided simulator that preserves personal characteristics while updating mental states through interaction to generate coherent conversations, and trains models on the Oracle's well-informed responses to assist users without direct access to their mental states at deployment.
该工作表明,专用优化隐式地从此潜空间轨迹中进行选择,并建立了一种新的专用训练视角:当缺少 gold reasoning 时,调参选择直接控制传递给下游模型的潜在监督信号。This work shows that specialist optimization implicitly selects from this latent trajectory space, and establishes a new view of specialist training: when gold reasoning is absent, tuning choices directly control the latent supervision passed to downstream models.
该工作利用 Headroom-Closed Index(HCI)揭示现有 LLM 的问题,并提出 RSI 概念及其发展路线图:从改进执行自主性、改进策略自主性、经验获取自主性、环境适应自主性,到递归元改进。This work uses the Headroom-Closed Index (HCI) to reveal the problems of existing LLMs, and introduces the RSI concept and its development roadmap: from improvement-execution autonomy, improvement-strategy autonomy, experience-acquisition autonomy, and environment-adaptation autonomy, to recursive meta-improvement.
本工作提出 HarnessVLN,一个零样本、无需训练的框架:通过共享的 Agent Harness 统一指令跟随与物体目标导航,并展示了其在真实世界中两类导航任务上的适用性This work introduces HarnessVLN, a zero-shot, training-free framework that unifies instruction-following and object-goal navigation through a shared Agent Harness and demonstrates its applicability to both navigation tasks in real-world environments.
本文提出 DualPath,一种 inference 系统,通过引入 dual-path KV-Cache loading 打破瓶颈,并实现一条新的 storage-to-decode 路径:KV-Cache 先加载到 decode engine,再通过 compute network 上的 RDMA 高效转发至 prefill engine。DualPath is presented, an inference system that breaks this bottleneck by introducing dual-path KV-Cache loading and enables a novel storage-to-decode path, in which the KV-Cache is loaded into decoding engines and then efficiently transferred to prefill engines via RDMA over the compute network.
StepAudio 3 Realtime 是一个围绕持续 listen-converse-think-act 循环组织的音频-语言基础模型,在实时语音输出的同时达到与专用推理模型相当的对话与推理性能,并通过 Think-While-Speaking 机制化解深度思考与延迟之间的矛盾。StepAudio 3 Realtime, an audio-language foundation model organized around a continuous listen-converse-think-act loop, achieves dialogue and reasoning performance comparable to dedicated reasoning models while speaking in real time, and resolves the tension between deep deliberation and latency via Think-While-Speaking.
该工作提出了 StepAudio 3 Music,一个支持显式音乐规划与开放域文本控制生成的大规模长篇音乐生成模型,在所评估系统中取得最高的 AudioBox Content Enjoyment、Content Usefulness 与 Production Quality 分数,以及最高的 MuQ-MuLan 相似度。This work introduces StepAudio 3 Music, a large-scale, long-form music generation model that supports explicit musical planning and open-domain text-controlled generation, and achieves the highest AudioBox Content Enjoyment, Content Usefulness, and Production Quality scores and the highest MuQ-MuLan similarity among the evaluated systems.
本文重新审视联合多模态表示学习与生成,旨在产生可直接被生成式解码器使用的线性可插值嵌入,并确保其表示同时充当判别性语义描述符和生成条件。This work revisits joint multimodal representation learning and generation to produce linearly interpolatable embeddings that are directly consumable by generative decoders and ensures its representations function as both discriminative semantic descriptors and generative conditions.