LIT(Latent Interface Training)是一个与框架无关的两阶段策略:先在无图像条件下建立空间目标条件化的动作先验,再通过姿态监督的潜在接口约束视觉条件化,可在保持或提升 LIBERO 平均成功率的同时改善 LIBERO-Plus 综合成功率。Latent Interface Training (LIT), a framework-agnostic two-stage strategy that first establishes a spatial-goal-conditioned action prior without images, then constrains visual conditioning through a pose-supervised latent interface, improves overall LIBERO-Plus success while preserving or improving average LIBERO success.
论文
1753 张论文卡片
在推理、长上下文理解和 Agent 任务中,SAS 在不同注意力预算下均优于可训练的稀疏注意力基线,在紧预算下增益尤为显著,表明其上下文排序对下游任务更有效。Across reasoning, long-context understanding, and agentic tasks, SAS consistently outperforms trainable sparse attention baselines across attention budgets, with especially large gains under tight budgets, demonstrating more effective context ranking for downstream tasks.
提出一种物理引导框架,用于提升单图像部件感知 3D 生成效果,使其几何结构物理相容且连接稳定,并在部件接触面引入参数化连接器。This work proposes a physics-guided framework for improving single-image part-aware 3D generation with physically compatible geometry and stable connections, and introduces parameterized connectors at their contact surfaces.
本文提出 Benchmark Radar,一个用于 AI 基准检索与发现的活体数据库和搜索引擎,覆盖 LLM 评测、Agent 与工具使用基准、代码、推理、安全及领域评测。Benchmark Radar is presented, a living database and search engine for retrieval and discovery of AI benchmarks, covering LLM evaluation, agentic and tool-use benchmarks, coding, reasoning, safety, and domain-specific evaluations.
COBRA-Skills 是一个高效框架,将技能优化建模为在动态演化的候选空间上的预算式序贯优化,对 Agent harness 变化保持鲁棒,并在目标模型自身用于技能生成与优化时依然有效。COBRA-Skills is introduced, an efficient framework that formulates skill optimization as budgeted sequential optimization over a dynamically evolving candidate space and remains robust to changes in the agent harness and performs effectively when the target model itself is used for skill generation and refinement.
本文聚焦 LLM 驱动的 kernel generation 领域,给出现有方法的结构化综述,涵盖 LLM-based 方法与 agentic optimization workflow,并系统梳理了支撑该领域学习与评测的数据集与 benchmark。This survey addresses the gap in LLM-driven kernel generation by providing a structured overview of existing approaches, spanning LLM-based approaches and agentic optimization workflows, and systematically organizing the datasets and benchmarks that underpin learning and evaluation in this domain.
PLC-DPO 将每个偏好对的训练信号路由为 clean、flip 或 tie 三类,将噪声偏好学习从单纯过滤可疑样本重构为主动修正监督方向与强度。PLC-DPO is proposed to robustly optimize preferences by routing each pair's training signal as a clean, flip, or tie case, which reframes noisy preference learning as actively correcting supervision direction and strength rather than merely filtering suspicious examples.
本研究提出了 StepAudio 3 Gen,一个通用音频生成模型,在统一框架下支持零样本文本到语音合成(TTS)、声音设计、人声生成、音效、音乐、风格化语音以及多种音频类型的混合生成。This study introduces StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types within a unified framework.
这是首个端到端证明:一个七人独立团队可以训练出在 agentic 网络安全能力上领先的开源权重模型,且全部三个 checkpoint 在相近参数规模下均排名第一。This is the first end-to-end demonstration that a seven-person independent team can train open-weight models with leading agentic cyber capability and all three checkpoints rank 1st among models at comparable parameter scales.
提出了 DataFlex-RL,一个在统一 GRPO 方案下比较数据策略的评测平台;研究发现,改变数据策略会显著影响训练过程,但相对于均匀训练并未带来可复现的提升。DataFlex-RL, an evaluation platform for comparing choices under a common GRPO recipe, is introduced, finding that changing the data policy measurably changes the training process but does not produce a reproducible improvement over uniform training.
本研究探索在反馈有限的在线场景下,将 prompt 自适应路由到大语言模型专家以最大化响应质量,并提出了策略性地选择和观察奖励以最小化遗憾的算法。This work studies adaptive routing of prompts to large language model experts to maximize response quality in an online setting with limited feedback and proposes algorithms that strategically select and observe rewards to minimize regret.
提出了 Diverse Skill Routing,一个具备多样性感知能力的重排序框架,使用 Determinantal Point Process 在相关性与非冗余性之间取得平衡,在强 pointwise 重排序基线之上提升了召回率与完整覆盖率,且在多技能 query 上增益更大。Diverse Skill Routing is proposed, a diversity-aware reranking framework that uses a Determinantal Point Process to balance relevance and non-redundancy and improves recall and full coverage over a strong pointwise reranking baseline, with larger gains on multi-skill queries.
最终得到的 82M 参数模型 Wayu-Paxa-TTS-Edge 实现了无需参考音频的设备端泰语 TTS,并在三个系统中取得了最低的停顿位置错误率与词内停顿率。The resulting 82M-parameter model, Wayu-Paxa-TTS-Edge, enables on-device Thai TTS without reference audio and achieves the lowest pause-placement error and intra-word pause rates among the three systems.
研究发现,仅使用合成数据的训练即可具备竞争力;将 0.9B 参数的 PaddleOCR-VL-1.6 适配为 Wayu-Paxa-OCR-Zero,一个无需真实泰语文档页 OCR 标签即可适配的泰语 OCR 模型,表明仅合成数据训练即可具备竞争力。It is found that synthetic-only training can be competitive, and the 0.9B-parameter PaddleOCR-VL-1.6 is adapted into Wayu-Paxa-OCR-Zero, a Thai OCR model adapted without OCR labels from real Thai document pages, showing that synthetic-only training can be competitive.
提出面向检索增强生成(RAG)的通用微调方案——RAG 模型融合预训练参数化记忆与非参数化记忆进行语言生成;研究发现,相较 SOTA 的纯参数化 seq2seq 基线,RAG 模型生成的文本更具针对性、更多样且更符合事实。A general-purpose fine-tuning recipe for retrieval-augmented generation (RAG) -- models which combine pre-trained parametric and non-parametric memory for language generation, and finds that RAG models generate more specific, diverse and factual language than a state-of-the-art parametric-only seq2seq baseline.
提出了一个能力门控(competence gate),根据已解决的结果估计领域级的源权重,将不确定的估计向全局权重收缩,并重新校准合并后的预测,为基于已测边际价值的选择性模型使用提供了实用方案。A competence gate is introduced that estimates domain-level source weights from resolved outcomes, shrinks uncertain estimates toward a global weight, and recalibrates the pooled forecast and provides a practical approach for selective model use based on measured marginal value.
本文提出 CATS,一种 self-speculative decoding 框架,在内存受限设备上结合 memory budget 与 parameter offloading pattern 进行级联式 verify 与 correction,在设备峰值显存与单独运行 target model 相当的前提下,最大化 token acceptance rate 与端到端加速比。CATS, a self-speculative decoding framework that conducts cascaded verification and correction based on the memory budget and parameter offloading patterns on memory-limited devices, is proposed, which maximizes token acceptance rate and end-to-end speedup while keeping the peak memory footprint on the device equal to that of the target model alone.
Pelican-Sim 在轨迹、场景、物体、具身和视角变化上的定性泛化能力,凸显其作为通用世界模型模拟器的潜力。Qualitative generalization across trajectory, scene, object, embodiment, and viewpoint shifts highlights Pelican-Sim's potential as a general-purpose world model simulator.
提出即插即用的 Feature Recovery Module (FRM),可在保持宿主网络冻结的前提下,将退化编码器特征映射到与干净图像对齐的表征;该模块提升了场景级检测、CLIP/SigLIP2 特征恢复以及全部四项物体级 VLM 任务,且退化越严重增益越大。The Feature Recovery Module (FRM), a plug-and-play module that maps degraded encoder features to pristine-aligned representations while keeping the host frozen, is proposed, which improves scene-level detection, CLIP/SigLIP2 feature recovery, and all four object-level VLM tasks, with larger gains under more severe degradation.
该工作提交于 ECCV 2026 Wearable AI Challenge 的 EgoProactive 赛道,在大模型组排名第一、≤2B 组排名第二,表明对当前任务而言视觉定位比标注量更为重要。This submission to the EgoProactive track of the ECCV 2026 Wearable AI Challenge is presented, which ranked first in the large-model division and second in the<=2B division, suggesting that visual grounding is more important than annotation volume for this task.
该系统仅用一个 2B 视觉-语言模型,通过一次贪心前向传播即可回答关于十分钟第一人称视频的多选题,并仅用大型 Agentic pipeline 1.1% 的参数即达到其 89% 的准确率。The system is a single 2B vision-language model that answers multiple-choice questions about ten-minute egocentric videos in one greedy forward pass, and reaches 89% of the accuracy of the large agentic pipeline using 1.1% of its parameters.
对突发物理危险的反应既是具身智能的重要检验,也是将多模态大语言模型 (MLLMs) 部署为家庭机器人决策核心的硬性要求;ReactHuman 是首个面向类人反应式决策的物理驱动基准。Reacting to sudden physical hazards is both a meaningful test of embodied intelligence and a hard requirement for deploying multimodal large language models (MLLMs) as the decision core of household robots, and ReactHuman is the first physics-grounded benchmark for human-like reactive decision-making.
即便在公式、变量与笔记访问条件一致时,向 Program-Solve 接口加入 executor 对不同开源权重模型的帮助程度各异,且无论哪种情况都无法替代经过验证的公式或可靠的变量抽取。Adding an executor to the Program-Solve interface helps some open-weight models more than others even under matched formula, variable and note access, and is not a substitute for verified formulas or reliable variable extraction either way.
提出结构感知的 RAG 框架 ReMoMask-2:耦合 Hierarchical Bidirectional Momentum 对比学习以对齐全局与部件级特征与文本;采用 Semantic Spatial-Temporal Attention (SSTA) 实现拓扑感知的融合;通过 Topology Structured Masking (TSM) 借助自适应掩码强化鲁棒的部件级 grounding。ReMoMask-2, a structure-aware RAG framework coupling Hierarchical Bidirectional Momentum contrastive learning to align global and part-level features with text; Semantic Spatial-Temporal Attention (SSTA) for topology-aware fusion; and Topology Structured Masking (TSM) to force robust part-level grounding via adaptive masking are introduced.
提出 MobileVLA-R1 2.0,一个 RL 增强的 VLA 框架,显式地将结构化具身推理与可执行的移动机器人控制耦合,并引入推理条件化的动作解码器,将多模态推理表征映射到任务级动作目标,再由机器人控制器翻译为具身特定的指令。This work proposes MobileVLA-R1 2.0, an RL-enhanced VLA framework that explicitly couples structured embodied reasoning with executable mobile robot control, and introduces a reasoning-conditioned action decoder that maps multimodal reasoning representations to task-level action targets, which are subsequently translated into embodiment-specific commands by robot controllers.
提出成本高效的协同模型 Occamy-1.0,由后训练 checkpoint Qwen3.6-35B-A3B 继续训练得到,位于所观测成本-性能帕累托前沿的低成本拐点处。Occamy-1.0, a cost-efficient co-work model obtained by further training the post-trained Qwen3.6-35B-A3B checkpoint, is presented and placed at the low-cost knee of the observed cost--performance Pareto frontier.
推出 Atria Dawn Preview,一个面向科学研究与工程工作流的基础 Agent 语言模型,旨在拓展真实场景下 Agent 生产力的前沿。Atria Dawn Preview is introduced, a foundation agentic language model designed for scientific research and engineering workflows, with the goal of expanding the frontier of agent productivity in the real world.
本文探讨 multi-agent system,并指出当前尚未被充分解决的问题,同时探索了 multi-agent system 在区块链系统中的潜在应用,为其在真实分布式系统中的未来发展与落地提供启示。This paper explores multi-agent systems and identifies challenges that remain inadequately addressed, and explores potential applications of multi-agent systems in blockchain systems to shed light on their future development and application in real-world distributed systems.
提出 RSIAgent,一种无需训练、通过自主构建记忆实现递归自我改进的多 Agent 框架,显著增强了强开源模型,使 Kimi-K3 与 GLM-5.3 超越包括 GPT-6.3 在内的前沿闭源模型。RSIAgent is introduced, a training-free multi-agent framework for recursive self-improvement through autonomous memory construction that substantially improves strong open-source models, enabling Kimi-K3 and GLM-5.3 to outperform frontier closed-source models including GPT-6.3.
提出 HazardAuditor,一个执行驱动的框架,在受控环境中运行异构 Agent 并将其交互归一化为规范事件表示以实现跨框架监督;观察到 token 级后训练目标与生成式守卫存在结构性失配,导致更长的推理链主导梯度更新。HazardAuditor is introduced, an execution-grounded framework that runs heterogeneous agents in controlled environments and normalizes their interactions into a canonical event representation for cross-framework supervision, and observes that token-level post-training objectives create a structural mismatch for generative guards, causing longer rationales to dominate gradient updates.
将规模化拐点定义为边际 Elo 增益等于独立采样参考时的单次会话预算;提出 Elo-per-token 分析,跟踪每个 token 预算下找到的最优解,并使用 Bradley-Terry 模型将各任务内的排序聚合为跨不同评分尺度任务的 Elo 评分。This work defines the scaling inflection point as the per-session budget where marginal Elo gains match the independent-sampling reference, and proposes Elo-per-token analysis, which tracks the best solution found at each token budget and uses a Bradley-Terry model to aggregate within-task orderings into Elo ratings across tasks with different score scales.
将投毒选择形式化为 oracle 预算下的集合优化问题,提出 SAILS (Set-level Audit-Informed Iterative Learned Selection),通过数百次微调-评估运行学习集合打分器,对百万级候选集合排序,仅审计少量入围集合。This work formalizes poison selection as oracle-budgeted set optimization and introduces SAILS (Set-level Audit-Informed Iterative Learned Selection), which learns a set scorer from a few hundred finetune-and-evaluate runs, ranks millions of candidate sets, and audits only a small shortlist.
提出结构化设计规范作为探索替代 UI 概念的实用控制点,同时通过将设计方向显式化为中间决策的推理架构,保持下游生成设置固定。This work identifies structured design specifications as a practical control point for exploring alternative UI concepts while keeping downstream generation settings fixed through an inference architecture that makes design direction an explicit intermediate decision.
Realtime-Venus 是一个主动全双工交互系统,由两个独立训练的 9B 模型组成:Realtime-Venus-Omni 用于音视频交互,Realtime-Venus-Audio 用于语音交互;在 MMAU、Llama Questions 与 Speech CMMLU 上均领先于对比模型Realtime-Venus, a proactive full-duplex interaction system with two separately trained 9B models: Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for spoken interaction, which leads the compared models on MMAU, Llama Questions, and Speech CMMLU.
本工作证明通用 agent 可在整个任务执行过程中直接驱动物理机器人,无需任何任务或环境专属训练;并提出 Agent as Policy(AGP),将任务规划与执行置于 agent 的控制之下This work demonstrates that a general-purpose agent can directly drive a physical robot throughout task execution without any task-specific or environment-specific training and introduces Agent as Policy (AGP), which places task planning and execution under the agent's control.
视频扩散模型具有随机性且难以控制:精确内容往往需要反复采样且无法保证成功,长时场景在外观、交互和时间一致性上会发生漂移。Agent式视觉创作可提供显式参考、可编辑的3D场景或可执行的游戏状态以实现稳定控制,但本身并不能保证高对象或角色保真度。二者结合可实现稳定且高质量的生成。为实现该结合,我们提出 LynnReal-Omni,一个基于32B共享多模态Video diffusion models are stochastic and hard to control: precise content often requires repeated sampling without guaranteed success, and long-horizon scenes drift in appearance, interactions, and temporal coherence. Agentic visual creation provides explicit references, editable 3D scenes, or executable game states for stable control, but does not by itself guarantee high object or character fidelity. Combining the two can enable stable, high-quality generation. To realize this combination, we present LynnReal-Omni, a native multimodal video generation framework built on a 32B shared multimo