提出成本高效的协同模型 Occamy-1.0,由后训练 checkpoint Qwen3.6-35B-A3B 继续训练得到,位于所观测成本-性能帕累托前沿的低成本拐点处。Occamy-1.0, a cost-efficient co-work model obtained by further training the post-trained Qwen3.6-35B-A3B checkpoint, is presented and placed at the low-cost knee of the observed cost--performance Pareto frontier.
论文
1640 张论文卡片 · OA 绿色
推出 Atria Dawn Preview,一个面向科学研究与工程工作流的基础 Agent 语言模型,旨在拓展真实场景下 Agent 生产力的前沿。Atria Dawn Preview is introduced, a foundation agentic language model designed for scientific research and engineering workflows, with the goal of expanding the frontier of agent productivity in the real world.
本文探讨 multi-agent system,并指出当前尚未被充分解决的问题,同时探索了 multi-agent system 在区块链系统中的潜在应用,为其在真实分布式系统中的未来发展与落地提供启示。This paper explores multi-agent systems and identifies challenges that remain inadequately addressed, and explores potential applications of multi-agent systems in blockchain systems to shed light on their future development and application in real-world distributed systems.
提出 RSIAgent,一种无需训练、通过自主构建记忆实现递归自我改进的多 Agent 框架,显著增强了强开源模型,使 Kimi-K3 与 GLM-5.3 超越包括 GPT-6.3 在内的前沿闭源模型。RSIAgent is introduced, a training-free multi-agent framework for recursive self-improvement through autonomous memory construction that substantially improves strong open-source models, enabling Kimi-K3 and GLM-5.3 to outperform frontier closed-source models including GPT-6.3.
提出 HazardAuditor,一个执行驱动的框架,在受控环境中运行异构 Agent 并将其交互归一化为规范事件表示以实现跨框架监督;观察到 token 级后训练目标与生成式守卫存在结构性失配,导致更长的推理链主导梯度更新。HazardAuditor is introduced, an execution-grounded framework that runs heterogeneous agents in controlled environments and normalizes their interactions into a canonical event representation for cross-framework supervision, and observes that token-level post-training objectives create a structural mismatch for generative guards, causing longer rationales to dominate gradient updates.
将规模化拐点定义为边际 Elo 增益等于独立采样参考时的单次会话预算;提出 Elo-per-token 分析,跟踪每个 token 预算下找到的最优解,并使用 Bradley-Terry 模型将各任务内的排序聚合为跨不同评分尺度任务的 Elo 评分。This work defines the scaling inflection point as the per-session budget where marginal Elo gains match the independent-sampling reference, and proposes Elo-per-token analysis, which tracks the best solution found at each token budget and uses a Bradley-Terry model to aggregate within-task orderings into Elo ratings across tasks with different score scales.
将投毒选择形式化为 oracle 预算下的集合优化问题,提出 SAILS (Set-level Audit-Informed Iterative Learned Selection),通过数百次微调-评估运行学习集合打分器,对百万级候选集合排序,仅审计少量入围集合。This work formalizes poison selection as oracle-budgeted set optimization and introduces SAILS (Set-level Audit-Informed Iterative Learned Selection), which learns a set scorer from a few hundred finetune-and-evaluate runs, ranks millions of candidate sets, and audits only a small shortlist.
提出结构化设计规范作为探索替代 UI 概念的实用控制点,同时通过将设计方向显式化为中间决策的推理架构,保持下游生成设置固定。This work identifies structured design specifications as a practical control point for exploring alternative UI concepts while keeping downstream generation settings fixed through an inference architecture that makes design direction an explicit intermediate decision.
Realtime-Venus 是一个主动全双工交互系统,由两个独立训练的 9B 模型组成:Realtime-Venus-Omni 用于音视频交互,Realtime-Venus-Audio 用于语音交互;在 MMAU、Llama Questions 与 Speech CMMLU 上均领先于对比模型Realtime-Venus, a proactive full-duplex interaction system with two separately trained 9B models: Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for spoken interaction, which leads the compared models on MMAU, Llama Questions, and Speech CMMLU.
本工作证明通用 agent 可在整个任务执行过程中直接驱动物理机器人,无需任何任务或环境专属训练;并提出 Agent as Policy(AGP),将任务规划与执行置于 agent 的控制之下This work demonstrates that a general-purpose agent can directly drive a physical robot throughout task execution without any task-specific or environment-specific training and introduces Agent as Policy (AGP), which places task planning and execution under the agent's control.
视频扩散模型具有随机性且难以控制:精确内容往往需要反复采样且无法保证成功,长时场景在外观、交互和时间一致性上会发生漂移。Agent式视觉创作可提供显式参考、可编辑的3D场景或可执行的游戏状态以实现稳定控制,但本身并不能保证高对象或角色保真度。二者结合可实现稳定且高质量的生成。为实现该结合,我们提出 LynnReal-Omni,一个基于32B共享多模态Video diffusion models are stochastic and hard to control: precise content often requires repeated sampling without guaranteed success, and long-horizon scenes drift in appearance, interactions, and temporal coherence. Agentic visual creation provides explicit references, editable 3D scenes, or executable game states for stable control, but does not by itself guarantee high object or character fidelity. Combining the two can enable stable, high-quality generation. To realize this combination, we present LynnReal-Omni, a native multimodal video generation framework built on a 32B shared multimo
展示并论证了算法无损性与其在有限精度算术下实现之间存在差距,主张应从精确生成轨迹与下游任务性能两个层面评估无损投机解码A gap between algorithmic losslessness and its implementation under finite-precision arithmetic is demonstrated and motivated, to motivate evaluating lossless speculative decoding at the level of exact generation trajectories as well as downstream task performance.
ModaLens 是一种配对图像交换审计,用于衡量报告可用性如何改变图像敏感性,并限制对视觉正确性的结论;该方向在另外两个模型系列中得到复现。ModaLens, a paired image-swap audit, measures how report availability changes image sensitivity and limits conclusions about visual correctness; the direction replicates in two further model lineages.
本文在 Argonne Leadership Computing Facility 的 Polaris 超级计算机上对分布式向量数据库性能进行了实证研究,选取 Qdrant 评估在最多 32 个 worker 下的插入、索引构建与 query latency。This work presents an empirical study of distributed vector database performance on the Polaris supercomputer in the Argonne Leadership Computing Facility, and selects Qdrant to evaluate insertion, index construction, and query latency with up to 32 workers.
提出 ESRL,一种架构感知的框架,显式探索 MoE 模型的专家路由空间,将高置信度专家保留为锚点,并把随机路由限制在合理候选池内,从而保留可靠的计算路径。ESRL is introduced, an architecture-aware framework that explicitly explores the expert-routing space of MoE models, and preserves high-confidence experts as anchors, and restricts stochastic routing to a plausible candidate pool, thereby retaining reliable computation paths.
该工作提出了最小干预强化学习(Minimal Intervention Reinforcement Learning, MInTRL),通过在原本的 on-policy rollout 中引入稀疏的局部干预来扩展探索边界,确立了最小干预作为增强 on-policy RL 的有效范式。This work introduces Minimal Intervention Reinforcement Learning (MInTRL), which expands the exploration frontier through sparse, local interventions in otherwise on-policy rollouts, and establishes minimal intervention as an effective paradigm for enhancing on-policy RL.
Dynin-Robotics 在 LIBERO 和零样本 LIBERO-Plus 上取得了具有竞争力的性能,并在 Franka Research 3 机器人的四种操作条件下达到了 78.4% 的平均成功率。Dynin-Robotics achieves competitive performance on LIBERO and zero-shot LIBERO-Plus, together with a 78.4% average success rate across four manipulation conditions on a Franka Research 3 robot.
该工作提出了分组值注意力(Grouped Value Attention, GVA),通过存储分组值并在推理时利用可学习的线性映射重构内容键,该映射可被吸收进 query 中,从而在预期解码路径中无需显式物化内容键。This work introduces Grouped Value Attention (GVA), which stores grouped values and reconstructs content keys with a learned linear map at inference, which can be absorbed into the query, eliminating the need to materialize content keys in the intended decode path.
外部评估显示,尽管引用有效性保持稳健,但在领域偏移下证据利用、片段对齐与拒答校准变得更加困难,表明可信的 RAG 系统需要在检索与最终答案交付之间进行显式验证。External evaluation shows that while citation validity remains strong, evidence utilization, span alignment, and refusal calibration become harder under domain shift, indicating that trustworthy RAG systems require explicit validation between retrieval and final answer delivery.
该工作提出了 SCoRE(Selection and Consolidation for Robust Evidence),一个用于显式证据选择与整合的统一 agent 循环,将最终推理与探索式试错解耦,并通过索引化的声明-图像关联确保严格的视觉锚定。This work proposes SCoRE (Selection and Consolidation for Robust Evidence), a unified agent loop for explicit evidence selection and consolidation, which decouples final reasoning from exploratory trial-and-error while ensuring strict visual grounding via indexed claim-to-image linkages.
论文提出了 E2A-Bench,一个面向金融图表推理的 969 查询基准,由 323 个 HS300 成分股在三种输入模态下构建,并附带由 OHLCV 确定性派生的证据锚点;结果表明金融 VLM 评估应追溯从证据到决策的完整链路,而非依赖单一幻觉分数。E2A-Bench is introduced, a 969-query benchmark for financial chart reasoning, constructed from 323 HS300 constituents under three input modalities with deterministic OHLCV-derived evidence anchors, and results show that financial VLM evaluation should trace the full evidence-to-action chain rather than rely on a single hallucination score.
提出 Continual Search,一个迭代框架:在多轮对话中持续推动判别器搜索尚未解决的诊断证据,在多个基准测试套件和模型系列上一致提升归因性能Continual Search is introduced, an iterative framework that nudges the judge, over successive turns, to keep searching for unresolved diagnostic evidence, which consistently improves attribution performance across multiple benchmark suites and model families.
本文提出一个无偏的多进程 evaluation 框架,能够有效分散 client 端负载,从而在每秒数千次 query 以上的生产规模下实现对 LLM 的精确、可复现 profiling。This work proposes an unbiased, multi-process evaluation framework that effectively distributes client-side load, enabling accurate, reproducible profiling of LLMs at production scales exceeding thousands of queries per second.
结果表明,模型级对齐并不具备可组合性:单独能力强且看似安全的 agent 在组成系统后,可能随着 AI 的持续性与互联化而产生性质上不同的失效模式。The results suggest that model-level alignment is not compositional: individually capable and apparently safe agents can form systems with qualitatively different failure modes as AI becomes persistent and interconnected.
检索增强生成(RAG)管线通常依赖在预处理阶段确定的固定索引与检索配置。这种一刀切的设计难以适配领域专家场景,因为异构查询需要不同的分块粒度、元数据约束与来源选择策略。因此,针对某一类查询有效的配置,往往在其他类查询上表现欠佳。本文提出 ORDER(Optimal Routing for Dynamic Evidence Retrieval),一种查询条件化的 RAG 框架,可联合自适应地调整索引与……Retrieval-Augmented Generation (RAG) pipelines typically rely on a fixed indexing and retrieval configuration determined at preprocessing time. This one-size-fits-all design is ill-suited to domain-expert settings, where heterogeneous queries require different chunking granularities, metadata constraints, and source-selection strategies. As a result, configurations that are effective for one family of queries often perform poorly for others. In this paper, we introduce ORDER (Optimal Routing for Dynamic Evidence Retrieval), a query-conditioned RAG framework that jointly adapts indexing and ret
论文提出了 ModAR,这是首个在预测动作前以自回归方式对多种未来模态进行去噪的 WAM;研究发现 ModAR 的序列化生成优于现有 WAM 形式,并在所有评估数据规模下取得最高的平均成功率。ModAR is introduced, the first WAM to autoregressively denoise multiple future modalities before predicting actions, and it is found that ModAR's sequential generation outperforms existing WAM formulations, with the highest average success rate at all evaluated data scales.
PhysStream 是一个用于物理驱动图像到视频合成的自回归模型,通过引入结构化场景记忆并支持基于稀疏速度增量信号的细粒度运动控制(编码物理量),使模型能够学习底层动力学。PhysStream is an autoregressive model for physics-grounded image-to-video synthesis that incorporates structured scene memory and supports fine-grained motion control via sparse velocity-increment signals that encode physical quantities, letting the model learn the underlying dynamics.
Mind2Dialogue 框架提出了一个心理学引导的模拟器,在交互过程中保留个人特征并更新心智状态以生成连贯对话;通过对 Oracle 信息充分的回复进行训练,使模型在部署时无需直接访问用户心智状态即可提供帮助。The Mind2Dialogue framework proposes a psychology-guided simulator that preserves personal characteristics while updating mental states through interaction to generate coherent conversations, and trains models on the Oracle's well-informed responses to assist users without direct access to their mental states at deployment.
该工作表明,专用优化隐式地从此潜空间轨迹中进行选择,并建立了一种新的专用训练视角:当缺少 gold reasoning 时,调参选择直接控制传递给下游模型的潜在监督信号。This work shows that specialist optimization implicitly selects from this latent trajectory space, and establishes a new view of specialist training: when gold reasoning is absent, tuning choices directly control the latent supervision passed to downstream models.
该工作利用 Headroom-Closed Index(HCI)揭示现有 LLM 的问题,并提出 RSI 概念及其发展路线图:从改进执行自主性、改进策略自主性、经验获取自主性、环境适应自主性,到递归元改进。This work uses the Headroom-Closed Index (HCI) to reveal the problems of existing LLMs, and introduces the RSI concept and its development roadmap: from improvement-execution autonomy, improvement-strategy autonomy, experience-acquisition autonomy, and environment-adaptation autonomy, to recursive meta-improvement.
本工作提出 HarnessVLN,一个零样本、无需训练的框架:通过共享的 Agent Harness 统一指令跟随与物体目标导航,并展示了其在真实世界中两类导航任务上的适用性This work introduces HarnessVLN, a zero-shot, training-free framework that unifies instruction-following and object-goal navigation through a shared Agent Harness and demonstrates its applicability to both navigation tasks in real-world environments.
本文提出 DualPath,一种 inference 系统,通过引入 dual-path KV-Cache loading 打破瓶颈,并实现一条新的 storage-to-decode 路径:KV-Cache 先加载到 decode engine,再通过 compute network 上的 RDMA 高效转发至 prefill engine。DualPath is presented, an inference system that breaks this bottleneck by introducing dual-path KV-Cache loading and enables a novel storage-to-decode path, in which the KV-Cache is loaded into decoding engines and then efficiently transferred to prefill engines via RDMA over the compute network.
论文提出了 ModularRSI,一个与基准解耦、对比式、模块化的可泛化 harness 进化框架,对同一任务下成功与失败的轨迹进行对比,并跨任务聚合证据以识别反复出现的行为缺陷。ModularRSI is proposed, a benchmark-disjoint, contrastive, and modular framework for generalizable harness evolution that contrasts successful and failed trajectories for the same task and aggregates evidence across tasks to identify recurring behavioral deficiencies.
论文报告了来自 OpenAI、Anthropic、xAI 与 Google DeepMind 的六类前沿模型实验,并使用 epistemic jailbreak 一词来指称随请求具体性增加而伴随出现的技术溯源严谨性丧失现象。This paper reports experiments across six frontier model types from OpenAI, Anthropic, xAI, and Google DeepMind, and uses the term epistemic jailbreak for the accompanying loss of discipline in technical provenance as requested specificity increases.
StepAudio 3 Realtime 是一个围绕持续 listen-converse-think-act 循环组织的音频-语言基础模型,在实时语音输出的同时达到与专用推理模型相当的对话与推理性能,并通过 Think-While-Speaking 机制化解深度思考与延迟之间的矛盾。StepAudio 3 Realtime, an audio-language foundation model organized around a continuous listen-converse-think-act loop, achieves dialogue and reasoning performance comparable to dedicated reasoning models while speaking in real time, and resolves the tension between deep deliberation and latency via Think-While-Speaking.
该工作提出了 StepAudio 3 Music,一个支持显式音乐规划与开放域文本控制生成的大规模长篇音乐生成模型,在所评估系统中取得最高的 AudioBox Content Enjoyment、Content Usefulness 与 Production Quality 分数,以及最高的 MuQ-MuLan 相似度。This work introduces StepAudio 3 Music, a large-scale, long-form music generation model that supports explicit musical planning and open-domain text-controlled generation, and achieves the highest AudioBox Content Enjoyment, Content Usefulness, and Production Quality scores and the highest MuQ-MuLan similarity among the evaluated systems.