本文提出 TTKV,一种 KV cache 管理框架,将人类记忆系统映射到具备异构容量与精度的 KV cache 上,在 128K 上下文任务上将跨层流量降低 5.94 倍。TTKV is proposed, a KV cache management framework that maps the human memory system onto the KV cache with heterogeneous capacity and precision, and reduces cross-tier traffic by 5.94x on 128K-context tasks.
论文
1173 张论文卡片 · 方法
提出 PARTS(Policy Adaptation with RL on Targeted Subtasks),一种真实世界子任务强化学习框架,能够将训练集中于瓶颈环节,同时以最小的人工干预推进训练 rollout。PARTS (Policy Adaptation with RL on Targeted Subtasks), a real-world subtask RL framework that concentrates practice at bottlenecks while allowing training rollouts to proceed with minimal human intervention, is presented.
设计了一个工具增强的 BI-Agent,将 BI 工作流分解为针对结构化数据的子任务(如 search、join、transform),并在各 BI 阶段编排专门的数据管理方法;与此同时开发了一套后训练框架,从真实 BI 项目中合成训练轨迹,使 BI-Agent 能够通过监督微调(SFT)和强化学习(RL)进行进一步后训练。A tool-augmented BI-Agent is designed that decomposes BI workflows into subtasks on structured data, such as search, join, and transform, and orchestrates specialized data management methods across BI stages, and a post-training framework is developed that synthesizes training trajectories from real BI projects, enabling BI-Agent to be further post-trained using both supervised fine-tuning (SFT) and reinforcement learning (RL).
改进表明,显式的图世界模型调度层能够显著提升长程具身规划在紧凑型和前沿托管型 LLM 能力下的可靠性和效率。Improvements show that an explicit graph world model harness can substantially improve the reliability and efficiency of long-horizon embodied planning across compact and frontier hosted LLM capabilities.
在更大规模 dense 模型与 MoE 模型上的比较进一步凸显了在持续生成过程中调度 prompt 处理的重要性:在并发负载下,vllm-metal 的 packed prefill-decode 路径维持了低于 omlx 的首 token 延迟。Comparisons on larger dense and MoE models reinforce the importance of scheduling prompt processing alongside ongoing generation: vllm-metal's packed prefill-decode path maintains lower first-token latency than omlx under concurrent load.
提出 SteerDuplex,一种基于 Moshi 的全双工语音模型,在自然对话和针对指令遵循、语音表达、推理与双工交互的合成对话上进行了微调,并通过两阶段混合奖励强化学习改善时序与回复连贯性。SteerDuplex is introduced, a Moshi-based full-duplex speech model fine-tuned on natural conversations and synthetic dialogues targeting instruction following, vocal delivery, reasoning, and duplex interaction, and two-stage reinforcement learning with hybrid rewards to improve timing and response continuity.
提出 CARE(Corrective Atomic Robotic Execution),一种通过从执行中遇到的失败进行学习以提升恢复能力的框架,并引入 Failure State Recovery Benchmark (FSR-Bench),用于评估在局部偏差与结构异常下从中间失败态恢复的表现。This work proposes CARE (Corrective Atomic Robotic Execution), a framework that improves recovery by learning from failures encountered during execution, and introduces the Failure State Recovery Benchmark (FSR-Bench), which evaluates recovery from intermediate failure states under local deviations and structural anomalies.
提出 Grounded Action Models (GAMs),一种基于 3D grounding 构建的机器人基础模型新范式,可自主运行,并作为底层控制器由高层规划器通过多种输入模态进行控制,支持长时序与依赖记忆的操作。Grounded Action Models (GAMs), a new paradigm of robot foundation models built with 3D grounding, are proposed, which can be run autonomously and serve as a low-level controller that a high-level planner controls using its various input modalities, allowing for long-horizon and memory-dependent manipulation.
证明了对通用语言模型适配教育任务(涵盖问题求解、课程理解与教学辅助)的精选、能力均衡的监督数据的价值。The value of curated, capability-balanced supervision for adapting general language models to educational tasks spanning problem solving, curriculum understanding, and instructional support is demonstrated.
提出 Document Retrieval-Aware Chunking (D-RAC),将 Web Retrieval-Aware Chunking (W-RAC) 框架扩展至任意文档格式,保留 W-RAC 的成本、确定性与可观测性优势,同时将每种可渲染格式解锁为一类输入。Document Retrieval-Aware Chunking (D-RAC), an extension of the Web Retrieval-Aware Chunking (W-RAC) framework to arbitrary document formats, is presented, preserving W-RAC's cost, determinism, and observability benefits while unlocking every renderable format as a first-class input.
提出 SVEET 框架,仅需在预训练的双向视频扩散模型上训练,即可支持高质量自回归流式视频编辑,并提出解耦训练方案,显式强制视频可控性与模型因果性优化方向之间的正交性。This paper proposes SVEET, a framework that requires merely training on a pretrained bidirectional video diffusion model but supports high-quality streaming video editing in an auto-regressive fashion and proposes a decoupled training scheme that explicitly enforces the orthogonality between the optimization directions of video controllability and model causality.
表明 Kimi Delta Attention (KDA) 可通过单次 delta-rule 变换与其通道门控提供的反射组合实现 2D 旋转,并刻画 CKDA 的表达能力,证明每个正交的对角加秩-1矩阵恰为一个 CKDA 转移矩阵。This work shows that Kimi Delta Attention (KDA) can realize 2D rotations by combining a single delta-rule transformation with a second reflection supplied by its channel-wise gate, and characterize the expressivity of CKDA and prove that every orthogonal diagonal-plus-rank-one matrix is exactly a CKDA transition matrix.
提出 Deep Persona,一种心理学驱动的三层架构,将 persona 组织为可观察表达、潜在信念与核心动机驱动的层级结构,用于构建高可信度的角色扮演 Agent。Deep Persona is introduced, a psychologically grounded, three-layered architecture that organizes personas into hierarchical levels of observable expression, latent beliefs, and core motivational drives, for constructing highly convincing role-playing agents.
本文提出 KV Policy (KVP),一种仅基于 key 与 value 向量、在预计算生成轨迹上训练的轻量级 per-head RL Agent 框架,证明学习预测未来 token 效用是自适应 KV cache 管理中强大且可扩展的范式。KV Policy (KVP), a framework of lightweight per-head RL agents trained on pre-computed generation traces using only key and value vectors, is introduced, demonstrating that learning to predict future token utility is a powerful and scalable paradigm for adaptive KV cache management.
提出 Calibrated Clipping,一种动态方法,通过匹配下界裁剪分位数并相应重平衡上界,将 FP8 裁剪边界与高精度 BF16 分布对齐,消除熵激增并恢复与 BF16 基线可比的性能。Calibrated Clipping is proposed, a dynamic method that aligns the FP8 clipping bounds with high-precision BF16 distributions by matching the lower-bound clipping quantile and rebalancing the upper bound accordingly, which eliminates entropy surges and restores performance comparable to the BF16 baseline.
提出 UltraTex,一种用于高分辨率多视图扩散 3D 纹理生成的高效端到端框架,并引入 Background Token Dropping(在 DiT 主干前移除背景 token)与 Block-Sparse Attention(降低前景序列上的注意力计算)。This paper presents UltraTex, an efficient end-to-end framework for high-resolution multi-view diffusion-based 3D texturing and introduces Background Token Dropping, which removes background tokens before the DiT backbone, and Block-Sparse Attention, which reduces attention computation over the retained foreground sequence.
内部识别探针(Probe of Internal Recognition, PIR)可区分"不愿回答"与"无法回答"的模型,支持装傻审计与遗忘验证,并从选择题扩展至自由生成。Probe of Internal Recognition (PIR) separates a model that will not answer from one that cannot, which supports sandbagging audits and unlearning verification, and extends from multiple-choice questions to free-form generation.
结果表明,将部分建模负担从网络参数转移至结构化输入表示,可在更低的数据、内存与算力需求下支持紧凑的多任务视觉系统。The results suggest that shifting part of the modeling burden from network parameters to a structured input representation can support compact multi-task vision systems with reduced data, memory, and compute requirements.
提出 SkillSpec,一种 Hoare 风格框架,将技能正确性建模为规约推理问题,并将异构技能仓库转化为统一的图表示,对齐描述、指令与代码构件。This work proposes SkillSpec, a Hoare-style framework that formulates skill correctness as a specification reasoning problem, and transforms a heterogeneous skill repository into a unified graph representation that aligns descriptions, instructions and code artifacts.
提出 StableVQ,重新审视每个模块的合适学习目标,解决各模块独立训练以承担各自角色时产生的问题;其基于共享投影 codebook 构建,轻量且不引入可学习参数。StableVQ is proposed, which revisits the proper learning objective of each module and resolves the problems that arise when each is trained to fulfill its own role independently, and Built on top of shared-projection codebooks, is lightweight and introduces no learnable parameters.
提出 HyperQ,在冻结的掩码扩散语言模型上添加 token 条件化的量子残差分支,支持 token 条件化电路发射,作为一种可处理的量子增强语言建模架构方法。HyperQ is introduced, which adds token-conditioned quantum residual branches to a frozen masked-diffusion language model, and support token-conditioned circuit emission as a tractable architectural approach to quantum-augmented language modelling.
Lean Pool 是一个形式化数学仓库,由 AI agent 进行生长、维护和优化。Lean Pool is a repository of formalized mathematics. It is grown, maintained and optimized by AI agents.
本文提出了 SEGMENT-SNAP,通过 part-handle coupling 融合几何与语义证据,在 Articulate3D Challenge 中获得第一名。This work presents SEGMENT-SNAP, which combines geometric and semantic evidence through part-handle coupling and achieved first place in the Articulate3D Challenge.
本文提出一种面向 few-shot 图像分类的超轻量级空间-关系架构,结合固定的 Gabor 边缘能量引导与窗口化的、内容自适应的 patch locator。实验表明,该架构中虽然存在 pairwise relational computation,但它并非性能的主要驱动因素;真正起决定作用的是 content-adaptive patch locator。An ultra-lightweight spatial-relational architecture for few-shot image classification that combines fixed Gabor edge-energy guidance with a windowed, content-adaptive patch locator is presented, showing that the architecture's pairwise relational computation, while present, is not the primary driver of its performance; the content-adaptive patch locator is.
本文揭示 Agent 数量是多 Agent 组织扩展通用智能边界的新 scaling 维度,为硬延迟约束或时间预算下的复杂任务提供了实用方案。The number of agents is revealed as a new scaling dimension for multi-agent organizations to expand the frontier of general intelligence, offering a practical solution for complex tasks under hard latency constraints or time budgets.
LLM-as-a-judge 支持跨任务评估,但在大规模场景下推理成本与置信度可靠性成为关键问题。本文研究仅做判断的 judge 能否提供经济高效的首轮判别,并识别何时需要更强的评估。在与十六种生成式与奖励模型 judge 的对比中(采用盲法人类裁定作为参照),我们发现 jev-as-a-judge 在普通偏好与有证据支撑的事实性任务上,与作为最强对照的 SOTA LLM judge 仅相差 3 个百分点,成本仅为后者的 0.36%。在需要核查LLM-as-a-judge enables evaluation across diverse tasks, but inference cost and confidence reliability become critical at scale. We study whether a decision-only judge can provide an economical first pass and identify when stronger evaluation is needed. Comparing jev-as-a-judge with sixteen generative and reward-model judges, with blinded human adjudication, we find it within three percentage points of a state-of-the-art LLM judge, our strongest comparator, on ordinary preference and evidence-grounded factuality at 0.36% of the comparator's fee. Larger gaps arise when judgments require checking
一种近期且有效的方法:通过一个小型 low-rank adapter,将大型预训练图像编辑模型适配到图像恢复任务,并以源自退化图像本身的 instruction 替代 text prompt;在同等条件下该方法优于文本条件。A recent and effective recipe adapts a large pretrained image-editing model to restoration using a small low-rank adapter with a text prompt to an instruction derived from the degraded image itself that outperforms text conditioning under a matched comparison.
本文展示了在一对架构匹配的 Qwen3.5 4B-to-9B 兄弟模型之间实现有用的持久 hybrid-state transfer,是首次在大小不同的混合语言模型之间完成持久循环推理状态的跨模型交接,且无需 target prefix replay。This work demonstrates useful persistent hybrid-state transfer across one architecture-matched Qwen3.5 4B-to-9B sibling pair, the first demonstrated cross-model handoff of persistent recurrent inference state between differently sized hybrid language models without target prefix replay.
TritonForge 是一个面向自动化 Triton kernel 优化的 profiling 引导框架,融合 kernel 分析、运行时 profiling 与迭代式代码转换以简化优化流程,并为自动化 GPU 性能优化领域的未来研究奠定基础。TritonForge, a profiling-guided framework for automated Triton kernel optimization that integrates kernel analysis, runtime profiling, and iterative code transformation to streamline the optimization process and provides a foundation for future research in automated GPU performance optimization.
Just-in-Time Memory (JitMem) 一致优于无 memory 的 Agent 以及启发式与学习式写入 memory 方法,相较最强基线在成功率上分别提升了 16.2、16.3 和 3.9 个绝对百分点。Just-in-Time Memory (JitMem) consistently outperforms no-memory agents as well as heuristic and learned write-time memory methods, improving over the strongest baseline by 16.2, 16.3, and 3.9 absolute success-rate points, respectively.
论文表明,在该高维空间中使用 flow matching 的标准速度预测要求模型拟合低维信号流形之外的正交噪声方向,导致优化效率低下,由此提出改用清晰数据参数化($\boldsymbol{x}_{0}$-prediction),使学习聚焦于底层信号流形。It is shown that using the standard velocity prediction in flow matching in this high-dimensional space requires the model to fit orthogonal noise directions outside the low-dimensional signal manifold, making optimization inefficient, and motivated using the clean data parameterization ($\boldsymbol{x}_{0}$-prediction) instead, which focuses learning on the underlying signal manifold.
介绍 RecCAR(Reciprocal Cross-modal Attention Regularization),一种 KL 正则化器,利用已确立的视频到模态对应关系作为固定参考,将较弱的模态到视频对应关系向其对齐。RecCAR, standing for Reciprocal Cross-modal Attention Regularization, is introduced, a KL regularizer that uses the well-established video-to-modality correspondence as a fixed reference and aligns the weaker modality-to-video correspondence toward it.
论文证实,AI 辅导在 GRE 学习收益上与专家人类辅导具有统计等效性(p = .015),且在 GRE 的七个领域中,有五个领域的最佳 AI 导师平均超越了人类导师。It is established that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains (p = .015), and in five of the seven GRE domains, the best performing AI tutor surpassed the human tutor, on average.
PackLab 是一个用于开发、训练与评估闭环机器人装箱 MLLM 的综合框架,在不同物体集合与容器配置下均优于传统装箱启发式方法、经典强化学习方法以及通用 MLLM,展现了 MLLM 在长时任务机器人装箱中的潜力。PackLab is a comprehensive framework for developing, training, and evaluating MLLMs for closed-loop robotic bin packing that outperforms conventional packing heuristics, traditional reinforcement learning methods, and general-purpose MLLMs across object sets and container configurations, highlighting the potential of MLLMs for long-horizon robotic packing.
介绍 Spatial-Interactor,一个通过交互训练 VLM 建模物理世界状态转换的框架,并将学习过程组织为三级课程:L1 被动世界状态转换、L2 主动自我状态转换、L3 长时交互轨迹。Spatial-Interactor is introduced, a framework that trains VLMs to model physical-world state transitions through interaction, organizing this learning process into a three-level curriculum covering L1 passive world-state transitions, L2 active self-state transitions, and L3 long-horizon interaction trajectories.
本文探讨容器化、微服务、动态调度等 Cloud Native 技术如何从根本上提升 LLM 推理服务,并展示 Cloud Native 系统在高需求场景下实现更高效资源分配、降低延迟与提升吞吐的能力。This article explores how Cloud Native technologies, such as containerization, microservices, and dynamic scheduling, can fundamentally improve LLM inference serving and demonstrates how a Cloud Native system enables more efficient resource allocation, reduces latency, and enhances throughput in high-demand scenarios.