FLEET 将每次生成表示为穿越状态的稀疏轨迹(状态的熵超过预设阈值),并基于这些轨迹推断逐 token 的效用分数以调整 logits,是一种将 memory 机制融入生成过程的新方法。FLEET represents each generation as a sparse trajectory through states whose entropy exceeds a predefined threshold and uses these trajectories to infer per-token utility scores that adjust the logits, a novel method that integrates a memory mechanism into the generation process.
论文
1173 张论文卡片 · 方法
本文提出一种方法,通过逐层剔除后 WER(Word Error Rate)的变化对编码器层进行排序,并给出剪枝后的模型——即一个层数更少的更浅编码器。This work presents an approach that ranks encoder layers by the leave-one-layer-out change in Word Error Rate (WER), and presents the pruned model, which is simply a more shallow encoder with fewer layers.
本文提出一个基于原理的、无需训练的框架,依次优化跨层权重配对与共享字典分解,并识别结构兼容的投影、学习一种能更好保留各层独立校准几何的共享表征。This work introduces a principled, training-free framework that sequentially optimizes cross-layer weight pairings and shared-dictionary factorizations, and identifies structurally compatible projections and learns a shared representation that better preserves each layer's distinct calibration geometry.
结果支持将生成式记忆视为对直接检索的选择性修正,并强调何时、以何种方式、以何种强度进行路由干预是核心挑战。The results support generated memory as a selective correction to direct retrieval and highlight routing when, which, and how strongly to intervene as the central challenge.
本文提出 X-Planner,一个面向具身推理的规划前端,同时解决监督与表征问题,并描述了规划文本质量与下游执行情况。This work presents X-Planner, a planning front-end that addresses both the supervision and representation of embodied reasoning, and describes planning-text quality and downstream execution.
结果发现 Astra 表现出 token 高效的有向推理,能更早选择正确轨迹,在内部解决基础步骤,仅外化关键推理,为前沿模型推理提供了超越基准分数的行为视角。It is found that Astra exhibits token-efficient directed reasoning, selecting a correct trajectory earlier, while resolving elementary steps internally and externalizing only crucial reasoning, which provides a behavioral lens on frontier-model reasoning beyond benchmark scores.
Knowledge Pull Requests 相比从来源重写或从零重新生成,能整合更多信息并更好地保留已有内容,同时每生成一个 token 增加的信息量最多。Knowledge Pull Requests integrate more information and better preserve existing content than rewriting from sources or regenerating from scratch, while adding the most information per token generated.
本文首次对 vLLM 启动延迟进行了详细的性能表征,并构建了一个轻量级分析模型,能够针对给定硬件配置准确预测 vLLM 的启动延迟,为大规模推理环境中的资源规划提供了可操作的指导。This paper presents the first detailed performance characterization of vLLM startup latency and develops a lightweight analytical model that accurately predicts vLLM's startup latency for a given hardware configuration, providing actionable guidance for resource planning in large-scale inference environments.
本文提出 PUBG Ally,一个面向 PUBG: BATTLEGROUNDS 的具身代理,能推理、自主行动并作为语音队友与玩家并肩作战,将代理式工具使用与实时游戏控制相结合。PUBG Ally is introduced, an embodied agent for PUBG: BATTLEGROUNDS that can reason, act autonomously, and play alongside players as a voice-enabled teammate that combines agentic tool use with real-time game control.
本文提出 World Action Agent,一个多代理框架,通过它 VLM 可借助基础工具操控机器人,所有决策均在可视化动作工作空间内完成,在相同骨干下优于端到端 VLA、code-as-policy 代理以及一个可视化框架基线。World Action Agent is presented, a multi-agent harness through which VLMs pilot robots with basic tools, making every decision within a visual action workspace, outperforming end-to-end VLAs, code-as-policy agents, and a visual-harness baseline with the same backbone.
主要发现是:多样且高质量的 SFT 奠定坚实的能力下限,难度过滤将 RL 提示维持在有效的学习区间内,而奖励可靠性为阶段排序提供了实用原则。The main findings are that diverse, high-quality SFT establishes a strong capability floor and difficulty filtering keeps RL prompts within a productive learning range, and reward reliability provides a practical principle for ordering stages.
本文提出 ViRDM,一种无需教师与评论家网络的视频后训练方案,将三网络蒸馏转化为仅生成器的后训练,在降低 GPU 显存与训练时间的同时提升视频质量。ViRDM, a teacher- and critic-free video post-training recipe that turns three-network distillation into generator-only post-training, reducing GPU memory use and training time while improving video quality, is introduced.
本文提出 Neural Spectral Capacity (NSC),一种以每个权重矩阵的奇异值谱为依据的闭式标量,在七类 Transformer 与 CNN 系列上的排序效果优于 #Params、#FLOPs 及代表性免训练代理指标。This work proposes Neural Spectral Capacity (NSC), a closed-form scalar grounded in the singular-value spectrum of each weight matrix, which outperforms #Params, #FLOPs, and representative training-free proxies in ranking across seven Transformer and CNN families.
结果显示 SAE 以分布式、类别依赖的形式定位形态句法信息,而非通过原子化的语法特征。The results show that SAEs localise morpho-syntactic information in a distributed and category-dependent form rather than through atomic grammatical features.
本文将数据安全形式化为 provenance monomials 上的聚合谓词,并提出 Passant——一个无需物化 provenance 即可强制执行 DFC 策略的可移植查询重写层。This paper formalizes data safety as aggregate predicates over provenance monomials and presents Passant, a portable query rewriting layer that enforces DFC policies without materializing provenance.
本文发现编码代理在通用 TAMP 上表现出惊人的有效性:在平均成功率上,三种代理配置均优于人工设计的规划器、一次性生成以及基于 LLM 的通用规划基线。This work finds that coding agents are surprisingly effective at generalized TAMP: all three agent configurations outperform hand-engineered planners, one-shot generation, and an LLM-based generalized planning baseline in mean success.
本文提出 DeltaWAM,通过 dense-anchor、sparse-delta 与 action 三条流联合预测视觉 delta 与动作,并设计了三种在表征与计算共享上有所不同的架构;同时开发了 Streaming Delta Memory (SDM),使用紧凑的观测 delta 更新缓存的 anchor 上下文,从而减少繁重的 video-expert 处理。This work proposes DeltaWAM, which jointly predicts visual deltas and actions using dense-anchor, sparse-delta, and action streams, with three architectures that differ in representation and computation sharing, and develops Streaming Delta Memory (SDM), which updates cached anchor context with compact observed deltas, reducing heavy video-expert processing.
本文提出 AV-GRPO,一种以模态为锚点的在线 diffusion RL 框架,以及 5DAV,一种解耦的、难度可调节的训练数据集;在 LoRA 与全量 fine-tuning 条件下,其在生成质量、语义对齐和跨模态同步性上均优于 LTX-2.3。This work proposes AV-GRPO, a modality-anchored online diffusion RL framework, and 5DAV, a decoupled, difficulty-controllable training dataset that outperforms LTX-2.3 in generation quality, semantic alignment and cross-modal synchronization under LoRA and full fine-tuning.
近年来,大语言模型(LLMs)解决高级数学问题的能力日益增强,包括许多悬而未决数十年的难题。这为以空前规模扩展数学知识打开了大门。然而,尽管 LLM 能够猜想并证明越来越多的定理,这些新数学知识是否有趣或有用仍属未知。我们将定理的内在有趣度定义为其证明长度与陈述长度之比。证明该指标与下载量的外在度量高度相关。Recently, Large Language Models (LLMs) have been increasingly able to solve advanced mathematical problems, including many that have been open for decades. This opens the door to expansion of mathematical knowledge at unprecedented scale. Yet, while LLMs may be able to conjecture and prove more and more theorems, it remains open whether this new mathematical knowledge is interesting or useful. We define intrinsic interestingness of a theorem as the ratio between the length of its proof and the length of its statement. We show that this correlates strongly with an extrinsic measure of the downs
本文提出一种可调深度网络架构,借助适配器残差模块可在线切换至不同视觉域;同时引入 Visual Decathlon Challenge 基准,用于评估表征同时捕获十个差异显著视觉域的能力,并衡量其跨域均匀识别的能力。This paper develops a tunable deep network architecture that, by means of adapter residual modules, can be steered on the fly to diverse visual domains and introduces the Visual Decathlon Challenge, a benchmark that evaluates the ability of representations to capture simultaneously ten very differentVisual domains and measures their ability to recognize well uniformly.
本文提出 PISA,一种采用金字塔式 Top-K 选择策略的 block-sparse attention 机制,并为训练与推理开发了硬件感知的 Triton kernel,将层级路由与 LogSumExp 评分融合,无需显式构造 query-key score 矩阵。PISA is proposed, a block-sparse attention mechanism that employs a pyramid Top-$K selection strategy, and develops hardware-aware Triton kernels for both training and inference, fusing hierarchical routing and LogSumExp scoring without materializing the query-key score matrix.
本文提出 ZooWork-ShopRanker,一族对齐到裁判标注购物偏好的电商 reranker,以及 ShopRank-Bench,一个包含约 10,000 条私有流量偏好对的低污染 benchmark,按承诺该标注的裁判家族数量分档呈现,覆盖多种文本格式。ZooWork-ShopRanker, a family of e-commerce rerankers aligned to judge-labeled shopping preference, and ShopRank-Bench, a contamination-limited benchmark of ~10,000 private-traffic preference pairs in both text formats, tiered by how many judge families committed to each label.
TrackEverything 是首个在 40 GB GPU 显存内即可在超过 1000 帧的视频中跟踪所有可见点的 3D tracker;本文还提出 3D WAFT,用场景点云内的高效特征采样取代了显存开销巨大的 4D correlation volumes。TrackEverything is the first 3D tracker capable of tracking all visible points across videos exceeding 1000 frames within 40 GB of GPU memory, and 3D WAFT is proposed, replacing memory-prohibitive 4D correlation volumes with efficient feature sampling in the scene cloud.
本文提出一条全自动的数据生成流程,可在无需任何人工标注的情况下生成图像对之间的逐点感知距离标签,并证明 diffusion trajectory 与人类视觉系统高度一致。This paper proposes a fully automated data generation pipeline that generates pointwise perceptual distance labels between image pairs without any human annotation, and demonstrates that the diffusion trajectory aligns well with the human visual system.
本文提出 PsPLUG,一个轻量级插件式模块,在刻画目标风格后学习用户专属的残差,能更好地保留用户偏好,并对个性化与风格遵循之间的平衡提供更精确的控制。PsPLUG is proposed, a lightweight plug-in that learns a user-specific residual after accounting for the requested style, and better preserves user preferences while providing precise control over the balance between personalization and style adherence.
本文提出一种改进的 Stable Diffusion 3 架构,采用裁剪后的文本流与逐 patch 的归一化策略,使其能在 LiDAR 数据上稳定训练,并实现从自然图像到高程图的迁移;研究表明多模态条件输入可提升高程精度。A modified Stable Diffusion 3 architecture with a pruned text stream and a patch-wise normalization strategy is introduced, enabling stable training on LiDAR data and transfer from natural images to elevation maps, demonstrating that multimodal conditioning improves elevation accuracy.
本文提出首个通过 continuous depth batching (CDB) 实现 depth-adaptive looped LM 的高效方法:在 loop 步骤之间重组 batch,动态调度架构中的 looped 与非 looped 部分,管理 looped KV-caching,并提前预测将退出 loop 的 token,以便异步准备 batch。This work introduces the first efficient method for depth-adaptive looped LMs via continuous depth batching (CDB), which forms new batches between loop steps, and dynamically schedules looped and non-looped parts of the architecture, manages looped KV-caching, and predicts which tokens will exit the loop in advance so it can prepare batches asynchronously.
Continnum,一种通过为 KV cache 保留引入 TTL 机制来优化多轮 Agent 工作负载任务完成时间的服务系统,与程序级 FCFS 结合时可保持多轮连续性,并降低 Agent 工作流的延迟。Continnum, a serving system to optimize job completion time for multi-turn agent workloads by introducing time-to-live mechanism for KV cache retention, and when combined with program-level first-come-first-serve preserves multi-turn continuity, and reduces delay for agentic workflows.
使用一种分层旋转位置编码(hRoPE),将段落、句子和 token 索引表示为独立通道,保持 token 序列固定,对段落坐标 $p_1$ 进行干预,并使用 token 距离精确估计器测量跨段落注意力。A hierarchical rotary positional encoding (hRoPE) that represents paragraph, sentence, and token indices as separate channels, hold the token sequence fixed, intervene on the paragraph coordinate $p_1$, and measure cross-paragraph attention with a token-distance-exact estimator is used.
BoundInk 是一种以书写者为条件的框架,将字符间的边界视为显式生成单元,既能保留书写者特有的字形外观,又能在完整文本行内改善连接性与字距。BoundInk is introduced, a writer-conditioned framework that treats inter-character boundaries as explicit generation units and preserves writer-specific glyph appearance while improving connectivity and spacing across complete text lines.
本文提出 VLA-Precision,一个面向真实世界的高效在线 RL 框架,包含 Asymmetric Co-Bootstrapping (ACoB) 算法与 ACoB-Stream 架构;其中 ACoB-Stream 以不变状态解耦与按需流式传输为设计原则,构建了经验-策略的闭环架构,可实现最高 10.9% 的吞吐与计算效率提升。VLA-Precision, an efficient real-world online RL framework featuring the Asymmetric Co-Bootstrapping (ACoB) algorithm and the ACoB-Stream architecture, and ACoB-Stream, a closed-loop experience--policy architecture that establishes invariant-state decoupling and on-demand streaming as design principles, deliver up to 10.9% improvements in throughput and computational efficiency.
QReason 是一种解耦框架,将面向 query 的推理与针对窗口的 passage 相关性评估分离,显著减少冗余推理,在取得与 reasoning-based reranker 相当乃至更优的排序性能的同时,超越了现有的 query rewriting 模型。QReason is a decoupled framework that separates query-focused reasoning from window-specific passage relevance assessment, and significantly reduces redundant reasoning, achieves ranking performance comparable to or better than strong reasoning-based rerankers, and outperforms existing query rewriting models.
Intent2Tc 是一个由语言模型驱动的闭环框架,将业务级流量整形意图转换为声明式子意图,进而生成经过验证的可执行 Linux traffic control (tc) 配置,并展示了该框架的实际适用性。Intent2Tc is presented, a closed-loop language-model-driven framework that translates business-level traffic-shaping intents into declarative sub-intents and subsequently into validated, executable Linux traffic control (tc) configurations and demonstrates the practical applicability of the proposed framework.
本文提出 softmax reparameterization,一种训练后方法,在量化前搜索功能等价的输出头,并展示了在总体 logit 误差增大的情况下保真度仍可提升。This work introduces softmax reparameterization, a post-training method that searches over functionally equivalent output heads before quantization and shows how fidelity can improve despite greater total logit error.
LightMIS 通过 Scale-Aligned Projection 块将五级编码器的输出对齐至统一分辨率,仅聚合一次,并利用结合 Adaptive Kernel Fusion 与所提 Progressive Receptive Fusion 模块的 Adaptive Fusion Cascade 精炼融合表示。LightMIS aligns the outputs of a five-level encoder to a common resolution using Scale-Aligned Projection blocks, aggregates them once, and refines the fused representation with an Adaptive Fusion Cascade, which combines Adaptive Kernel Fusion with the proposed Progressive Receptive Fusion module.
SarseX 模型无关、无需训练,并与 Prefix Cache 兼容,可为多轮对话、检索增强生成 (RAG) 和 Agent 工作流等常见在线服务场景提供统一支持。SarseX is model-agnostic, training-free, and compatible with Prefix Cache, and it provides unified support for common online serving scenarios including multi-round chat, retrieval-augmented generation (RAG), and agent workflows.