受 Apache Spark 等分布式计算框架启发,DISCO 将长上下文切分到一组专职 Worker LLM 上,进行并行、局部的 grounding,建立了鲁棒长上下文推理的高效范式。Inspired by distributed computing frameworks like Apache Spark, DISCO partitions long context across a fleet of Worker LLMs dedicated exclusively to parallel, localized grounding, establishing a highly efficient paradigm for robust long-context inference.
论文
1640 张论文卡片 · OA 绿色
提出 GAGAR,一个面向代码 agent RL 的质量感知信用再分配框架;结果表明,将基于测试的验证与成组的 agentic 评分相结合,可提升代码 agent RL 的质量与稳定性。GAGAR, a framework for quality-aware credit redistribution in code agent RL, is introduced and results support combining test-based verification with groupwise agentic grading to improve the quality and stability of code agent RL.
介绍 Relic,它将反复出现的协作失败转化为组织拥有的、可执行的协议,使协作经验成为超越其创建者仍持续有用的组织级持久状态。The introduction of Relic, which turns recurring collaboration failures into organization-owned, executable protocols, and how collaboration experience can become persistent organizational state that remains useful beyond the members who created it are shown.
本工作提出 LMCACHE,首个也是目前最高效的开源 KV 缓存方案,可将现代 LLM 引擎生成的 KV 缓存从 GPU 显存中提取并存储,并跨引擎和查询共享。This work presents LMCACHE, the first and so far the most efficient open-source KV caching solution, which extracts and stores KV caches generated by modern LLM engines out of the GPU memory and shares them across engines and queries.
提出 G²PTQ,一个具有广义梯度补偿的统一 PTQ 框架,在全局监督的分块优化目标下融合一阶与二阶信息,能更好地对齐全精度模型,性能优于 SOTA 基线。G$^2$PTQ is presented, a unified PTQ framework with Generalized Gradient Compensation that integrates both first- and second-order information under a globally supervised, block-wise optimization objective and enables better alignment with the full-precision model, outperforming state-of-the-art baselines.
本报告建议不同利益相关方可采取多种措施,提升关于 AI 系统及其相关开发流程声明的可验证性,重点是为 AI 系统的安全性、安保、公平性与隐私保护提供证据。This report suggests various steps that different stakeholders can take to improve the verifiability of claims made about AI systems and their associated development processes, with a focus on providing evidence about the safety, security, fairness, and privacy protection of AI systems.
注意力往往集中在上下文中一小部分 token 上,但每个 query 关注的关键子集各不相同。为利用这种动态结构,我们提出 SANTA++,一种免训练的随机注意力方法,通过代表性 key 进行内存高效的选择,无需扫描整个 KV cache。缓存的 key 被组织成若干 team,query 对每个 team 中的代表性 key 打分以决定采样哪些 team。我们在采样得到的 team 内计算精确的注意力分数,并通过其采样概率的倒数对各 team 的贡献进行重新加权。Attention often concentrates on a small subset of tokens in the context, but which subset matters changes from one query to the next. To exploit this changing structure, we introduce SANTA++, a training-free stochastic attention method that uses representative keys for memory-efficient selection without scanning the entire key-value (KV) cache. Cached keys are organized into teams, and the query scores one representative from each team to decide which teams to sample. We compute exact attention scores within the sampled teams and reweight each team's contribution by the inverse of its inclusio
提出反事实树上的注意力搜索(ASCT),将训练时的多步搜索转化为局部动作信用,将反事实评估与策略学习相连接,同时部署时只需使用 actor。Attentive Search over Counterfactual Trees (ASCT) is introduced, a framework that turns training-time multi-step search into local action credit and connects counterfactual evaluation to policy learning while deploying the actor alone.
OLIVE 通过从持续演化的 student 生成前缀,在离线蒸馏进入平台期后继续提升,同时更好地保持了 student 的通用能力与可塑性。By regenerating prefixes from the evolving student, OLIVE continues improving after offline distillation plateaus while better preserving the general capabilities and plasticity of the student.
提出解耦信用自蒸馏,理论上将信用方向与幅度解耦为两个可靠信号,并据此校准特权教师监督,从而完成策略优化的 step-to-token 信用分配。Decoupled Credit Self-Distillation is introduced, which theoretically decouples credit direction and magnitude into two reliable signals and uses them to calibrate privileged teacher supervision, enabling step-to-token credit assignment for policy optimization.
结果表明,推理路径多样性可作为筛选 SFT 数据的实用准则,能更好地为 RL 准备模型,并据此提出一种轻量级、基于规则的指纹方法用于筛选。These results identify reasoning-route diversity as a practical criterion for selecting SFT data that better prepares models for RL, and propose a lightweight, rule-based fingerprint to select for it.
NVV-SuperBench 与人工听力评测的结果显示,NVAlign 在标签跟随准确率上优于 SFT 与 Flow-GRPO 基线,证明直接对奖励梯度进行优化可提升连续自回归流匹配 TTS 中的非语言控制能力。Results from NVV-SuperBench and human listening evaluations show that NVAlign improves tag-following accuracy over SFT and Flow-GRPO baselines, demonstrating that direct reward-gradient optimization can improve non-verbal control in continuous autoregressive flow-matching TTS.
本文提出一个新颖的初步框架,可定量评估输出格式不同(如列表与消息)的检索系统的 IR 准确度,为客观评估基于关键词与基于语义的对话式检索方法奠定坚实基础。This work introduces a novel, preliminary framework to quantitatively assess IR accuracy of search systems that produce different output formats, such as lists and messages, and provides a strong foundation for objectively assessing keyword-based and semantic chat-based search methods.
提出 Context Language Models,能原生管理自身上下文,并可通过经标准 skill 优化循环演化的自然语言指令进行引导,在上下文管理任务上将未见数据的准确率最高提升 35.9 分,同时降低计算开销。Context Language Models are introduced, language models that natively manage their own context and can be steered with natural-language instructions evolved through a standard skill-optimization loop, improving held-out accuracy by up to 35.9 points on a context-management task while reducing compute.
构建了 RoboGenesis,一个基于仿真的工作流与数据引擎,可从原子技能组合配置好的实验工作流,对 rollout 进行验证与过滤,并跨支持的机器人配置导出结构化演示数据。RoboGenesis is built, a simulation-based workflow and data engine that composes configured laboratory workflows from atomic skills, validates and filters rollouts, and exports structured demonstrations across supported robot profiles.
提出 CorpusMap,一个以语料中重复出现的实体为核心的导航层;这些实体可从文档自身识别,并能在不同来源间把单个文档与众多其他文档相连,表明实体可作为大型文档集合导航的有效锚点。CorpusMap is introduced, a navigation layer that organizes the corpus around its recurring entities, which are identifiable from the documents themselves and can link a single document to many others across sources, suggesting that entities serve as effective anchors for navigating large document collections.
在所测试的每一对基座与指令微调模型中,指令微调都强化了模型对保留标记的偏好,且该差距在该通道上持续存在。In every base and instruction-tuned pair the authors test, instruction tuning strengthens the model's preference for reserved markers, and the gap persists on that channel.
本文推导了对应的反向传播规则(包括校准项的导数),并在模型、数据、优化器均一致的预训练实验中对不同选择进行了比较。This work derives the corresponding backward rules, including calibration derivatives, and compares these choices in pretraining experiments matched on model, data, and optimizer, and compares these choices in pretraining experiments matched on model, data, and optimizer.
结果表明,harness 层面的能力组合是构建广泛、可演进 Omni 系统的实用路径,且无需改变宿主 agent 的推理核心。The results establish harness-level capability composition as a practical route to broad, evolvable Omni systems without changing the host agent's reasoning core.
提出 Omni-Decision,一个基于证据账本规划的 omni-modal agent:用显式的证据账本替代不断膨胀的对话历史,记录仍缺失的证据、已确认的内容以及记录间的冲突。Omni-Decision is presented, an omni-modal agent built on evidence-ledger planning: it replaces the growing dialogue history with an explicit evidence ledger that records what evidence is still missing, what has been confirmed, and where records conflict.
KUPAS MASTER 是一个围绕九层认知语料构建的经验工程平台,将异构的工作记录与从业者访谈转化为 agent 可用的可追溯、可复用经验语料,提供了从个体隐性经验到组织知识与 agent 能力的可行路径。KUPAS MASTER, an experience engineering platform built around nine-layer cognitive corpus construction, turns heterogeneous work records and practitioner interviews into traceable, reusable experience corpora for agents, and provides a practical path from individual tacit experience to organizational knowledge and agent capabilities.
APM-Bench 将真实场景的流式交互重构为多会话生命轨迹,并揭示了一个清晰的效用—延迟—存储权衡:现有方法仍难以同时实现可靠的长程记忆、低开销以及跨会话有效的主动协助。APM-Bench is introduced, which reformulates real-world streaming interaction as multi-session life trajectories and reveals a clear utility--latency--storage trade-off: existing methods still struggle to simultaneously achieve reliable long-term recall, low overhead, and effective proactive assistance across sessions.
一种随机少步生成器,其 FID 在采样预算从 4 增加到 50 NFE 时不会下降,且相同配方可迁移到 text-to-image 生成。A stochastic few-step generator whose FID does not degrade as the sampling budget grows from 4 to 50 NFE, and the same recipe transfers to text-to-image generation.
提出 Budgeted ATTA:测试批次中仅有一小部分可获得标签,且监督时机是 active test-time adaptation 中一个关键但尚未充分探索的方面。Budgeted ATTA is introduced in which labels are available for only a fraction of test batches, and the timing of supervision is a key, yet underexplored, aspect of active test-time adaptation.
本文提出一种异步 LLM 框架,允许用户(或 Agent 自身)定义具有重叠 memory 状态的推理协程,并展示 Qwen 3.x 模型无需任务专属训练即可在流式视频理解、电子游戏和监控中实现异步运行。This work develops an asynchronous LLM framework that lets users (or the agents themselves) define inference coroutines with overlapping memory states and showcases that Qwen 3.x models are capable of asynchronous operation for streaming video understanding, videogames, and monitoring, without task-specific training.
本文提出 AutoRef,在保持两侧模型冻结的前提下自动优化 harness:一个 coding Agent 迭代重写 harness 代码,并发现 AutoRef-Harness,可改进开源权重模型 FLUX。This work proposes AutoRef, which optimizes the harness automatically while keeping both models frozen: a coding agent iteratively rewrites the harness code, and discovers AutoRef-Harness, which improves the open-weight FLUX.
在包括多视角空间推理和碰撞预测在内的六项多样化视觉推理任务中,ReImaGin 始终优于纯文本推理和专用 vision-tool 基线,提升幅度高达 25%,证明了灵活、可生成的视觉推理的优势。Across six diverse visual reasoning tasks including multi-view spatial reasoning and collision prediction, ReImaGin consistently outperforms both text-only reasoning and specialist vision-tool baselines, with gains of up to 25\%, demonstrating the advantage of flexible, generative visual reasoning.
本文提出 LibraryDesignBench,这是一个两阶段基准,Agent 根据一份定义所需能力和潜在用例但不规定具体设计的规范,实现一个功能完备的 library。LibraryDesignBench is introduced, a two-phase benchmark in which an agent implements a full-featured library from a specification that defines required capabilities and potential use cases without prescribing the design.
通道级 history-state 注入结合 timestep 条件化构成了一种低成本且更有效的设计,能够在更长 unrolling 下更好地保留知识,同时提升在不同推理预算下的鲁棒性。It is shown that channel-wise history-state injection combined with timestep conditioning offers a low-cost and more effective design, better preserving knowledge under extended unrolling while improving robustness across inference budgets.
本文揭示使用 chunked KV-cache 压缩的模型中存在一种系统性不对称:相同信息在一个阶段易于检索,在另一阶段却难以获取,暴露出平均 benchmark 分数可能掩盖的周期性薄弱点。A systematic asymmetry in models using chunked KV-cache compression is uncovered: the same information can be easy to retrieve at one phase and difficult at another, revealing periodic weak spots that average benchmark scores can conceal.
本文提出 FRAC,一种源自分数阶动力学的选择性 SSM 架构,用幂律长记忆替代指数衰减,在长上下文性能上持续优于 SOTA SSM 基线,同时在短上下文上保持竞争力。FRAC is introduced, a selective SSM architecture derived from fractional dynamics that replaces this exponential decay with power-law long memory and consistently improves long-context performance over state-of-the-art SSM baselines while staying competitive on short-context.
研究发现,早期 OPD 训练动态均呈现一种规律的 *useful-transfer* 区间,其中留出准确率(即 *gold score*, $G$)随 $d=\sqrt{\mathrm{KL}(\pi_\theta \Vert \pi_{\mathrm{ref}})}$(学生初始化在 token 级反向 KL 散度的平方根)近似线性上升。It is found that early OPD training dynamics uniformly exhibit a regular *useful-transfer* regime, in which held-out accuracy (the *gold score*, $G$) rises approximately linearly in $d=\sqrt{\mathrm{KL}(\pi_\theta \Vert \pi_{\mathrm{ref}})}$, the square root of token-level reverse KL divergence from the student initialization.
本工作提出 12 个具体指标,从一致性、鲁棒性、可预测性和安全性四个关键维度分解 Agent 可靠性,可与传统评估互补,并提供用于分析 Agent 表现、退化与失败方式的工具。This work proposes twelve concrete metrics that decompose agent reliability along four key dimensions: consistency, robustness, predictability, and safety, which complement traditional evaluations while offering tools for reasoning about how agents perform, degrade, and fail.
本文提出 UniEvo-VL,一种面向多模态模型的自演化框架,可在 test-time compute 阶段从这种建设性的自纠错反馈中学习,在无外部监督或指导的情况下提升用户使用多模态模型的体验。UniEvo-VL, a self-evolving framework for multimodal models to learn from this constructive self-correction feedback during test-time compute, is introduced to enhance the user experience when using multimodal models without external supervision or guidance.
本文提出 Mid-Harness,在执行前对候选动作进行采样和验证,同时保持 generator 和 harness 不变,并将 action scaling 识别为终端 Agent 中 test-time compute scaling 的一个有前景的目标。Mid-Harness is introduced, which samples and verifies candidate actions before forwarding one for execution, while keeping the generator and harness unchanged, and identifies action scaling as a promising target for test-time compute scaling in terminal agents.
本文提出多样本验证(MSV),对同一模型在有源信息和无源信息下各查询三次,以决定任务接纳并替换不可靠的伪标签,这部分减少了虚假一致性,但仍残留大量 co-cheating。This work introduces multi-sample verification (MSV), which queries the same model three times with the source and three times without it to decide task admission and replace unreliable pseudo-labels, which partially reduces false agreement but leaves substantial residual co-cheating.