最小执行契约可在生成程序中引发主动的结构适应,将计算从无约束分配中移开,在执行前显著改善观察到的资源时间分布,建立了 substrate-aware Agent 规划的受控概念验证。A minimal execution contract induces proactive structural adaptation in generated programs, shifting computation away from unconstrained allocations and substantially improving observed resource-time profiles before execution, establishing a controlled proof of concept for substrate-aware agent planning.
论文
1753 张论文卡片
研究结果强调了方向特定的迁移测试、严格的 embedding space 隔离,以及在 memory migrations 中为 memory repair 保留源历史的必要性。Findings highlight the necessity of direction-specific migration testing, strict embedding space isolation, and the retention of source histories for memory repair in memory migrations.
该工作提出了 OR-Clarify,一个用于预表述澄清的 benchmark,并提出了 Interactive Optimization (InterOPT),这是一个两阶段框架,能识别未解决的、影响表述的关键缺口,并据此引导系统决定是提出下一个问题还是停止提问。This work introduces OR-Clarify, a benchmark for pre-formulation clarification and proposes Interactive Optimization (InterOPT), a two-stage framework that identifies unresolved formulation-critical gaps and uses them to guide whether to ask the next question or to stop.
Tangram 是一种 serving 框架,将先前系统动态处理的内容静态解析,可作为现有 non-uniform 压缩方法的即插即用底座,在匹配其精度的同时,端到端吞吐量较 full-KV 基线最高提升 2.6×。Tangram is a serving framework that statically resolves what prior systems handle dynamically, and serves as a drop-in substrate for existing non-uniform compression methods, matching their accuracy while improving end-to-end throughput by up to $2.6\times over the full-KV baseline.
涵盖数学推理、多领域 STEM、代码生成以及多轮 Agent 任务的实验表明,RISE 在所有设置下均优于仅使用 RLVR 的训练以及 on-policy self-distillation。Experiments spanning mathematical reasoning, multi-domain STEM, code generation, and multi-turn agentic tasks show that RISE outperforms RLVR-only training and on-policy self-distillation across all settings.
该工作提出了 MaxKernel,一个多 Agent 系统,实现了 TPU kernel 开发的 three distinct paradigms:Human-in-the-Loop (HITL) Agent,用于协作式分步设计;Autonomous (Auto) Agent,执行全自动、由指标和 trace 驱动的优化循环;以及 Graph-Based Autonomous Search,将 Auto Agent 扩展以对设计空间进行全局探索。This work presents MaxKernel, a multi-agent system that implements three distinct paradigms for TPU kernel development: a Human-in-the-Loop (HITL) agent for collaborative, step-by-step design; an Autonomous (Auto) agent that executes a fully automated, metric/trace-driven optimization loop; and a Graph-Based Autonomous Search that scales the Auto agent for global exploration of the design space.
该工作通过向参考解中注入受控的 AST 级 corruption,从 400 个 BigCodeBench 问题构建了一个评估框架,为每个修复任务赋予已知的最小 patch,并将 edit fidelity 定位为 code-repair 质量的一个独立维度,表明其可被度量与学习。This work constructs an evaluation framework from 400 BigCodeBench problems by injecting controlled AST-level corruptions into reference solutions, giving each repair task a known minimal patch, and positions edit fidelity as a distinct axis of code-repair quality and shows that it can be measured and learned.
循环状态写回被确立为低精度循环动态的关键决定因素,state-storage interface 被识别为量化循环推理的核心设计考量。R recurrent-state write-back is established as a key determinant of low-precision recurrent dynamics and the state-storage interface is identified as a central design consideration for quantized recurrent inference.
提出了 Motion-Omni,一个端到端框架,其中的 spoken dialogue model 原生输出显式的 facial expression 以及手部、上半身和下半身运动,这些输出直接由生成语音的 hidden states 生成。Motion-Omni is presented, an end-to-end framework in which a spoken dialogue model natively outputs explicit facial expression together with hand, upper-body and lower-body motion, generated directly from the hidden states that produce the speech.
该工作发现 reasoning operations 在 held-out 表征中是可分的,且 separability 在中间层达到峰值,并验证该结构无法由词汇或位置混淆因素解释。This work finds that reasoning operations are separable in held-out representations, with separability peaking in middle layers, and verify that this structure is not explained by lexical or positional confounds.
提出了 τ^τ-bench(读作 hyper-tau-bench),一个将 Agent 构建作为任务的 benchmark,将协作式 Agent 构建工作转化为面向 coding agent 的可度量目标。The $\tau^\tau$-bench (pronounced hyper-tau-bench), a benchmark that makes agent construction the task, is introduced to turn the work of cooperative agent building into a measurable target for coding agents.
研究揭示沿音视频边的路由是双向的:音频可影响视频生成,视频也可影响音频生成;模型参数中编码的偏差是泄漏的主要来源之一。It is revealed that routing along the audio-video edge is bidirectional: audio can influence video generation, while video can influence audio generation, and biases encoded in the model's parameters and emerges as a major contributor to leakage.
提出随机反思记忆提升(SRMA),仅在 grounding 后的评估风险严格下降时才接受候选记忆,为随机评估提供置信度门控,并为分段平稳环境提供重新锚定保证。Stochastic Reflective Memory Ascent (SRMA), which accepts a candidate memory only after a grounded evaluation risk strictly decreases, is introduced and provides confidence gating for stochastic evaluation and re-anchoring guarantees for piecewise-stationary environments.
ShallowStream 是一个利用 MLLM 浅层同时进行帧编码与检索索引构建的新框架,性能与当前最强流式方法相当,同时将单帧 prefill 延迟与 10 秒端到端延迟分别降低至多 52.1 倍和 11.9 倍。ShallowStream is a novel framework that leverages the shallow layers of an MLLM to simultaneously perform frame encoding and retrieval index building and achieves performance on par with the strongest existing streaming methods, while reducing per-frame prefill latency and 10-second end-to-end latency by up to 52.1x and 11.9x, respectively.
本文认为,AI agent——即以大语言模型作为主要推理引擎、动态生成与丢弃代码作为工具性资源的系统——的出现构成了对"软件"本身的根本性重构,而非渐进式的工具改进。This paper argues that the emergence of AI agents -- systems where large language models serve as the primary reasoning engine, dynamically generating and discarding code as an instrumental resource -- constitutes a fundamental restructuring of what software is, not an incremental tool improvement.
本文提出 Enoki,一个面向多级幻觉检测的开放信息抽取框架,支持基于 LLM、基于编码器与基于规则的三类抽取模式,通过统一接口平衡准确率与推理成本。This work proposes Enoki, an Open Information Extraction framework for multi-level hallucination detection that supports LLM-based, encoder-based, and rule-based extraction regimes, balancing accuracy and inference cost through a common interface.
HoloWorld 是首个在统一连贯的 3D 城市世界中同时支持室内与室外生成的框架,构建于持续更新的跨尺度世界上下文之上。HoloWorld is the first framework to unify indoor and outdoor generation within a coherent 3D urban world, built on a continuously updated cross-scale world context.
提出 UniMate,一个统一的 foundation model,可从绑定骨骼的 3D 资产与文本提示合成任意骨骼的关节运动,无需测试时优化或针对每个骨骼的重新训练,在质量、泛化性与效率上均超越 SOTA 基线。UniMate is presented, a unified foundation model that synthesizes articulated motion for arbitrary skeletons from a rigged 3D asset and a text prompt, with no test-time optimization or per-skeleton retraining, and outperforms state-of-the-art baselines in quality, generalization, and efficiency.
提出 KoNA,一个评估 VLM 选择性不遵从的基准,覆盖五类情形:False Premise、Visual Inaccessibility、Universal Unknown、Task Feasibility 与 Safety。实验结果显示,微调后的模型能够区分可回答部分与需要不遵从的部分,并以符合任务要求的方式作答。KoNA is introduced, a benchmark for evaluating selective non-compliance in VLMs across five categories: False Premise, Visual Inaccessibility, Universal Unknown, Task Feasibility, and Safety, and results suggest that the fine-tuned models can distinguish between answerable components and those requiring non-compliance and respond in a task-appropriate manner.
本文将安全微调数据集中的回复拆分为两个独立部分:模板化的拒答声明与解释拒答的理由,并表明拒答声明会诱导模型依赖表层线索,从而妨碍对有害与良性查询的准确区分。This paper decomposes a response in the safety-tuning dataset into two distinct components: a boilerplate refusal statement and a rationale explaining the refusal, and shows that refusal statements impede accurate discrimination between harmful and benign queries by inducing reliance on superficial cues.
视频编辑涵盖多种编辑范式,但在单一统一框架中同时实现高质量的指令引导与主体引导编辑仍具挑战性。我们提出 EditVid,一个免训练框架,结合用于局部一致性的稀疏因果记忆、用于长程身份保持的基于对应关系的后注意力 token 注入,以及用于编辑局部性的软潜在融合。同一框架支持指令引导和参考引导的编辑,包括风格迁移、属性修改、对象插入、部分级编辑和主体替换。在Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guided editing within a single unified framework remains challenging. We introduce EditVid, a training-free framework combining sparse causal memory for local coherence, correspondence-based post-attention token injection for long-range identity preservation, and soft latent blending for edit locality. The same framework supports instruction-guided and reference-guided edits, including style transfer, attribute modification, object insertion, part-level editing, and subject replacement. On
HarvestBench 是首个为避免 side effect 标定价格并将 side effect 命名为生物的 benchmark;在六个模型中有四个的每次作答遭遇 kill rate 对价格变化敏感。HarvestBench is the first benchmark to put a price on avoiding a side effect and name the side effect as a living creature and four out of six models'kill rate per answered encounter were sensitive to price changes.
为推进任务课程聚焦于大规模词分类,本竞赛设置两条互补赛道:Deep 赛道面向同一被试内部的大规模词分类,目标是追求最佳性能;Broad 赛道面向跨被试泛化。Advancing the curriculum of tasks to focus on word classification to focus on word classification at scale, two complementary tracks are presented in this competition: the Deep track targets within-subject word classification at scale, aiming at the best possible performance; the Broad track targets cross-subject generalisation.
提出 FactoSR,一个因子化强化学习框架,显式解读视觉投影所塌缩的维度,并指出强化显式的、因子化的 4D 一致性是将 VLM 演化为稳健、具有世界感知能力的推理器的关键一步。FactoSR is presented, a factorized reinforcement learning framework that explicitly interpret the dimensions collapsed by visual projection, and suggests that reinforcing explicit, factorized 4D consistency is a critical step toward evolving VLMs into robust, world-aware reasoners.
本文提出 Dr. Claw,一个开源工作空间,将现有编码 Agent 可执行文件封装在可控且可审计的人机协同工作流中,而非引入另一个自主 Agent。Dr. Claw is presented, an open-source workspace that wraps existing coding-agent executors in a controllable and auditable human-in-the-loop workflow rather than introducing another autonomous agent.
本文提出 Memanto,一种面向 agentic 人工智能的通用记忆层,挑战了"必须依赖知识图谱复杂度才能实现高保真 agent 记忆"的普遍假设,并取得 SOTA 准确率。Memanto is introduced, a universal memory layer for agentic artificial intelligence that challenges the prevailing assumption that knowledge graph complexity is necessary to achieve high fidelity agent memory and achieves state of the art accuracy scores.
FlowBalance 是一种以验证器为锚定的自改进方法,学习完整响应上的归一化分布,在 Qwen3-4B 和 Qwen3-8B 上相对 FlowRL 提升了平均性能,同时改善了训练速度与稳定性,避免了直接 OPSD 响应长度坍缩,并在受控的 AIME24 诊断中表现出更高的正确策略多样性。FlowBalance, a verifier-grounded self-improvement method that learns a normalized distribution over complete responses, improves average performance over FlowRL on both Qwen3-4B and Qwen3-8B, while also improving training speed and stability, avoiding direct OPSD's response-length collapse, and exhibiting higher correct-strategy diversity in a controlled AIME24 diagnostic.
本文为该设定引入新基准,并基于该基准评估了九种修订方法,包括序贯反思与并行采样变体,使用 gpt-oss-20b/120b、gpt-5.4-mini 以及 qwen3.5-9b/27b/122b 进行测试。A new benchmark for this setting is introduced, and nine revision methods are evaluated, including sequential reflection and parallel sampling variants, using gpt-oss-20b/120b, gpt-5.4-mini, and qwen3.5-9b/27b/122b on the benchmark.
本文提出 Teacher-Gated On-Policy Distillation(教师门控的在线策略蒸馏),其核心原则是在引入密集监督前以 prompt 级别验证教师可靠性;在全部六个单领域设定下优于 Vanilla OPD,并在多领域训练下于两种规模上取得更高的七项基准平均成绩。Teacher-Gated On-Policy Distillation is introduced, built on the principle that teacher reliability should be verified at the prompt level before dense supervision is admitted, and outperforms Vanilla OPD in all six single-domain settings and achieves higher seven-benchmark averages at both scales under multi-domain training.
本文提出 KnowChange,一个知识引导的变化数据合成框架,利用预训练视觉语言模型作为知识源,从变化前场景与期望变化类型推理合理的变化位置与类别转移,在统一框架下灵活合成多样变化类型。KnowChange is introduced, a knowledge-guided change data synthesis framework that leverages pretrained vision-language models as knowledge sources to reason about plausible change locations and class transitions from pre-change scenes and desired change types and enables flexible synthesis of diverse change types within a unified framework.
大语言模型(LLM)的成功很大程度上归功于 next-token prediction(NTP),但其自回归(AR)结构需要缓慢的串行 token 生成。为克服这一瓶颈,我们提出扩散增强 LLM,一类新模型,在使用扩散从该分布中并行采样多个 token 的同时定义 AR 模型分布。我们将这些模型的参数解耦为两组:AR 权重,使用标准 NTP 目标训练;轻量扩散权重,训练用于同时生成多个 token。扩散权重Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressive (AR) structure requires slow, sequential token generation. To overcome this bottleneck, we introduce diffusion-augmented LLMs, a new class of models that defines an AR model distribution while using diffusion to draw multiple tokens in parallel from that distribution. We decouple the parameters of these models into two sets: AR weights, trained using the standard NTP objective, and lightweight diffusion weights, trained to generate multiple tokens simultaneously. The diffusion weight
我们提出 ENEAS,一种用于实例追踪与语义发现的统一且文本可提示的方法。包括 SAM 3 在内的文本可提示分割模型仍存在时间幻觉、空间碎片化与语义误分类问题:目标离开视野时无法报告目标缺失;极端特写下只分割局部纹理而非完整目标;将视觉特征置于本体事实之上,从而把雕像、绘画或反射等视觉相似的物体误分割为目标。We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-promptable segmentation models, including the latest foundation models such as SAM 3, still suffer from temporal hallucinations, spatial fragmentation, and semantic misclassification: they fail to report target absence when an object leaves the field of view, segment local textures instead of the complete object during extreme close-ups, and prioritize visual features over ontological reality, so that visually similar artifacts such as statues, paintings, or reflections are segmented as targe
EmbodiedSkills 是一个统一框架,将每项技能决策视为执行提案:运行时先检查前置条件再执行,执行后再验证结果,为将底层 VLA 策略转化为闭环具身系统提供可训练且可检视的 Agent 层。EmbodiedSkills is a unified framework that treats each skill decision as an execution proposal: the runtime checks its prerequisites before execution and verifies the outcome afterward and provides a trainable and inspectable agent layer for turning low-level VLA policies into closed-loop embodied systems.
本综述将崩溃视为受三处杠杆共同支配的症状:信号施加位置(即 token 加权方式)、教师所见信息(即特权信息的性质)、以及信号变化时机(即教师的动态性与引导衰减)。This review treats collapse as a symptom governed by three levers: where the signal is applied, that is, how tokens are weighted; what the teacher is shown, that is, the nature of the privileged information; and when the signal changes, that is, the teacher's dynamics and the decay of guidance.
本文对一个双节点 split-LLM 训练系统进行系统安全案例研究:其隐私评估通过,却遗留一条未被测试的可观测信道;系统因此并不安全——包括跨训练步骤累积观测在内的五类攻击从未被测量。A systems-security case study of a two-node split-LLM training system whose privacy evaluation passed while leaving an observable channel untested, but the system is not thereby safe: five classes of attack, including those accumulating observations across training steps, were never measured.
结果表明数据组成控制安全性与可用性的权衡,且安全对齐应在预期拒答边界的两侧进行评估。The results show that data composition controls the safety and usability trade-off, and that safety alignment should be evaluated on both sides of the intended refusal boundary.