本文提出 IterSynth,一种角色解耦、基于摘要的范式,在用于识别信息需求的 Planner 与用于将证据整合进演化摘要状态的 Synthesizer 之间交替,作为一种模型无关的提示范式,在前沿闭源模型上相对 ReAct 及类似提示范式取得显著的零样本增益。IterSynth is proposed, a role-decoupled and summary-based paradigm that alternates between a Planner for identifying information needs and a Synthesizer for integrating evidence into an evolving summary state, and serves as a model-agnostic prompting paradigm, delivering substantial zero-shot gains over ReAct and similar prompting paradigms on frontier proprietary models.
论文
72 张论文卡片 · 观点 · OA 绿色
本文给出证据,说明 superposition 是 Transformer 架构的内在属性,而非训练过程中涌现的结果,并证明通过轻量级 fine-tuning 可以在很大程度上恢复线性性。This work provides evidence that superposition is an intrinsic property of the Transformer architecture rather than an emergent consequence of training, and demonstrates that linearity can be substantially restored through lightweight fine-tuning.
本文提出 ExpertAlign,一种无需域标签或独立 routing 模型的框架,可在每个 token 上对全体 teacher 池进行监督路由;研究表明 token 级路由能够利用跨域互补监督,并减少对 prompt 级域指派的单一依赖。This work proposes ExpertAlign, a framework that routes supervision over the full teacher pool at each token, without domain labels or training a separate routing model, and demonstrates token-level routing can exploit cross-domain complementary supervision, and reduce exclusive reliance on prompt-level domain assignment.
Boxoffice 是一个程序化生成评测数据集的工具,用以覆盖具有挑战性的 KV cache 复用模式;研究表明现有数据集并未展现出充分评估此类技术所需的复用动态特性。Boxoffice, a tool that programmatically generates evaluation datasets that exercise challenging KV cache reuse patterns, is introduced and it is shown that existing datasets do not exhibit the reuse dynamics needed to thoroughly evaluate such techniques.
本报告建议不同利益相关方可采取多种措施,提升关于 AI 系统及其相关开发流程声明的可验证性,重点是为 AI 系统的安全性、安保、公平性与隐私保护提供证据。This report suggests various steps that different stakeholders can take to improve the verifiability of claims made about AI systems and their associated development processes, with a focus on providing evidence about the safety, security, fairness, and privacy protection of AI systems.
本文揭示使用 chunked KV-cache 压缩的模型中存在一种系统性不对称:相同信息在一个阶段易于检索,在另一阶段却难以获取,暴露出平均 benchmark 分数可能掩盖的周期性薄弱点。A systematic asymmetry in models using chunked KV-cache compression is uncovered: the same information can be easy to retrieve at one phase and difficult at another, revealing periodic weak spots that average benchmark scores can conceal.
转移层面的结果表明,state adaptation 选择性而非统一地应用时最为有效,且 adaptation 的价值取决于策略反转的频率与幅度。The transition-level results suggest that state adaptation is most useful when applied selectively rather than uniformly, and that adaptation value depends on both the frequency and magnitude of strategy reversals.
实验表明,屏蔽高熵偏移位置相比朴素的 SD 提升了分布外泛化能力,由此得到的自蒸馏 judges 在所评估的主观子类别上比基于结果监督 RL 训练的 judges 高出 2-9 个百分点,同时在客观子类别上保持竞争力。Experiments show that masking higher-entropy-shift positions improves out-of-distribution generalization over naive SD, and the resulting self-distilled judges outperform judges trained with outcome-supervised RL by 2-9 percentage points on the evaluated subjective subcategories, while remaining competitive on objective ones.
本文提出 JevSpawn,一种将自然语言任务规范连接到有限概率探索的组合策略,并将 JevSpawn 确立为结构化 Agent 推理的一种有前景的方法,在任务性能和导航速度上均有所提升。This work introduces JevSpawn, a compositional policy that connects natural language task specifications to finite probabilistic exploration, and establishes JevSpawn as a promising approach to structured agentic inference, with improved task performance and faster navigation.
WARP:一个直接从已发布权重还原微调模型训练数据混合比例的框架,抽取几何特征并映射至各领域占比,可采用无参数 softmax 读出器,或基于合成混合训练的 MLP 投影器。WARP is introduced, a framework that recovers a fine-tuned model's training mixtures directly from its released weights and extracts geometric features and maps them to domain proportions using either a parameter-free softmax readout or an MLP projector trained on synthetic mixtures.
本文提出 Randomized YaRN,一种通过将基于 YaRN 的位置外推与随机位置编码和长度课程相结合来提升长度泛化能力的训练方法,表明渐进式地将模型暴露于分布外位置分布是实现可泛化长上下文推理的有效方案。Randomized YaRN is proposed, a training method that improves length generalization by combining YaRN-based positional extrapolation with randomized positional encoding and a length curriculum, and suggests that progressively exposing models to OOD positional distributions provides an effective recipe for generalizable long-context reasoning.
本文通过收益–成本视角研究 PLT 循环次数选择:额外循环可精炼表示,但 CLP 也会在每次循环边界引入位置错配,由此解释 PLT 在两次循环时趋于饱和的现象,并为循环次数选择提供诊断依据。This study studies PLT loop-count selection through a gain--cost view: an extra loop may refine representations, but CLP also introduces a positional mismatch at each loop boundary, explaining PLT's saturation at two loops and providing diagnostics for loop-count selection.
在 MS MARCO 与真实生产流量上的结果表明,自适应检索 cardinality 能够在不降低检索质量的前提下提升检索效率。Results on both MS MARCO and real-world production traffic suggest that adaptive retrieval cardinality can improve retrieval efficiency without degrading retrieval quality.
本文在视觉语言导航、具身问答和语言条件操控任务上评估了三种 AAS 变体,覆盖四个具身执行器,结果表明架构级搜索能在具身任务上产生可部署且具有方向性的成功率提升,而其中一个看似得分较高的候选因存在泄漏而被判定为无效。This work evaluates three AAS variants across four embodied executors spanning vision-language navigation, embodied question answering, and language-conditioned manipulation, and shows that architecture-level search can produce deployable and directional success-rate gains on embodied tasks, while one apparent high-scoring candidate is rejected as leak-bearing.
本文论证了稀疏组合监督与动宾学习的不对称性会助长物体驱动的捷径学习,并指出减少捷径诊断可提升组合泛化能力。This work argues that sparse compositional supervision and verb-object learning asymmetry can promote object-driven shortcut learning and reduces shortcut diagnostics and consequently improves compositional generalization.
通过考察包含意识形态话语的 RAG 框架对 LLM 生成答案的影响,发现 RAG 框架倾向于将意识形态话语传递到 LLM 响应中,且采样温度对这种传递的强度有可测量的影响。Examining the influence of the RAG framework, comprising ideological discourses, in LLM-generated answers shows that the RAG framework is prone to transferring ideological discourses into LLM responses, with sampling temperature having a measurable impact on the strength of this transfer.
本文从玩家动作控制、游戏状态动态、状态-观测持久性与实时交互生成四个维度审视交互式游戏世界建模,并针对《Black Myth: Wukong》提出可扩展的数据引擎,采集超过 90 小时的游戏画面作为状态感知型游戏世界建模的资源。This paper examines interactive game world modeling along four dimensions: player action control, game state dynamics, state-observation persistence, and real-time interactive generation, and presents a scalable data engine for Black Myth: Wukong that collects over 90 hours of gameplay as a resource for state-aware game world modeling.
提出 AI Prototyper,一个开源 Figma 插件,通过分解与 RAG 流水线自动完成 GUI 原型设计,并引入人在回路编辑步骤,允许用户在渲染前审查、修改或扩展生成的功能列表。AI Prototyper is presented, an open-source Figma plugin that automates GUI prototyping through a decomposition and retrieval-augmented generation (RAG) pipeline, and introduces a human-in-the-loop editing step that lets users review, modify, or extend the generated feature list before rendering.
提出 PA-HDP(Prompt-Aware Dynamic Hierarchical Differential Privacy)框架,通过 prompt 感知的风险分层动态评估不同查询下的隐私风险,并采用自适应敏感实体替换与基于指数机制的文本选择,在保留语义可用性的同时提供差异化的隐私保护。A Prompt-Aware Dynamic Hierarchical Differential Privacy framework (PA-HDP) is proposed, which performs a prompt-aware risk hierarchy to dynamically assess privacy risks under different queries and applies adaptive sensitive entity replacement and exponential mechanism-based text selection to provide differentiated privacy protection while preserving semantic utility.
该报告主张并例证了一种严谨且基于经验的方法来研究 AI 意识:依据获得最佳支持的神经科学意识理论,详细评估现有 AI 系统。This report argues for, and exemplifies, a rigorous and empirically grounded approach to AI consciousness: assessing existing AI systems in detail, in light of best-supported neuroscientific theories of consciousness.
提出了ShotPlan,一个基于视频扩散基础模型构建的、用于显式多镜头电影级视频生成的框架,显著优于现有的电影级视频生成方法,提供更灵活的镜头管理和更强的跨镜头一致性。ShotPlan is proposed, a framework for explicit multi-shot cinematic video generation built upon a video diffusion foundation model that significantly outperforms existing cinematic video generation methods, offering more flexible shot management and stronger inter-shot consistency.
本文认为,可解释 AI 研究总体上有助于推动 AI/ML 在医疗领域的落地,并特别有助于增强透明性与信任。It is argued that research in explainable-AI would generally help to facilitate the implementation of AI/ML in the medical domain, and specifically help to facilitates transparency and trust.
提出结构化动态模型(SDM),通过未来特征预测,显式地将时间变化的主导来源与残差动态分离开来,而非使用单一纠缠的隐变量或非结构化的、空间密集的转移 token 来表示视频变化。The Structured Dynamics Model (SDM) is proposed, which explicitly separates the dominant source of temporal change from residual dynamics through future-feature prediction, rather than representing video change with a single entangled latent or with unstructured, spatially dense transition tokens.
将 Positive--Direction Matching (PDM)——一种分支感知的 OPD 目标,分别约束正预测方向与 CFG 条件方向——引入 dense-to-sparse 视频控制;由于朴素的 guided matching 对推理 guidance 尺度极为敏感,分支感知监督可实现更鲁棒、更有效的知识迁移。Positive--Direction Matching (PDM), a branch-aware OPD objective that separately constrains the positive prediction and the CFG conditional direction, is introduced to dense-to-sparse video control, where naive guided matching is highly sensitive to inference guidance scales, while branch-aware supervision enables more robust and effective knowledge transfer.
本文提出 ReDesign,一种 agentic 框架,通过跨模态选择与组合专用工具来构建可编辑的层级结构(layer hierarchy),在取得强视觉保真度的同时,于布局、颜色与文本编辑上提供最高的可编辑性。ReDesign is presented, an agentic framework that grows an editable layer hierarchy by selecting and composing specialized tools across modalities, and achieves strong visual fidelity while delivering the highest editability across layout, color, and text edits.
提出一种自适应推理时引导框架,利用VLA Latents上的强化学习,发现推理时引导在成功与失败状态下遵循根本不同的scaling laws:动作多样性在基础VLA可能失败时最为有益,但在成功可能性高时可能不必要地扰动已准确的动作。This work introduces an adaptive inference-time steering framework that leverages Reinforcement Learning on VLA Latents, and discovers that inference-time steering follows fundamentally different scaling laws under success and failure states, revealing that action diversity is most beneficial when the base VLA is likely to fail, but can unnecessarily perturb already-accurate actions when success is likely.
情感对话研究包含两种颇具影响力的策略传统。共情对话优先理解说话者的情绪体验;情感支持对话则选择并排序以满足求助者当前的需求。持续使用引入了更进一步的目标:有效的支持应在整个交互生命周期中维持用户进行情绪调节、应对、自我认同决策以及社会联结的能力。我们提出能力维持型情感对话(CSED)作为一种纵向研究范式,将支持策略与上述目标对齐,并...Emotional dialogue research includes two influential strategy traditions. Empathetic dialogue prioritizes understanding a speaker's emotional experience. Emotional support conversation selects and sequences support for the seeker's current needs. Sustained use introduces a further goal. Effective support should sustain users' capacities for emotion regulation, coping, self-endorsed decisions, and social connection across the interaction lifecycle. We propose capability-sustaining emotional dialogue (CSED) as a longitudinal research paradigm that aligns supportive strategy with this goal and or
发现 ALiBi 的失效模式会显著损害 token 检索,而对标准 decoder 基准影响较小;提出四种训练时缓解策略,在 passkey 检索上获得最一致的提升。It is found that ALiBi's failure mode can substantially impair token retrieval while having only a minor effect on standard decoder benchmarks, and proposes four training-time mitigation strategies that yield the most consistent improvements in passkey retrieval.
EffectLearner 是一个语义推理增强框架,结合基于 VLM 的 Object-Effect Reasoner 与基于 DiT 的 Video Eraser,在 EffectWorld-Eval 和具有挑战性的 EffectWorld-Wild 上均取得明显优势,证明其能在复杂真实场景中实现高质量的视频物体擦除。EffectLearner is proposed, a semantic-reasoning-enhanced framework that combines a VLM-based Object-Effect Reasoner with a DiT-based Video Eraser that achieves clear advantages on both EffectWorld-Eval and the challenging EffectWorld-Wild, demonstrating its ability to deliver high-quality video object removal in complex real-world scenes.
在客观与主观指标上的重建和生成结果匹配或超越前沿开源 tokenizer,包括 Wan-2.2、HunyuanVideo-1.5、FLUX.2、MovieGen、StableAudio 和 MMAudio 的 VAE。It is demonstrated that reconstruction and generation results on objective and subjective metrics matches or surpasses frontier opensource tokenizers, such as VAEs from Wan-2.2, HunyuanVideo-1.5, FLUX.2, MovieGen, StableAudio and MMAudio.
HiSparse 已合入上游 SGLang,并在 H200、B200 与 GH200 平台上针对 DSA、NSA、Quest 三类稀疏注意力家族进行评估:在长上下文负载下峰值生成吞吐提升最高达 4.7×,同时保持可比的 per-token 延迟,并降低高负载下的 time-to-first-token。HiSparse is merged into upstream SGLang and evaluated across three sparse-attention families (DSA, NSA, and Quest) on H200, B200, and GH200 platforms: it improves peak generation throughput by up to 4.7x on long-context workloads while preserving comparable per-token latency and reducing time-to-first-token at high load.
与现有句法语言模型在推理时对众多句法树求边际或在运行时丢弃句法不同,SiPE 以单一句法树为条件,在句法监督与推理成本之间建立了新的 Pareto 前沿。Unlike existing syntactic language models that marginalize over many parses at inference or discard syntax at runtime, SiPE conditions on a single parse, establishing a new Pareto frontier between syntactic supervision and inference cost.
这篇立场论文定义了可解释性,阐述了何时需要(以及何时不需要)可解释性,并提出了一种用于严格评估的分类法,同时指出了迈向更严谨的可解释机器学习科学所面临的开放性问题This position paper defines interpretability and describes when interpretability is needed (and when it is not), and suggests a taxonomy for rigorous evaluation and exposes open questions towards a more rigorous science of interpretable machine learning.
引入了一个框架,通过提供简洁接口来跟踪实时能耗与碳排放、生成标准化的在线附录来简化核算,并为节能的强化学习算法建立排行榜以激励负责任的研究A framework is introduced that makes accounting easier by providing a simple interface for tracking realtime energy consumption and carbon emissions, as well as generating standardized online appendices, and creates a leaderboard for energy efficient reinforcement learning algorithms to incentivize responsible research.
本文提出 Syfer,一种用于多语言多跳问答的 synthesizer-folding 框架,默认推迟翻译而非直接应用翻译,在保持具有竞争力准确性的同时,在性能与计算成本之间取得良好平衡。The method Syfer is introduced, a synthesizer-folding framework for multilingual multi-hop question answering that defers translation rather than applying it by default and attains competitive accuracy while striking a favourable balance between performance and computational cost.
2023 年,一位纽约法官在 Mata v. Avianca 案中制裁了两位律师,因其提交的 brief 包含了由 ChatGPT 生成的虚构引用。此类失误大多能被数据库检索发现;但更棘手的问题在于检测那些指向真实案例、却不支持其所述命题的引用——这一失效模式是现有面向法律场景的 LLM 评测基本忽略的。本文通过对来自两个法律语料库的真实法律引用进行受控扰动(替换引用In 2023, a New York judge sanctioned two attorneys in Mata v. Avianca for filing a brief with hallucinated citations generated by ChatGPT. Such failures are largely caught by database lookups; the harder problem is detecting citations that point to real cases but do not support the propositions for which they are offered -- a failure mode that existing evaluations of LLMs for legal use cases largely overlook. In this paper, we study proposition-level citation support verification through controlled perturbations of real legal citations obtained from two legal corpora, either replacing the cite