为每个可查询的条件分布推导出一个尾部影响度量,聚合其不确定性对每一次 Bellman 复用中 CVaR 的影响,并刻画了尾部最优与均值最优分配重合的精确网格(exact-grid)机制。A tail influence for each queryable conditional law that aggregates how its uncertainty affects CVaR across every Bellman reuse is derived, and an exact-grid regime in which tail- and mean-optimal allocations coincide is characterized.
论文
1094 张论文卡片 · 方法 · OA 绿色
提出 WM-VLM,为预训练 VLM 配备一个轻量级世界模型分支以生成中间视觉状态,表明内部世界模型为 VLM 在语言与视觉双重空间中的推理提供了一条有前景的路径。WM-VLM is introduced, which equips a pretrained VLM with a lightweight world model branch for generating intermediate visual states and suggests that internal world models offer a promising path toward VLMs that reason in both language and visual space.
提出 Nexus,一个将每次调用视为持久任务的云-边平台,结果表明局部性、操作作用域权限与持久结果标识如何支持具有工作负载相关成本的云-边 Agent 服务。Nexus, a cloud-edge platform that treats each invocation as a persistent task, is presented and results demonstrate how locality, operation-scoped authority, and persistent result identity support cloud-edge agent services with workload-dependent costs.
提出 LaMem-VLA,一种以潜在记忆为核心框架的方法,将历史经验重建为潜在记忆 token,并直接与 VLA 推理交织,使记忆能够在有界上下文下直接参与 VLA 推理并引导动作生成。LaMem-VLA is introduced, a latent-memory-native framework that reconstructs historical experience into latent memory tokens and directly interweaves them with VLA reasoning, and enables memory to directly participate in VLA reasoning and guide action generation under a bounded context.
提出 MemAdapter,一种自适应整合检索记忆以支持客观、可靠推理的新框架,证明 MemAdapter 在多种场景下持续提升记忆可靠性。This work proposes MemAdapter, a novel framework that adaptively integrates retrieved memories to support objective and reliable reasoning and demonstrates that MemAdapter consistently improves memory reliability across diverse scenarios.
实验表明,相较于编码 Agent 直接生成游戏世界以及现有基线方法,Code2Games 一致性地提升了生成游戏世界的视觉质量、交互保真度,以及经引擎适配后游戏成品的质量。Experiments demonstrate that, compared with direct gaming-world generation by coding agents and existing baseline methods, Code2Games consistently improves the visual quality and interactive fidelity of generated gaming worlds, as well as the quality of the resulting games after engine adaptation.
Perturbot 和 GroundingFscore 提供了一套训练与评估框架,用于将 VLA 决策与捷径先验解耦,同时保持对任务相关证据的响应性。Perturbot and GroundingFscore provide a training-and-evaluation framework for disentangling VLA decisions from shortcut priors while preserving responsiveness to task-relevant evidence.
提出 SourceLearn,结合两种互补的学习机制以捕获源的可复用理解,包括其知识的结构、解读与应用方式,并以一个捕获源可复用理解的持久 source model 来表示该能力。This work proposes SourceLearn, which combines two complementary learning mechanisms that captures reusable understanding of the source, including how its knowledge is structured, interpreted, and applied, and represents this competence with a persistent source model that captures reusable understanding of the source.
提出一种"失败即关闭"的临时可见性协议:在截止时间前接纳新内容,隐藏验证未及时提交的项目,并支持按谱系范围的遏制,同时量化新鲜度与可用性之间的权衡。A fail-closed provisional-visibility protocol is presented that admits new content under a deadline, hides items whose verification has not committed in time, and supports lineage-scoped containment, and quantifies the freshness and availability trade-offs.
提出基于 Agentic AI 的零信任架构(Agentic-ZTA),通过协调的多 Agent 决策流水线将 NIST SP 800-207 ZTA 架构的控制循环落地,并证明使用 AI agent 执行零信任的可行性。This paper presents Agentic AI based zero trust architecture (Agentic-ZTA) that operationalizes the NIST SP 800-207 ZTA architecture control loop through coordinated multi- agent decision pipeline and demonstrates the feasibility of enforcing zero trust using AI agents.
提出 Behavior-Preserving KV Cache Compression,一种无需训练的框架,通过估计候选条目被移除后压缩缓存所诱导的 logits,并利用驱逐前的前向统计量评估其与完整缓存下一 token 分布之间的 KL 散度来为候选驱逐打分。This work proposes Behavior-Preserving KV Cache Compression, a training-free framework that scores candidate evictions by estimating the compressed-cache logits induced by their removal and evaluating the resulting KL to the full-cache next-token distribution using pre-eviction forward statistics.
DynaKRAG 是一种统一的证据-动作框架,通过学习共享的状态条件策略来协调上述操作,证明统一的状态条件证据控制能在保证强回答质量的同时,实现高效检索与紧凑的答案生成上下文。DynaKRAG, a unified evidence-action framework that learns a shared state-conditioned policy for coordinating these operations, demonstrates that unified, state-conditioned evidence control supports strong answer quality, efficient retrieval, and compact answer-generation contexts.
结果表明,基于查询的循环记忆组合能够提升长上下文建模能力,并在超出训练上下文范围之后依然有效,同时每个 chunk 仅使用紧凑的仿射摘要。Results demonstrate that query-dependent composition of recurrent memory improves long-context modeling and remains effective beyond the training context while using compact per-chunk affine summaries.
提出基于流的自潜在推理(Flow-based Latent Reasoning, FLaRe),一份简洁方案,涵盖潜空间编码内容及其塑形方式、流训练位置、答案读取方式,以及最终在模型自身已验证思考上进行训练的阶段。Flow-based Latent Reasoning (FLaRe), a simple recipe covering what the latent space encodes and how to shape it, where to train the flow, how to read out the answer, and a final stage of training on the model's own verified thoughts, is presented.
表格基础模型通过以标注训练样本为条件进行上下文学习 (ICL) 来预测。与分离训练与推理的传统模型不同,这些模型必须在每次前向传播中处理所有训练样本,导致每次预测成本高昂。限制训练样本数量可降低成本,但会显著降低性能。我们不丢弃上下文,而是提出激活对齐 (activation alignment),利用完整上下文来教会模型在仅看到子集时如何行为。这是通过训练一个……Tabular foundation models perform in-context learning (ICL) by conditioning predictions on labeled training examples provided as context. Unlike traditional models that separate training from inference, these models must process all training examples in every forward pass, making each prediction expensive. Restricting the number of training examples reduces this cost but substantially degrades performance. Instead of discarding context, we propose activation alignment, a method that leverages the full context to teach a model how to behave when seeing only a subset. This is achieved by trainin
提出 LLM-as-Jev,一种保留架构的框架,直接从带括号数字标识符的下一 token 概率中提取校准决策,并同时给出无需训练的推理方案与一种微调目标——通过树分解的列表式损失优化候选选择,并使用 KL 散度惩罚将辅助预测锚定到基模型。This work presents LLM-as-Jev, an architecture-preserving framework that extracts calibrated decisions directly from next-token probabilities over bracketed numeric identifiers, and provides both a training-free inference recipe and a fine-tuning objective that optimizes candidate selection via a tree-factorized listwise loss while anchoring auxiliary predictions to the base model using KL divergence penalties.
将 Clalit Health Services 涵盖超过 540 万患者的专家组中所蕴含的证据提炼为一个已发布的评分工具,支持隐私保护的生物标志物假设生成,并将 Agent 的"提出—评分—精化"循环锚定于真实世界数据,同时不暴露任何患者数据。This work distilled the evidence held in the Clalit Health Services panel of over 5.4 million patients into a released scoring tool, which supports privacy-preserving biomarker hypothesis generation and grounds an agent's propose-score-refine loop in real-world data without exposing any patient data.
一项预注册研究,基于留出的 LoCoMo 对话与 LongMemEval 数据,发现 Jev 以 LLM reranker 三分之一的延迟实现了同等的选取准确率,并且优于多轮 graph traversal 调用。A pre-registered study on held-out LoCoMo conversations and LongMemEval finds Jev selects as accurately as an LLM reranker at a third of the latency, and more accurately than a multi-call graph traversal.
论文介绍了 MiniCorp,一个用于研究 Agent 如何协作运营公司并规模化生成企业数据的办公模拟器,并对其端到端保真度相对于真实市场实证研究所报告的模式进行了评估。MiniCorp, an office simulator for studying how agents can collectively run a company while generating enterprise data at scale, is introduced and end-to-end fidelity against patterns reported in empirical studies of real markets is evaluated.
提出一种新颖的进化系统,专注于利用 RAG 增强的 LLM 自动生成空域隐写的代码级代价函数;该方法在隐写安全性上一致优于现有自动设计方法,同时提高了平均代码执行率并降低了搜索成本。A novel evolutionary system focused on exploiting Retrieval-Augmented Generation enhanced LLMs for the automatic code-level generation of spatial steganography cost functions, which consistently achieves higher steganographic security than existing automatically designed methods and increases the average code execution rate while reducing the search cost.
本文提出 DiffGate,一种将 GRPO 与选择性、有界教师指导相结合的结果门控目标,在四种模型–领域设置下均提升了 pass@8,表明在该评估协议下解的覆盖度得到改善。DiffGate is introduced, an outcome-gated objective that combines GRPO with selective, bounded teacher guidance, and improves pass@8 across all four model--domain settings, demonstrating improved solution coverage under the evaluation protocol.
自回归视频扩散支持流式生成与交互控制,但其 KV cache 随生成历史持续增长。现有压缩策略要么采用固定窗口丢弃历史,要么依据局部注意力与相似度信号选择 token,均无法直接衡量当前块是否贡献了超出已保留上下文的信息。本文提出 DeCoPrune,一种将 cache 压缩视为去噪一致性问题的免训练方法。经验上发现,去噪难度可作为 token 价值的有用代理……Autoregressive video diffusion supports streaming generation and interactive control, but its KV cache grows continuously with the generated history. Existing compression strategies either discard history using fixed windows or select tokens through local attention and similarity signals, which do not directly measure whether the current chunk contributes information beyond the retained context. We introduce DeCoPrune, a training-free method that treats cache compression as a denoising-consistency problem. We find empirically that denoising difficulty provides a useful proxy for a token's valu
分布匹配蒸馏(DMD)通过分别估计的目标分数与学生分数之差训练少步学生模型,因而必须维护一个拟合学生动态分布的辅助扩散模型,带来额外的显存与计算开销。本文提出 DMAD——分布匹配对抗蒸馏,将分布匹配重构为分类任务,直接学习所需的 log-density 比。共享主干上的两个判别头分别区分真实数据与教师样本是否来自学生模型,并对其 logits 施加线性损失以训练……Distribution Matching Distillation (DMD) trains a few-step student from the difference between separately estimated target and student scores, so it must keep an auxiliary diffusion model fitted to the student's evolving distribution at extra memory and computation cost. We introduce DMAD, Distribution Matching as Adversarial Distillation, which recasts distribution matching as classification and learns the required log-density ratios directly. Two discriminator heads on a shared backbone distinguish real data and teacher samples from the student's, and linear losses on their logits train the
将音乐转录为人类可读的乐谱需要对节奏、和声、旋律与曲式的整体理解。两大障碍限制了这一目标的实现:标注录音稀缺,以及精确的局部预测仍可能产生不一致的音乐序列。本文提出 SheetSage2,一个统一的音乐转录框架,结合合成数据、任务特定的结构化解码与自回归蒸馏。自动标注的 MIDI 渲染为音频后,可为各类音乐理解任务提供可扩展的监督。任务特定的结构化解码器整合互补的音乐……Transcribing music into a human-readable score requires a coherent understanding of rhythm, harmony, melody, and form. Two obstacles limit this goal: annotated recordings are scarce, and accurate local predictions can still produce inconsistent musical sequences. We present SheetSage2, a unified music transcription framework that combines synthetic data, task-specific structured decoding, and autoregressive distillation. Automatically annotated MIDI, rendered into audio, provides scalable supervision across music understanding tasks. Task-specific structured decoders integrate complementary mu
稀疏视图新视角合成是三维内容创作中的核心问题,但基于扩散的方法受迭代去噪限制,多视图生成在推理时开销高昂。本文提出 NAMVIS,一种无扩散框架,将多视图图像合成重新表述为几何条件下的下一尺度自回归过程。NAMVIS 不通过反复去噪生成目标视图,而是通过少量由粗到细的尺度步预测离散视觉 token,并在同一尺度内以及跨目标视图间并行采样所有 token。为锚定此过程……Sparse-view novel view synthesis is a central problem in 3D content creation, but diffusion-based approaches remain limited by iterative denoising, making multi-view generation expensive at inference time. We introduce NAMVIS, a diffusion-free framework that reformulates multi-view image synthesis as geometry-conditioned next-scale autoregression. Instead of generating target views through repeated denoising, NAMVIS predicts discrete visual tokens through a small number of coarse-to-fine scale steps, while sampling all tokens within each scale and across target views in parallel. To anchor thi
WildCity 是一个由自动驾驶车队在复杂城市环境中采集的真实多模态数据集,旨在推动城市级渲染的进展,并更广泛地推动 AI 在空间感知、记忆与推理方面达到与人类认知相当规模的能力。WildCity, a real-world multimodal dataset collected by autonomous fleets traversing complex urban environments, aims to catalyze progress not only in city-scale rendering, but more broadly in the pursuit of AI that can perceive, remember, and reason across space at a scale comparable to human cognition.
提出了用于复杂图像创建与编辑的大规模多模态工具调用数据集 CanvasCraft,以及通过多轮交互学习编排异构视觉工具的工具增强多模态 Agent CanvasAgent。CanvasCraft, a large-scale multimodal tool-use dataset for complex image creation and editing, and CanvasAgent, a tool-augmented multimodal agent that learns to orchestrate heterogeneous visual tools through multi-turn interaction are introduced.
提出 KVpop,通过直接监督保留或丢弃决策来学习固定预算的 KV 淘汰策略,并引入基于延迟记忆的打分器,在已学习的淘汰方法中独有地延迟固定步数的打分,以利用近期上下文。KVpop is introduced, which learns a fixed-budget KV eviction policy by directly supervising the keep-or-drop decision, and introduces a delayed memory-based scorer that, uniquely among learned eviction methods, defers scoring for a fixed number of steps to exploit near-future context.
提出 MV-Forcing 框架,通过在顺序生成视角之间引入 4D 几何桥梁,在单一扩散模型中组合时间与视角自回归,并弥合时间与视角序列自回归中训练-推理的曝光偏差差距。MV-Forcing, a framework that composes temporal and view-wise autoregression within a single diffusion model by introducing a 4D geometric bridge between sequentially generated views, and closes the train-inference exposure bias gap for both temporal and view-sequential autoregression.
该工作通过轻量级 Gated MLP 将固定的 Gemini Embedding 2 视觉-语言表征映射到目标空间,使用 KL 散度与同方差不确定性加权进行训练,并展示了 AI Wizards 参加 EXIST 2026 多模态迷因性别歧视识别任务的提交方案。This work maps fixed Gemini Embedding 2 vision-language representations through a lightweight Gated MLP trained with KL divergence and homoscedastic uncertainty weighting, and presents the AI Wizards submission to EXIST 2026 for multimodal sexism identification in memes.
该工作提出了一个端到端的档案处理与检索框架,将大语言模型(LLM)集成到档案流程中,并证明将 LLM 与成熟的文档处理与检索流程相结合,可将数字图书馆从静态存储库提升为可交互、可语义检索的档案系统。This work presents an end-to-end archival processing and retrieval framework that integrates large language models (LLMs) into the archival pipeline and demonstrates that integrating LLMs with established document processing and retrieval pipelines can elevate digital libraries from static repositories to interactive, semantically searchable archival systems.
提出流形约束假设(Manifold Constraint Hypothesis, MCH):若自然表征集中于一个结构化的低维流形上,则干预应被约束在该流形内,从而在干预过程中更好地保留表征中编码的其他信息。The Manifold Constraint Hypothesis (MCH) is proposed: if natural representations concentrate on a structured, lower-dimensional manifold, then interventions should be constrained to that manifold and better preserve other information encoded in the representation during interventions.
CONFLUX 是一个面向胸部 CT 的潜在扩散模型:由 3D 变分自编码器压缩每个体数据,整流流 Transformer 在潜在空间中生成,并以分类器从生成体中恢复所请求病征的可靠性作为奖励。CONFLUX, a latent diffusion model for chest computed tomography (CT): a 3D variational autoencoder compresses each volume, and a rectified-flow transformer generates in the latent space, that rewards how reliably a classifier recovers the requested findings from each generated volume.
本工作推出 Gemma 4——Gemma 模型系列中新一代开源权重、原生多模态的语言模型,在 STEM、多模态与长上下文基准上实现性能跃升,在人类评分任务上可与更大的前沿开源模型相媲美。This work introduces Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family that establishes a leap in performance across STEM, multimodal, and long-context benchmarks, and rivals larger, frontier open models in human-rated tasks.
本文提出 OrbitQuant,一种数据无关的权重量化器,通过在归一化旋转基空间中进行量化以绕过范围估计,将图像扩散 Transformer 的 PTQ 推进到 W2A4 并保持可用生成质量。OrbitQuant is presented, a data-agnostic weight-activation quantizer that bypasses range estimation by quantizing in a normalized, rotated basis and pushes PTQ of image diffusion transformers to W2A4 with usable generation quality.
本文提出 SOLAR,一种学习增强框架,从 regret 累积中推导修改时机,并基于隐式检索反馈的贝叶斯在线学习进行内容选择,实现与缓存大小和时域无关的常数竞争比。SOLAR is proposed, a learning-augmented framework that derives modification timing from regret accumulation and content selection from Bayesian online learning over implicit retrieval feedback and achieves a constant competitive ratio, independent of cache size and horizon.