Manifold Mixup 在监督学习、对单步对抗攻击的鲁棒性、半监督学习以及留出样本的负对数似然(NLL)上,相较强基线均取得了大幅提升。Manifold Mixup achieves large improvements over strong baselines in supervised learning, robustness to single-step adversarial attacks, semi-supervised learning, and Negative Log-Likelihood on held out samples.
论文
1130 张论文卡片 · 方法 · OA 绿色
本文提出新的计算机视觉基础模型 Florence,通过融入来自 Web 规模图文数据的通用视觉-语言表示,将表征范围从粗粒度(场景)扩展到细粒度、从静态(图像)扩展到动态(视频)、从 RGB 扩展到多种模态(描述、深度等)。This work introduces a new computer vision foundation model, Florence, to expand the representations from coarse (scene) to fine, from static (images) to dynamic (videos), and from RGB to multiple modalities (caption, depth), by incorporating universal visual-language representations from Web-scale image-text data.
本文提出了不变风险最小化(IRM),一种用于在多个训练分布上估计不变相关性的学习范式,并展示了 IRM 学到的不变性如何与数据的因果结构相关,从而实现分布外泛化。This work introduces Invariant Risk Minimization, a learning paradigm to estimate invariant correlations across multiple training distributions and shows how the invariances learned by IRM relate to the causal structures governing the data and enable out-of-distribution generalization.
研究表明 BigBird 是序列函数的通用逼近器,且具备图灵完备性,从而保留了二次全注意力模型的这些性质。It is shown that BigBird is a universal approximator of sequence functions and is Turing complete, thereby preserving these properties of the quadratic, full attention model.
本文提出 Enfold,将构建未来的计算转移到由当前视觉上下文和语言指令预测出的表征中,并把世界生成器重塑为预测控制表征的来源,前提是其内部结构可被折叠(enfold)到当下。This work presents Enfold, which transfers this computation that constructs a future into a representation predicted from the current visual context and language instruction, and recast a world generator as a source of predictive control representations if its internal structure can be enfolded into the present.
训练一个紧凑分类器,对被省略的句子按其是否为理解保留文本所必需进行排序,并在推理时无需支撑标注地回插排名靠前的候选,同时优化相关性与指代完整性。A compact classifier is trained to rank omitted sentences by whether they are needed to interpret retained text and reinsert the top-ranked candidates without support annotations at inference, training a compact classifier to optimize both relevance and referential completeness.
本文提出 Self-Modulating QKAN-based FWPs,用低秩生成的逐元素调制(作用于新提案分支、有界旧状态分支或两者)替代该广播门,并提出 Complementary Matrix Gating (CMG),在保持标量门控的有界凸更新和仿射前缀扫描结构的同时提供逐坐标的记忆控制。This work introduces Self-Modulating QKAN-based FWPs, which replace this broadcast gate with low-rank-generated element-wise modulation of the new-proposal branch, a bounded old-state branch, or both, and proposes Complementary Matrix Gating (CMG), which provides coordinate-wise memory control while preserving the bounded convex update and affine prefix-scan structure of scalar gating.
本文提出 DuplexGen,一个通过少量 slot 级人类偏好标注对 LLM 预测进行校准从而生成具有场景自适应轮换特性的对话框架;结果表明,使轮换合成具备场景特异性的关键是人类校准,而非单纯的语料规模或提示设计。DuplexGen is introduced, a framework for generating dialogues with scenario-adaptive turn-taking by calibrating LLM predictions against a small set of slot-level human preference annotations, and results show that human calibration, not corpus scale or prompt design alone, is what allows turn-taking synthesis to be scenario-specific.
RynnValue 是一个用于机器人操控的开源价值基础模型,用时间距离(即从观测到语言指定目标的有向 cost-to-go)替代任务内部锚点,将时间距离确立为通用机器人策略的可扩展监督目标和实用奖励接口。RynnValue, an open-source value foundation model for robotic manipulation that replaces task-internal anchors with temporal distance, the directed cost-to-go from an observation to the language-specified goal, establishes temporal distance as a scalable supervision target and practical reward interface for generalist robot policies.
我们提出 Ouroboros,一个自进化的 Agent harness,其工具、提示、上下文组装与核心实现通过已评审的 commit 持续改进,并成为后续工作的运行时。核心演化以两种模式推进:在递归自由演化中,改进本身就是任务,完成一个演化周期即可调度下一周期;在经验驱动核心演化中,常规工作和社交交互暴露的 bug、粗糙之处及低效上下文构造会引发已评审的结构变更。在 Terminal-Bench 2.1 上,Opus 5 运行取得 86.74% 的得分,为该基准报告的最佳结果。We present Ouroboros, a self-developing agent harness whose tools, prompts, context assembly, and core implementation improve through reviewed commits that become the runtime for later work. Core evolution proceeds in two modes. In recursive free evolution, improvement is itself a task, and completing one evolution cycle can schedule the next. In experience-driven core evolution, ordinary work and social interaction expose bugs, rough edges, and inefficient context construction that lead to reviewed structural changes. On Terminal-Bench 2.1, an Opus 5 run scores 86.74%, the best result reporte
本文提出 OasisKV,一种以显存为中心的 LLM 推理系统设计,通过在 LLM 解码期间将完整 KV-cache 存储与 HBM 解耦来缓解 HBM 容量压力,并观察到未来重要 token 可借助推测解码(SD)所起草的前瞻 token 被提前准确预测。OasisKV is presented, a memory-centric LLM inference system design that alleviates HBM capacity pressure by decoupling full KV-cache storage from HBM during LLM decoding and observes that future important tokens can be predicted accurately in advance using lookahead tokens drafted by speculative decoding (SD).
本文提出 Agent Memory Distillation (AMD),一个无需训练、通过分层记忆将结构化知识从大型教师智能体迁移到小型学生智能体的框架,一致优于现有基于记忆的基线方法。Agent Memory Distillation (AMD), a training-free framework that transfers structured knowledge from a large teacher agent to a small student agent through hierarchical memory, is proposed, which consistently outperforms existing memory-based baselines.
Steerling-8B 在与训练计算量多 2–16 倍的开源同侪模型对比中仍保持竞争力,表明存在一种不同的可扩展范式:可解释性可以被设计进训练过程中,并随规模放大而提升。Steerling-8B remains competitive with open peer models trained on substantially 2-16x more compute, suggesting a different scaling paradigm: interpretability can be designed into training, and it improves with scale.
本文提出 Factorized Hypothesis Search (FHS),在多个命名语义维度上维护多个部分解释,以支持结构化查询渲染、多假设检索以及维度级候选验证。This work proposes Factorized Hypothesis Search (FHS), which maintains multiple partial interpretations over named semantic dimensions that support structured query rendering, multi-hypothesis retrieval, and dimension-level candidate verification.
一个三阶段框架,用于减少建议编辑与当前图像之间的视觉不一致性,并将推荐 CTR、图像带走率以及用户平均对话轮次分别显著提升 39.90%(所有 p<0.05)。A three-stage framework to reduce visual inconsistencies between suggested edits and the current image, which significantly improves recommendation CTR, image take-away rate, and average conversation turns per user by 39.90% (all p<0.05).
本文提出 BDH-CQ,一个结合上下文学习与循环潜在推理的推理模型,通过在高维潜在空间中的迭代计算来求解查询,且无需将中间推理外化为语言。BDH-CQ is introduced, a reasoning model that combines in-context learning with recurrent latent reasoning that solves a query through iterative computation in a high-dimensional latent space, without verbalizing its intermediate reasoning.
本文识别出一种绕过 anti-distillation 机制、允许攻击者窃取专有模型推理能力的架构漏洞,并提出具体的密码学与系统级缓解措施以保障客户端推理安全。An architectural vulnerability is identified that circumvents anti-distillation mechanisms, allowing adversaries to extract a proprietary model's reasoning, as well as proposing concrete cryptographic and system-level mitigations to secure client-side reasoning.
本文提出 Ego-OSCAR,一种用于野外自我中心数据采集的开源硬件、低成本、头戴式立体惯性采集设备,旨在成为众包自我中心采集中最廉价且可辩护的载体,降低任何团队大规模采集自我中心数据的启动门槛。Ego-OSCAR is presented, an open-hardware, low-cost, head-mounted stereo-inertial capture device for egocentric data collection in the wild that aims to be the cheapest defensible substrate for crowdsourced egocentric capture, and to lower the activation energy for any team that wants to collect egocentric data at scale.
该形式化方法将短语定位、指代表达定位以及开放词汇检测等常见 grounding 任务统一起来,将文本分割、图像分割以及跨模态对齐视为单一的对应预测问题。This formulation unifies common grounding tasks, including phrase grounding, referring expression grounding, and open-vocabulary detection, by treating text segmentation, image segmentation, and cross-modal alignment as a single correspondence prediction problem.
一个结构定理将无记忆等变规则刻画为恰好由 Gram 矩阵决定的左预处理子;一个迁移定理将梯度流的路径性质推广到 common-scalar 流。A structure theorem characterizes the memoryless equivariant rules as exactly the Gram-determined left preconditioners, and a transfer theorem carries gradient flow's pathwise properties to common-scalar flows.
本工作研究图结构输入的特征学习技术,并在程序验证任务上取得 SOTA 性能,该任务需将子图与抽象数据结构进行匹配。This work studies feature learning techniques for graph-structured inputs and achieves state-of-the-art performance on a problem from program verification, in which subgraphs need to be matched to abstract data structures.
本文提出 KGCaRe,一种将神经检索与基于 LLM 生成 KG 的符号推理相结合的混合方法,在 Vanilla LLM、Code Prompt、Text Prompt、Think-on-Graph、Vanilla RAG 和 HybridContextQA 等 baseline 上 consistently 取得更优表现。KGCaRe is proposed, a hybrid approach that combines neural retrieval with symbolic reasoning over LLM-generated KGs that consistently outperforms existing baselines, including Vanilla LLM, Code Prompt, Text Prompt, Think-on-Graph, Vanilla RAG, and HybridContextQA.
Atelier 是一种面向艺术家风格图像生成的捷径感知控制状态规划框架,可提升艺术家级风格保真度,更忠实地保持源结构,并相较提示工程、检索增强与通用 agent 基线大幅减少捷径替换。Atelier, a shortcut-aware control-state planning framework for artist-grounded image generation, improves artist-level style fidelity, preserves source structure more faithfully, and substantially reduces shortcut substitution compared with prompt-engineered, retrieval-augmented, and general-purpose agent baselines.
本文提出 Combodied Agents——一种以人为中心的范式,借助软件工具、传感器、可穿戴设备、机器人与人工服务作为行动通道而非终极目标,在时间维度上感知、建模、预测并支持个体的人体状态轨迹。Combodied Agents is introduced, a human-centered paradigm that perceives, models, predicts, and supports individual human-state trajectories over time, using software tools, sensors, wearables, robots, and human services as action channels rather than end goals.
对抗性 Fréchet 距离 (AdvFD) 用经过校准的对抗性学习表示来补充 FD-Loss 中的静态表示目标,通过对抗方式最大化真实样本与生成样本之间的 Fréchet 差异;同时引入真实特征白化,对尺度与协方差几何进行归一化,从而稳定极小极大优化。Adversarial Fr\'echet Distance (AdvFD), which complements the static representation targets in FD-Loss with a calibrated adversarially learned representation that adversarially maximizes the Fr\'echet discrepancy between real and generated samples, and introduces real-feature whitening, which normalizes its scale and covariance geometry and stabilizes the min--max optimization.
本文研究基于开源大语言模型的无参考多语言机器翻译后训练,发现 on-policy 蒸馏能够达到但无法超越结合 checkpoint 插值的强化学习所确立的质量前沿。This work studies reference-free post-training for multilingual machine translation with open large language models and finds that on-policy distillation reaches, but does not surpass, the quality frontier achieved by RL with checkpoint interpolation.
本文提出 DistilVDR,一个 524M 的端到端 VDR 系统,通过逐点余弦对齐损失从一个 8B 视觉-语言教师模型进行双向蒸馏,并以非对称的纯编码器学生模型匹配 VDR 的文本查询与图像-文档输入不对称性,将视觉容量集中于文档端,查询端保持 70M 参数。This work presents DistilVDR, a 524M end-to-end VDR system distilled bilaterally from a single 8B vision-language teacher under a pointwise cosine alignment loss and matches VDR's text-query and image-document input asymmetry with an asymmetric encoder-only student that concentrates visual capacity on the document side and keeps the query side at 70M parameters.
结果表明,剪枝效果更取决于剪枝应用的位置,而非具体的评分规则:早期剪枝带来最大的端到端节省,后期剪枝主要用于细化最终的合成上下文。The results show that pruning effectiveness depends more on where pruning is applied than on the specific scoring rule: early pruning yields the largest end-to-end savings, while later pruning mainly refines the final synthesis context.
iFAN 提出 Adjusted Probability-Mask Ranking (APMR),将查询竞争与预测的掩码质量对齐,抑制高置信度但不准确的竞争者;同时采用 Cross-Layer Self-Distillation (CLSD) 将更强的中间预测传递至最终层。iFAN introduces Adjusted Probability-Mask Ranking (APMR), which aligns query competition with predicted mask quality and suppresses high-confidence but inaccurate competitors, and employs Cross-Layer Self-Distillation (CLSD) to transfer stronger intermediate predictions to the final layer.
本文引入静止状态 (rest-state) 公式化方法,从单一闭合构型重建关节物体——这是一种固有的不适定设定,几何、语义与运动先验在此弥补运动线索的缺失。This work introduces a rest-state formulation that reconstructs articulated objects from a single closed configuration, an inherently ill-posed setting where geometry, semantics, and motion priors compensate for the absence of motion cues.
本文提出 InSight-doc,一种智能体视觉感知框架,将视觉分辨率视为一种自适应的推理时资源,从低分辨率起步,选择性地放大高分辨率区域以获取更细粒度的证据,且不依赖任何外部检索器。This work proposes InSight-doc, an agentic visual perception framework that treats visual resolution as an adaptive reasoning-time resource that starts from low resolution and selectively zooms into high-resolution regions for finer evidence, without relying on any external retriever.
本工作提出在海量文本语料上使用新的自监督目标 PEGASUS 对大型 Transformer 编码器-解码器模型进行预训练,并证明其在所有 12 个下游数据集上按 ROUGE 分数衡量均取得 SOTA 性能This work proposes pre-training large Transformer-based encoder-decoder models on massive text corpora with a new self-supervised objective, PEGASUS, and demonstrates it achieves state-of-the-art performance on all 12 downstream datasets measured by ROUGE scores.
尽管方法简单,Decision Transformer 在 Atari、OpenAI Gym 和 Key-to-Door 任务上达到或超过 SOTA 无模型离线 RL 基线的性能Despite its simplicity, Decision Transformer matches or exceeds the performance of state-of-the-art model-free offline RL baselines on Atari, OpenAI Gym, and Key-to-Door tasks.
提出了能够平衡"延迟告警的上下文敏感成本"与"打断成本"的模型与推理流程,并通过对用户活动与通知内容的分析,描述了在不确定性下推理此类成本所面临的挑战Models and inference procedures that balance the context-sensitive costs of deferring alerts with the cost of interruption are presented and the challenge of reasoning about such costs under uncertainty via an analysis of user activity and the content of notifications is described.
本文提出连续性核 (Continuity Kernel, CK),一种激活契约,将提交前候选评估与原子状态激活解耦,将连续性定义为已接受分支头的连续且经过授权的谱系。The Continuity Kernel (CK), an activation contract that decouples off-commit candidate evaluation from atomic state activation from atomic state activation is presented, defining continuity as an unbroken, authorized lineage of accepted branch heads.
本文提出一种框架,用于从系统描述自动构建 DML 模型并将其表示为知识图谱 (KG-DML),以 RAG 和 LLM 作为使能工具。This study presents a framework for automated construction of DML models from system descriptions and their representation as Knowledge Graphs (KG-DML), using Retrieval-Augmented Generation and Large Language Models as enabling tools.