为推进任务课程聚焦于大规模词分类,本竞赛设置两条互补赛道:Deep 赛道面向同一被试内部的大规模词分类,目标是追求最佳性能;Broad 赛道面向跨被试泛化。Advancing the curriculum of tasks to focus on word classification to focus on word classification at scale, two complementary tracks are presented in this competition: the Deep track targets within-subject word classification at scale, aiming at the best possible performance; the Broad track targets cross-subject generalisation.
论文
1640 张论文卡片 · OA 绿色
提出 FactoSR,一个因子化强化学习框架,显式解读视觉投影所塌缩的维度,并指出强化显式的、因子化的 4D 一致性是将 VLM 演化为稳健、具有世界感知能力的推理器的关键一步。FactoSR is presented, a factorized reinforcement learning framework that explicitly interpret the dimensions collapsed by visual projection, and suggests that reinforcing explicit, factorized 4D consistency is a critical step toward evolving VLMs into robust, world-aware reasoners.
本文提出 Dr. Claw,一个开源工作空间,将现有编码 Agent 可执行文件封装在可控且可审计的人机协同工作流中,而非引入另一个自主 Agent。Dr. Claw is presented, an open-source workspace that wraps existing coding-agent executors in a controllable and auditable human-in-the-loop workflow rather than introducing another autonomous agent.
本文提出 Memanto,一种面向 agentic 人工智能的通用记忆层,挑战了"必须依赖知识图谱复杂度才能实现高保真 agent 记忆"的普遍假设,并取得 SOTA 准确率。Memanto is introduced, a universal memory layer for agentic artificial intelligence that challenges the prevailing assumption that knowledge graph complexity is necessary to achieve high fidelity agent memory and achieves state of the art accuracy scores.
FlowBalance 是一种以验证器为锚定的自改进方法,学习完整响应上的归一化分布,在 Qwen3-4B 和 Qwen3-8B 上相对 FlowRL 提升了平均性能,同时改善了训练速度与稳定性,避免了直接 OPSD 响应长度坍缩,并在受控的 AIME24 诊断中表现出更高的正确策略多样性。FlowBalance, a verifier-grounded self-improvement method that learns a normalized distribution over complete responses, improves average performance over FlowRL on both Qwen3-4B and Qwen3-8B, while also improving training speed and stability, avoiding direct OPSD's response-length collapse, and exhibiting higher correct-strategy diversity in a controlled AIME24 diagnostic.
本文为该设定引入新基准,并基于该基准评估了九种修订方法,包括序贯反思与并行采样变体,使用 gpt-oss-20b/120b、gpt-5.4-mini 以及 qwen3.5-9b/27b/122b 进行测试。A new benchmark for this setting is introduced, and nine revision methods are evaluated, including sequential reflection and parallel sampling variants, using gpt-oss-20b/120b, gpt-5.4-mini, and qwen3.5-9b/27b/122b on the benchmark.
本文提出 Teacher-Gated On-Policy Distillation(教师门控的在线策略蒸馏),其核心原则是在引入密集监督前以 prompt 级别验证教师可靠性;在全部六个单领域设定下优于 Vanilla OPD,并在多领域训练下于两种规模上取得更高的七项基准平均成绩。Teacher-Gated On-Policy Distillation is introduced, built on the principle that teacher reliability should be verified at the prompt level before dense supervision is admitted, and outperforms Vanilla OPD in all six single-domain settings and achieves higher seven-benchmark averages at both scales under multi-domain training.
本文提出 KnowChange,一个知识引导的变化数据合成框架,利用预训练视觉语言模型作为知识源,从变化前场景与期望变化类型推理合理的变化位置与类别转移,在统一框架下灵活合成多样变化类型。KnowChange is introduced, a knowledge-guided change data synthesis framework that leverages pretrained vision-language models as knowledge sources to reason about plausible change locations and class transitions from pre-change scenes and desired change types and enables flexible synthesis of diverse change types within a unified framework.
大语言模型(LLM)的成功很大程度上归功于 next-token prediction(NTP),但其自回归(AR)结构需要缓慢的串行 token 生成。为克服这一瓶颈,我们提出扩散增强 LLM,一类新模型,在使用扩散从该分布中并行采样多个 token 的同时定义 AR 模型分布。我们将这些模型的参数解耦为两组:AR 权重,使用标准 NTP 目标训练;轻量扩散权重,训练用于同时生成多个 token。扩散权重Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressive (AR) structure requires slow, sequential token generation. To overcome this bottleneck, we introduce diffusion-augmented LLMs, a new class of models that defines an AR model distribution while using diffusion to draw multiple tokens in parallel from that distribution. We decouple the parameters of these models into two sets: AR weights, trained using the standard NTP objective, and lightweight diffusion weights, trained to generate multiple tokens simultaneously. The diffusion weight
我们提出 ENEAS,一种用于实例追踪与语义发现的统一且文本可提示的方法。包括 SAM 3 在内的文本可提示分割模型仍存在时间幻觉、空间碎片化与语义误分类问题:目标离开视野时无法报告目标缺失;极端特写下只分割局部纹理而非完整目标;将视觉特征置于本体事实之上,从而把雕像、绘画或反射等视觉相似的物体误分割为目标。We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-promptable segmentation models, including the latest foundation models such as SAM 3, still suffer from temporal hallucinations, spatial fragmentation, and semantic misclassification: they fail to report target absence when an object leaves the field of view, segment local textures instead of the complete object during extreme close-ups, and prioritize visual features over ontological reality, so that visually similar artifacts such as statues, paintings, or reflections are segmented as targe
EmbodiedSkills 是一个统一框架,将每项技能决策视为执行提案:运行时先检查前置条件再执行,执行后再验证结果,为将底层 VLA 策略转化为闭环具身系统提供可训练且可检视的 Agent 层。EmbodiedSkills is a unified framework that treats each skill decision as an execution proposal: the runtime checks its prerequisites before execution and verifies the outcome afterward and provides a trainable and inspectable agent layer for turning low-level VLA policies into closed-loop embodied systems.
本综述将崩溃视为受三处杠杆共同支配的症状:信号施加位置(即 token 加权方式)、教师所见信息(即特权信息的性质)、以及信号变化时机(即教师的动态性与引导衰减)。This review treats collapse as a symptom governed by three levers: where the signal is applied, that is, how tokens are weighted; what the teacher is shown, that is, the nature of the privileged information; and when the signal changes, that is, the teacher's dynamics and the decay of guidance.
本文对一个双节点 split-LLM 训练系统进行系统安全案例研究:其隐私评估通过,却遗留一条未被测试的可观测信道;系统因此并不安全——包括跨训练步骤累积观测在内的五类攻击从未被测量。A systems-security case study of a two-node split-LLM training system whose privacy evaluation passed while leaving an observable channel untested, but the system is not thereby safe: five classes of attack, including those accumulating observations across training steps, were never measured.
结果表明数据组成控制安全性与可用性的权衡,且安全对齐应在预期拒答边界的两侧进行评估。The results show that data composition controls the safety and usability trade-off, and that safety alignment should be evaluated on both sides of the intended refusal boundary.
本文提出 MatryoshkaLoRA,一种受 Matryoshka 启发、面向 LoRA 的通用训练框架,通过在已有 LoRA adapter 之间插入一个固定的、经精心设计的对角矩阵来按比例缩放其子秩,从而学习到准确的层次化低秩表示。MatryoshkaLoRA is proposed, a general, Matryoshka-inspired training framework for LoRA that learns accurate hierarchical low-rank representations by inserting a fixed, carefully crafted diagonal matrix between the existing LoRA adapters to scale their sub-ranks accordingly.
因果基础模型是经过预训练的神经网络,可在全新的数据集上通过上下文学习估计因果量(如平均处理效应),无需模型更新。Causal foundation models are pretrained neural networks that estimate causal quantities, such as the average treatment effect, on entirely new datasets using in-context learning without requiring model updates.
本文提出 Conformal Relevance 框架,利用上下文学习的示例筛选与集成构造打分函数,在保持覆盖的同时以极低人工成本提升简洁性。The Conformal Relevance framework is introduced which uses in-context learning example curation and ensembling to create a score function which maintains coverage while improving conciseness with minimal manual input.
提出了 Gander,一个原生多模态双工交互模型,基于 MiniCPM-o 4.5 构建,并通过异步 Agent 循环进一步适配实时交互,同时开源其模型、代码和数据,以推动社区进一步研究与开发。Gander is presented, a native multimodal duplex interaction model that builds on MiniCPM-o 4.5 and is further adapted for realtime interaction with an asynchronous agent loop and is released together with its models, code, and data to facilitate further research and development in the community.
Mask Forcing 是一种双噪声掩码展开策略,通过扰动自回归学生的自展开来缓解由 reverse-KL 模式寻求引发的模式坍缩,能以更高视觉质量高效改进多种自回归视频扩散蒸馏方法,且无需引入真实视频数据或额外后训练阶段。Mask Forcing, a Dual-Noise Masking Rollout strategy that perturbs the AR student self-rollout to mitigate mode collapse induced by reverse-KL mode seeking, improves multiple AR video diffusion distillation methods with higher visual quality efficiently, without incorporating real video data or additional post-training stages.
本文提出 AuK,一个开源基础模型,通过自然语言指令与音频上下文的统一接口整合语音生成与编辑,在零样本与指令控制的语音生成以及通用指令引导编辑上取得领先性能。AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context, is introduced and leading performance on zero-shot and instruction-controlled speech generation and general instruction-guided editing is demonstrated.
本文证明空间覆盖与三维推理性能相关,并提出 CoVeR,一种仅使用 token 坐标、不依赖学习信号的确定性、无训练选择器,在三项三维推理基准上均超越既有 SOTA,并可作为即插即用模块泛化到四种 VLM。It is shown that spatial coverage is associated with 3D reasoning performance and CoVeR, a deterministic, training-free selector that uses only token coordinates, with no learned signals is introduced, which outperforms prior SOTAs on all three 3D reasoning benchmarks and generalizes as a plug-and-play module tested across four VLMs.
本文提出 TANGO,首个面向语言条件人形机器人在杂乱环境中通行的全身视觉语言导航框架,在视觉语言导航任务中达到 SOTA,并在需要避障的困难场景中超越强模块化基线。This work introduces TANGO, the first whole-body vision-language navigation framework for language-conditioned humanoid traversal in cluttered environments, and demonstrates state-of-the-art performance in vision-language navigation, while outperforming strong modular baselines in navigating challenging scenes requiring obstacle negotiation.
TransNormal-2 是基于 FLUX.2 的整流流框架,采用单步确定性推理,针对 VAE 解码器两侧的重建退化问题,施加 RGB 引导的残差修正以降低局部于边界的解码误差,且不自由改写粗预测。TransNormal-2, a FLUX.2-based rectified-flow framework with single-step deterministic inference that addresses VAE reconstruction degradation on both sides of the VAE decoder, and applies an RGB-guided residual correction to reduce boundary-localized decoding errors without freely rewriting the coarse prediction.
本文提出 BeaconKV,一种免训练的 KV cache 压缩方法,通过为每个全局查询簇维护紧凑代表性 beacon query 来预测哪些 KV 对将被重访,无需存储完整查询历史。BeaconKV is proposed, a training-free KV cache compression method that maintains beacon queries, compact representatives for each global query cluster, to anticipate which KV pairs will be revisited without storing the entire query history.
本文发布大规模机器人操控数据集与基准,用于诊断 VLA 模型的具身推理能力,并将 RoboSPA 确立为面向更强、更可靠、更具泛化性具身 Agent 的挑战性诊断基准。A large-scale robotic manipulation dataset and benchmark for diagnosing embodied reasoning in VLA models, and establishes RoboSPA as a challenging diagnostic benchmark for developing more capable, reliable, and generalizable embodied agents.
本文提出通过让 evaluator 与解决方案协同进化来自动化 evaluator 的设计,并证明突破 evaluation 瓶颈可释放 ADRS 的潜力,为下一代数据系统生成高度优化、可部署的代码。This work proposes automating the design of evaluators by co-evolving them with the solutions, demonstrating that addressing the evaluation bottleneck unlocks the potential of ADRS to generate highly optimized, deployable code for next-generation data systems.
Marigold V2 在应用于其他稠密回归任务(如表面法向量估计与本征图像分解)时取得 SOTA 结果。Marigold V2 achieves state-of-the-art results when applied to other dense regression tasks, such as surface normals estimation and intrinsic image decomposition, and achieves state-of-the-art results when applied to other dense regression tasks, such as surface normals estimation and intrinsic image decomposition.
提出 A*-Thought-V2,一个由 LLM 引导的几何动力学框架,将 CoT 建模为隐状态轨迹,并以显式-隐式交错潜在架构替代硬删除,引入更广义的软目标以促进更丰富的步骤级特征学习。A*-Thought-V2 is presented, a geometric dynamics of LLM guided framework that models CoT as a hidden-state trajectory and replaces hard deletion with an explicit-implicit interleaved latent architecture that reflects broader soft targets that encourage richer step-level feature learning.
发现空间步态参数对视觉域偏移更敏感,且更好的 HMR 重建并不一定带来下游步态估计的改进;提出 GaitXFormer,作为直接基于 RGB 的参考模型用于步态参数估计。It is found that spatial gait parameters are more sensitive to visual domain shift and that improved HMR reconstruction alone does not necessarily translate to improved downstream gait estimation, and GaitXFormer is introduced as a direct RGB reference model for estimating gait parameters.
提出 OpenWAM,一个开源研究栈,将世界-动作预训练转化为可控的实验项目,并发布完整栈,包括基础设施、评估协议、预训练模型和数据配方,以促进未来研究。This work introduces OpenWAM, an open research stack that turns world-action pretraining into a controlled experimental program, and releases the full stack, including infrastructure, evaluation protocols, pretrained models, and data recipes, to facilitate future research.
发现合作者专业能力在模型早期层最易被解码,到网络中点前降至接近随机水平,并基于合成语料以单个模型作为初步验证。It is found that partner expertise is most decodable in the early layers and falls to near chance before the midpoint of the network, and one model on a synthetic corpus is used as an initial demonstration.
提出 UCF-Net,一个不确定性感知级联融合网络,利用 CLIP 语言对齐的语义先验与 DINO 自监督视觉结构先验,在域内与跨域评估中均取得最优平均 AUC。UCF-Net is proposed, an uncertainty-aware cascaded fusion network that harnesses CLIP's language-aligned semantic priors and DINO's self-supervised visual-structure priors and achieves the best mean AUC among the evaluated methods in both in-domain and cross-domain evaluations.
提出一个前馈式生成 Transformer,用于直接的单视图与多视图图像重光照,完全绕过显式本征属性估计,并采用置换不变的位置编码对称处理无序多视图输入,避免序列偏差。This work introduces a feed-forward generative Transformer for direct single- and multi-view image relighting that entirely bypasses explicit intrinsic property estimation and employs permutation-invariant positional encodings to symmetrically process unordered multi-view inputs without sequential bias.
我们提出 Cadence,一种面向数值时序的误差有界有损压缩器,将 3.3 亿参数的时序基础模型(Google TimesFM-3)与自适应算术编码器相结合,保证每个样本满足 |x_t - x̂_t| ≤ τ。一项负面结论限定了设计空间:在无损编码场景下,基础模型毫无价值,因为节省的比特数仅与预测器精度呈对数关系 Δb = log_2(MAE_old / MAE_new)。因此 TimesFM-3 相对 32 阶线性预测器 1.51 倍的精度优势,在 20.28 比特中仅换取 0.60 比特,中位数增益仅 +0.03%。误差有界编码仅在一点上突破了这一限制。We present Cadence, an error-bounded lossy compressor for numeric time series pairing a 330M-parameter time-series foundation model (Google TimesFM-3) with an adaptive arithmetic coder, guaranteeing |x_t-x_t|leτ on every sample. One negative result constrains the design space: for lossless coding a foundation model is worth nothing, because bits saved are logarithmic in predictor accuracy, Δb=log_2(MAE_{old}/MAE_{new}). So the 1.51times advantage TimesFM-3 holds over a 32-tap linear predictor buys 0.60 bits of 20.28, a median gain of +0.03%. Error-bounded coding escapes this at one point: once
提出一个自演化 Agent 框架,通过识别 Agent 轨迹中不稳定、低一致性的步骤并将其转化为情节记忆,以供后续运行调用,从而缩小一致性 gap。This work presents a self-evolving agent framework that reduces the consistency gap by identifying unstable, low-consistency steps in agent trajectories and converting them into episodic memory the agent can draw on in future runs.
Noah 是一个时间感知、任务无关的生成式 Transformer 模型,对完整多模态患者旅程进行表征与预测,是该领域首个真正整体化的生成模型,支持自回归预测,并具备可选的时间控制、零样本分类与反事实干预模拟能力。Noah is a time-aware, task-agnostic, generative transformer model representing and forecasting the full multimodal patient journey, and is the first truly holistic generative model in its field, enabling autoregressive forecasting with optional time control, zero-shot classification, and counterfactual intervention simulation.