提出 Negative Self-Distillation(NSD),一个通过偏离有缺陷的推理而非模仿特权解来优化 LLM 的新框架,并一致优于 OPSD 及其他无标签、自举式强化学习(RL)基线。Negative Self-Distillation (NSD) is introduced, a new framework that optimizes LLMs by diverging from flawed reasoning rather than imitating privileged solutions, and consistently outperforms OPSD and other label-free, self-bootstrapping reinforcement learning (RL) baselines.
论文
1640 张论文卡片 · OA 绿色
实验结果表明 DRG-MAPPO 达到了 87% 的 SOTA 胜率,表明该框架在合作空战中有效平衡了关系建模、可解释性和优化稳定性。Experimental results demonstrate that DRG-MAPPO achieves a state-of-the-art win rate of 87%, suggesting that the framework effectively balances relational modeling, interpretability, and optimization stability for cooperative air combat.
引入 MetroLLM-Bench,一个包含 955 个用例的基准,用于测试语言模型作为交通信息亭策略层的能力,并评估了来自六个厂商的 26 个模型,其中 23 个被排名。This work introduces MetroLLM-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk, and evaluates twenty-six models from six vendors, of which twenty-three are ranked.
Video-LLM 的时间推理基准常以语言为中介,留有来自选项措辞、答案相关性或语言先验带来的语言捷径空间,因此提出 TempCloze,一个用于评估 Video-LLM 视觉时间推理能力的视频完形填空基准。Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors, so TempCloze is introduced, a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs.
提出 Generative Verification(GenV),通过复用语言模型的原生词表空间,将离线 Z3 等价性预言机蒸馏为无参考、连续参考等价分数,并从理论上证明基于结构、仅判决的验证启发式在这些欺骗性合法轨迹上的检测能力在数学上有界于随机水平。Generative Verification (GenV) is introduced, which distills an offline Z3-equivalence oracle into a reference-free, continuous reference-equivalence score by repurposing the language model's native vocabulary space and theoretically proves that structural, verdict-only verification heuristics are mathematically bounded to chance-level detection on these deceptively valid traces.
一个简单、无训练的框架,其中具备推理能力的视觉语言模型在 Wikipedia 上迭代搜索与推理,动态收集证据,表明推理与检索在稀有实体上具有互补性。A simple, training-free framework in which a reasoning-capable vision-language model iteratively searches and reasons over Wikipedia, gathering evidence dynamically, shows that reasoning and retrieval are complementary on rare entities.
本文提出了 SeedRG,一个用于缓解 knowledge leakage 并应对 benchmark aging 问题的半合成 benchmark 生成 pipeline。SeedRG is introduced, a semi-synthetic benchmark generation pipeline that mitigates knowledge leakage and addresses the issue of benchmark aging.
表明将 L2 推理泛化到未见语言的关键路径在于更广泛的语言覆盖、现成可用的多语言非推理数据,以及足够强的英语推理骨干,表明推理是一种与语言无关的行为,可通过精心数据混合在类型多样的语言间迁移。It is shown the path to generalizing L2 reasoning to held-out languages goes through broader language coverage, readily available multilingual non-reasoning data, and a sufficient English reasoning backbone, indicating that reasoning is a language-agnostic behavior that can be transferred across typologically diverse languages through careful data mixing.
研究面向实现的研究方法规范的可编码化就绪度,定义为其是否为胜任的实现者或编码 Agent 提供足够的方法学信息,以在不引入未支持假设的情况下构建预期方法。This work studies the codification readiness of implementation-facing research-method specifications, defined by whether they provide sufficient methodological information for a competent implementer or coding agent to construct the intended method without unsupported assumptions.
构建一个受控的纯自回归测试床,在文本、图像、文生图(T2I)和图生文(I2T)预测的多模态持续预训练中跟踪任务特定验证损失,表明更好的重建并不一定带来更低的任务特定损失或更强的下游性能,且图像分词器的选择在联合优化下会影响文本建模。A controlled pure-autoregressive testbed is built and track task-specific validation losses during multimodal continual pretraining across text, image, text-to-image (T2I), and image-to-text (I2T) prediction, showing that better reconstruction does not necessarily yield lower task-specific losses or stronger downstream performance, and that image tokenizer choice can affect text modeling under joint optimization.
引入 ActReview,一个由反驳引导的后训练框架,将论文特定的诊断连接到具体、有依据的修订计划,并引入 ActReview-Bench,一个包含 1,000 个实例的人工整理基准,用于评估诊断质量和修订实用性。This work introduces ActReview, a rebuttal-guided post-training framework that connects paper-specific diagnoses to concrete, grounded revision plans, and introduces ActReview-Bench, a human-curated benchmark of 1,000 instances for evaluating diagnostic quality and revision usefulness.
结果表明,在评估 harness 中使用 Adaptive Bridge 可将关键的订阅者尾端 p95 延迟从最高 15 s 降至 1.55 ms,覆盖所有损伤严重度,同时保留发布者配置的吞吐量。The results show that using the Adaptive Bridge in the evaluation harness reduces the critical subscriber tail p95 latency from up to 15 s to 1.55 ms across all impairment severities while preserving the publisher's configured throughput.
LIT(Latent Interface Training)是一个与框架无关的两阶段策略:先在无图像条件下建立空间目标条件化的动作先验,再通过姿态监督的潜在接口约束视觉条件化,可在保持或提升 LIBERO 平均成功率的同时改善 LIBERO-Plus 综合成功率。Latent Interface Training (LIT), a framework-agnostic two-stage strategy that first establishes a spatial-goal-conditioned action prior without images, then constrains visual conditioning through a pose-supervised latent interface, improves overall LIBERO-Plus success while preserving or improving average LIBERO success.
在推理、长上下文理解和 Agent 任务中,SAS 在不同注意力预算下均优于可训练的稀疏注意力基线,在紧预算下增益尤为显著,表明其上下文排序对下游任务更有效。Across reasoning, long-context understanding, and agentic tasks, SAS consistently outperforms trainable sparse attention baselines across attention budgets, with especially large gains under tight budgets, demonstrating more effective context ranking for downstream tasks.
提出一种物理引导框架,用于提升单图像部件感知 3D 生成效果,使其几何结构物理相容且连接稳定,并在部件接触面引入参数化连接器。This work proposes a physics-guided framework for improving single-image part-aware 3D generation with physically compatible geometry and stable connections, and introduces parameterized connectors at their contact surfaces.
本文提出 Benchmark Radar,一个用于 AI 基准检索与发现的活体数据库和搜索引擎,覆盖 LLM 评测、Agent 与工具使用基准、代码、推理、安全及领域评测。Benchmark Radar is presented, a living database and search engine for retrieval and discovery of AI benchmarks, covering LLM evaluation, agentic and tool-use benchmarks, coding, reasoning, safety, and domain-specific evaluations.
COBRA-Skills 是一个高效框架,将技能优化建模为在动态演化的候选空间上的预算式序贯优化,对 Agent harness 变化保持鲁棒,并在目标模型自身用于技能生成与优化时依然有效。COBRA-Skills is introduced, an efficient framework that formulates skill optimization as budgeted sequential optimization over a dynamically evolving candidate space and remains robust to changes in the agent harness and performs effectively when the target model itself is used for skill generation and refinement.
本文聚焦 LLM 驱动的 kernel generation 领域,给出现有方法的结构化综述,涵盖 LLM-based 方法与 agentic optimization workflow,并系统梳理了支撑该领域学习与评测的数据集与 benchmark。This survey addresses the gap in LLM-driven kernel generation by providing a structured overview of existing approaches, spanning LLM-based approaches and agentic optimization workflows, and systematically organizing the datasets and benchmarks that underpin learning and evaluation in this domain.
PLC-DPO 将每个偏好对的训练信号路由为 clean、flip 或 tie 三类,将噪声偏好学习从单纯过滤可疑样本重构为主动修正监督方向与强度。PLC-DPO is proposed to robustly optimize preferences by routing each pair's training signal as a clean, flip, or tie case, which reframes noisy preference learning as actively correcting supervision direction and strength rather than merely filtering suspicious examples.
这是首个端到端证明:一个七人独立团队可以训练出在 agentic 网络安全能力上领先的开源权重模型,且全部三个 checkpoint 在相近参数规模下均排名第一。This is the first end-to-end demonstration that a seven-person independent team can train open-weight models with leading agentic cyber capability and all three checkpoints rank 1st among models at comparable parameter scales.
提出了 DataFlex-RL,一个在统一 GRPO 方案下比较数据策略的评测平台;研究发现,改变数据策略会显著影响训练过程,但相对于均匀训练并未带来可复现的提升。DataFlex-RL, an evaluation platform for comparing choices under a common GRPO recipe, is introduced, finding that changing the data policy measurably changes the training process but does not produce a reproducible improvement over uniform training.
本研究探索在反馈有限的在线场景下,将 prompt 自适应路由到大语言模型专家以最大化响应质量,并提出了策略性地选择和观察奖励以最小化遗憾的算法。This work studies adaptive routing of prompts to large language model experts to maximize response quality in an online setting with limited feedback and proposes algorithms that strategically select and observe rewards to minimize regret.
提出了 Diverse Skill Routing,一个具备多样性感知能力的重排序框架,使用 Determinantal Point Process 在相关性与非冗余性之间取得平衡,在强 pointwise 重排序基线之上提升了召回率与完整覆盖率,且在多技能 query 上增益更大。Diverse Skill Routing is proposed, a diversity-aware reranking framework that uses a Determinantal Point Process to balance relevance and non-redundancy and improves recall and full coverage over a strong pointwise reranking baseline, with larger gains on multi-skill queries.
最终得到的 82M 参数模型 Wayu-Paxa-TTS-Edge 实现了无需参考音频的设备端泰语 TTS,并在三个系统中取得了最低的停顿位置错误率与词内停顿率。The resulting 82M-parameter model, Wayu-Paxa-TTS-Edge, enables on-device Thai TTS without reference audio and achieves the lowest pause-placement error and intra-word pause rates among the three systems.
研究发现,仅使用合成数据的训练即可具备竞争力;将 0.9B 参数的 PaddleOCR-VL-1.6 适配为 Wayu-Paxa-OCR-Zero,一个无需真实泰语文档页 OCR 标签即可适配的泰语 OCR 模型,表明仅合成数据训练即可具备竞争力。It is found that synthetic-only training can be competitive, and the 0.9B-parameter PaddleOCR-VL-1.6 is adapted into Wayu-Paxa-OCR-Zero, a Thai OCR model adapted without OCR labels from real Thai document pages, showing that synthetic-only training can be competitive.
提出面向检索增强生成(RAG)的通用微调方案——RAG 模型融合预训练参数化记忆与非参数化记忆进行语言生成;研究发现,相较 SOTA 的纯参数化 seq2seq 基线,RAG 模型生成的文本更具针对性、更多样且更符合事实。A general-purpose fine-tuning recipe for retrieval-augmented generation (RAG) -- models which combine pre-trained parametric and non-parametric memory for language generation, and finds that RAG models generate more specific, diverse and factual language than a state-of-the-art parametric-only seq2seq baseline.
提出了一个能力门控(competence gate),根据已解决的结果估计领域级的源权重,将不确定的估计向全局权重收缩,并重新校准合并后的预测,为基于已测边际价值的选择性模型使用提供了实用方案。A competence gate is introduced that estimates domain-level source weights from resolved outcomes, shrinks uncertain estimates toward a global weight, and recalibrates the pooled forecast and provides a practical approach for selective model use based on measured marginal value.
本文提出 CATS,一种 self-speculative decoding 框架,在内存受限设备上结合 memory budget 与 parameter offloading pattern 进行级联式 verify 与 correction,在设备峰值显存与单独运行 target model 相当的前提下,最大化 token acceptance rate 与端到端加速比。CATS, a self-speculative decoding framework that conducts cascaded verification and correction based on the memory budget and parameter offloading patterns on memory-limited devices, is proposed, which maximizes token acceptance rate and end-to-end speedup while keeping the peak memory footprint on the device equal to that of the target model alone.
Pelican-Sim 在轨迹、场景、物体、具身和视角变化上的定性泛化能力,凸显其作为通用世界模型模拟器的潜力。Qualitative generalization across trajectory, scene, object, embodiment, and viewpoint shifts highlights Pelican-Sim's potential as a general-purpose world model simulator.
提出即插即用的 Feature Recovery Module (FRM),可在保持宿主网络冻结的前提下,将退化编码器特征映射到与干净图像对齐的表征;该模块提升了场景级检测、CLIP/SigLIP2 特征恢复以及全部四项物体级 VLM 任务,且退化越严重增益越大。The Feature Recovery Module (FRM), a plug-and-play module that maps degraded encoder features to pristine-aligned representations while keeping the host frozen, is proposed, which improves scene-level detection, CLIP/SigLIP2 feature recovery, and all four object-level VLM tasks, with larger gains under more severe degradation.
该工作提交于 ECCV 2026 Wearable AI Challenge 的 EgoProactive 赛道,在大模型组排名第一、≤2B 组排名第二,表明对当前任务而言视觉定位比标注量更为重要。This submission to the EgoProactive track of the ECCV 2026 Wearable AI Challenge is presented, which ranked first in the large-model division and second in the<=2B division, suggesting that visual grounding is more important than annotation volume for this task.
该系统仅用一个 2B 视觉-语言模型,通过一次贪心前向传播即可回答关于十分钟第一人称视频的多选题,并仅用大型 Agentic pipeline 1.1% 的参数即达到其 89% 的准确率。The system is a single 2B vision-language model that answers multiple-choice questions about ten-minute egocentric videos in one greedy forward pass, and reaches 89% of the accuracy of the large agentic pipeline using 1.1% of its parameters.
对突发物理危险的反应既是具身智能的重要检验,也是将多模态大语言模型 (MLLMs) 部署为家庭机器人决策核心的硬性要求;ReactHuman 是首个面向类人反应式决策的物理驱动基准。Reacting to sudden physical hazards is both a meaningful test of embodied intelligence and a hard requirement for deploying multimodal large language models (MLLMs) as the decision core of household robots, and ReactHuman is the first physics-grounded benchmark for human-like reactive decision-making.
即便在公式、变量与笔记访问条件一致时,向 Program-Solve 接口加入 executor 对不同开源权重模型的帮助程度各异,且无论哪种情况都无法替代经过验证的公式或可靠的变量抽取。Adding an executor to the Program-Solve interface helps some open-weight models more than others even under matched formula, variable and note access, and is not a substitute for verified formulas or reliable variable extraction either way.
提出结构感知的 RAG 框架 ReMoMask-2:耦合 Hierarchical Bidirectional Momentum 对比学习以对齐全局与部件级特征与文本;采用 Semantic Spatial-Temporal Attention (SSTA) 实现拓扑感知的融合;通过 Topology Structured Masking (TSM) 借助自适应掩码强化鲁棒的部件级 grounding。ReMoMask-2, a structure-aware RAG framework coupling Hierarchical Bidirectional Momentum contrastive learning to align global and part-level features with text; Semantic Spatial-Temporal Attention (SSTA) for topology-aware fusion; and Topology Structured Masking (TSM) to force robust part-level grounding via adaptive masking are introduced.
提出 MobileVLA-R1 2.0,一个 RL 增强的 VLA 框架,显式地将结构化具身推理与可执行的移动机器人控制耦合,并引入推理条件化的动作解码器,将多模态推理表征映射到任务级动作目标,再由机器人控制器翻译为具身特定的指令。This work proposes MobileVLA-R1 2.0, an RL-enhanced VLA framework that explicitly couples structured embodied reasoning with executable mobile robot control, and introduces a reasoning-conditioned action decoder that maps multimodal reasoning representations to task-level action targets, which are subsequently translated into embodiment-specific commands by robot controllers.