该工作表明带 sink 的 Sliding Window Attention(SWA)表现不逊于甚至优于后训练的 Linear Attention 模型,并建议改用 SWA 而非后训练线性模型。This work shows that Sliding Window Attention (SWA) with sinks performs as well or better than post-trained Linear Attention models, and recommends switching to SWA instead of post-training linear models.
论文
1640 张论文卡片 · OA 绿色
结果表明,可靠的 Agent 自我进化需要协同设计验证、状态接地、见证语义和恢复语言表达能力,而非仅依赖迭代提示。The results indicate that reliable agent self-evolution requires co-designing verification, state grounding, witness semantics, and recovery-language expressivity rather than relying on iterative prompting alone.
结果表明,广泛的 SFT 带来模型大部分能力提升;当失败检测精准时,turn-local 监督可发挥作用,且观察到的迁移主要集中在同族模型之间。The results suggest that broad SFT brings most of the model's capability improvement; turn-local supervision can be effective when failure detection is precise, with observed transfer concentrated primarily within-family.
一种神经符号架构,将逻辑知识图谱(LKG)与动态求解器路由相结合,并引入基于本体的 LKG,将逻辑规则和约束视为一等拓扑节点,从而支持对从文本中抽取的依赖关系进行显式建模。A Neuro-Symbolic architecture that integrates a Logical Knowledge Graph (LKG) with dynamic solver routing, and introduces an ontology-based LKG that treats logical rules and constraints as first-class topological nodes, enabling explicit modeling of dependencies extracted from text.
该候选名单可在不牺牲检索质量的前提下短路全语料库加密搜索:200–500 个候选即可在 5 个零样本语料库(规模从 25K 到 5.4M 文档)中与全语料库检索效果接近匹配。This shortlist short-circuits full-corpus cryptographic search without sacrificing retrieval quality: with 200-500 candidates, it closely matches full-corpus retrieval across five zero-shot corpora spanning 25K to 5.4M documents.
提出 MnIST-PRO 基准,将 Agent 感知能力隔离测试:把 MNIST 数字识别转换为带回看约束的序列式 glimpse 搜索任务,并评估了十个多模态模型,结果表明仅获取视觉证据是不够的,Agent 还必须能够构建并更新可靠的感知状态。MnIST-PRO is addressed, a benchmark that isolates agentic perception by converting MNIST digit recognition into a sequential, glimpse-based search task with lookback constraints, and ten multimodal models are evaluated, showing that simply acquiring visual evidence is not enough and agents must also be able to build and update a reliable perceptual state.
Pera 描述了一种持久化 Agent,围绕感知和控制组件组织,持续从情景任务执行、上下文及周围环境变化中感知服务相关信号,并利用这些信号构建生命周期任务。Pera describes a persistent agent organized around perception and control components that continually perceive service-relevant signals from episodic task executions, internal context, and changes in the surrounding environment, and use these signals to construct lifecycle tasks.
跨数据集分析表明,语义切分在具有显式关系线索的抽取数据集(如 GM-CIHT 和 DDI)上表现更优,而固定切分在密集生化抽取和二分类场景(如 ChemProt 和 ADE)下仍具竞争力甚至更强。Cross-dataset analysis shows that semantic chunking improves extraction datasets with explicit relation cues, such as GM-CIHT and DDI, while fixed chunking remains competitive or stronger for dense biochemical extraction and binary classification settings such as ChemProt and ADE.
GLM-5 在真实编码任务中展现出前所未有的能力,在端到端软件工程挑战的处理上超越既有基线,并提出了新颖的异步 Agent RL 算法,进一步提升了 RL 质量。GLM-5 demonstrates unprecedented capability in real-world coding tasks, surpassing previous baselines in handling end-to-end software engineering challenges and proposing novel asynchronous agent RL algorithms that further improve RL quality.
结果揭示了集成组合对精度–召回权衡的直接影响:异构跨范式集成通常提升精度,而同构 LLM 集成更常取得更高的整体 F1。It is revealed that ensemble composition directly affects the precision-recall trade-off: heterogeneous cross-paradigm ensembles generally improve precision, whereas homogeneous LLM ensembles more often achieve higher overall F1-scores.
通过发布紧凑的 7B 生成器和 2K Refiner,该工作致力于使原生音视频生成平民化,并为统一音视频生成建模的未来研究提供可及的基础。By releasing the compact 7B generator and 2K Refiner, this work seeks to democratize native audio-video generation and provide an accessible foundation for future research in unified audio-video generative modeling.
提出 MTPaperBananaBench,一个面向多轮图表生成的基准,包含 292 张图像和 3,518 条用户需求标注,并引入 PaperBanana-Interact,一个通过内部 critique-and-refine 循环来优化图表的多智能体系统。MTPaperBananaBench, a benchmark for multi-turn diagram generation containing 292 images annotated with 3,518 user requirements, is presented and PaperBanana-Interact, a multi-agent system that refines diagrams via an internal critique-and-refine loop is introduced.
提出 Verification-Aware Training,一种插件式框架,在每一步训练中模拟验证并将产生的 accept 与 reject 模式转化为监督信号,在数学、代码和聊天基准上提升了平均接受长度和实际加速比。Verification-Aware Training is introduced, a plug-in framework that simulates verification at every training step and turns the resulting accept and reject patterns into supervision and improves average acceptance length and wall-clock speedup across math, code, and chat benchmarks.
大量实验表明,SOTA embedding 模型和静态 decompose-and-rerank 范式存在关系盲区,而 BundleWeaver 取得了显著性能提升,凸显了从原子打分转向动态关系组合的必要性。Extensive experiments demonstrate that while state-of-the-art embedding models and static decompose-and-rerank paradigms suffer from relational blindness, BundleWeaver achieves substantial performance gains, highlighting the necessity of shifting from atomic scoring to dynamic relational composition.
该工作提出 Super Library Agent 问题:Agent 顺序生成 N 个相关应用组成的组合,同时维护一个共享的 Super Library 用于跨应用可复用组件,并解决了基于候选引导的代码块摘要抽取、抽取前的代码库整合,以及利用抽取 trace 和调用图信息的上下文感知迁移。This work introduces the Super Library Agent problem, where an agent sequentially generates a portfolio of N related applications while maintaining a shared Super Library of reusable cross-application components, and addresses candidate-guided extraction over code chunk summaries, pre-extraction codebase consolidation, and context-aware migration using extraction traces and call-graph information.
结果表明,在所测试的单调用、预填工具场景下,当偏好信息通过工具传入或需从原始产物中推断时,CoT 监控的可靠性可能下降。The results suggest that CoT monitoring may be less reliable when preference information arrives through tools or must be inferred from raw artifacts, within the single-call, prefilled-tool setting tested here.
该论文提出 SafeAtlas-VL,一个包含 1.5M 训练实例的数据集,将图像、请求和响应级判断置于五级有序量表上,并通过 target-conditioned tuning 训练 SafeAtlas Guard 系列模型,用于多模态安全检测。This paper introduces SafeAtlas-VL, a dataset of 1.5M training instances that places image-, request-, and response-level judgments on a five-level ordered scale, and trains the SafeAtlas Guard series of models via target-conditioned tuning for multimodal safety detection.
提出 Hash-Atlas 网络,将 3D 场景编辑重新表述为对 2D atlas 图像的操作,从而实现 2D 编辑与 3D 重建流程的解耦。The Hash-Atlas network is proposed, which reformulates 3D scene editing as operations on 2D atlas images, thereby achieving a workflow decoupling of the 2D editing and 3D reconstruction processes.
提出 PRISK,一个具备自动化数据生成和定制化指标的动态评估框架,用于揭示当前 LLM 个性化中的系统性局限以及个性化信息如何塑造其响应。PRISK is proposed, a dynamic evaluation framework with automated data generation and tailored metrics that uncovers systematic limitations in current LLM personalization and how personalized information shapes its responses.
本文利用 LSTM 网络学习视频序列表示,并通过在 UCF-101 与 HMDB-51 数据集上对人体动作识别这一监督学习任务进行微调来评估所学表示。This work uses Long Short Term Memory networks to learn representations of video sequences and evaluates the representations by finetuning them for a supervised learning problem - human action recognition on the UCF-101 and HMDB-51 datasets.
端到端天气预报系统直接从原始地球观测数据生成高质量的全球网格与站点预报,仅以数值天气预报流水线(含数据同化)一小部分的成本取而代之。这类系统是确定性的,不输出不确定性结果。本文通过在每个组件上附加一种随机机制,将 Aardvark Weather 模型概率化:在观测编码器中加入学习到的、依赖输入的噪声,以捕捉源自观测系统的偶然不确定性;在处理器中使用 Monte Carlo dropout,以捕捉认知不确定性。End-to-end weather forecasting systems produce skillful global gridded and station forecasts directly from raw Earth observations, replacing the numerical weather prediction pipeline, including data assimilation, at a fraction of its cost. These systems are deterministic and issue no uncertainty. Here we render the Aardvark Weather model probabilistic by attaching one stochastic mechanism to each component: learned, input-dependent noise at the observation encoder, capturing aleatoric uncertainty inherited from the observing system, and Monte Carlo dropout in the processor, capturing epistemic
评估四个 SOTA 模型发现,将角色置于语义无关的上下文中并不会抑制该角色的关联属性;相反,跨角色的属性集中度会上升(合并 BI $+0.047$)。Evaluating four state-of-the-art models finds that placing a role in a semantically unrelated context does not suppress role-linked attributes; instead, cross-role attribute concentration increases (pooled BI $+0.047$).
提出 SpanCalib-VLM,这是一种用于 SHROOM-Visions 共享任务的混合双系统,结合了多模态序列标注器(由 XLM-RoBERTa-Large 与 SigLIP 视觉编码器通过交叉注意力融合而成)与微调后的生成式 VLM(Qwen3.5-4B-SHROOM-SFT)。SpanCalib-VLM is presented, a hybrid dual-system for the SHROOM-Visions Shared Task that combines a multimodal sequence tagger, consisting of XLM-RoBERTa-Large fused with a SigLIP vision encoder via cross-attention, with the fine-tuned generative VLM (Qwen3.5-4B-SHROOM-SFT).
基于对数据的深入定性分析和既有理论研究,构建了一套新颖且全面的多模态错误信息分类法,并由此获得了关于社交媒体用户如何在实际场景中将图像与文本结合以传播错误信息的此前未被记录的洞见。A novel, comprehensive taxonomy of multimodal misinformation grounded in an in-depth qualitative analysis of the data and prior theoretical work is developed, which leads to previously undocumented insights about how social media users combine images with text to spread misinformation in the wild.
提出 EvoGenUI-Bench,一个面向多轮界面维护的基准,包含 150 个五轮任务,共计 750 轮,覆盖三种场景:信息呈现、可执行交互和工具驱动的外部状态。EvoGenUI-Bench is introduced, a benchmark for multi-turn interface maintenance comprising 150 five-turn tasks and 750 turns across three scenarios: information presentation, executable interaction, and tool-grounded external state.
本文提出 RECAP-Forcing,一种无需训练的推理方法,不增加任何可学习参数,在多个强基线上稳定提升视觉质量与语义保真度,并优于现有记忆方法。This work proposes RECAP-Forcing, a training-free inference method with no additional learnable parameters that consistently improves visual quality and semantic fidelity across multiple strong baselines and outperforms existing memory methods.
结果表明,即便在算力预算匹配的前提下,循环(looping)仍可提升 Transformer,提供了一种将深度复用转化为可衡量增益的实用方案。Results show that looping can improve Transformers even under budget matching, offering a practical recipe that turns depth reuse into measurable gains.
路由状态接口在模型的原生计算中统一了上下文记忆与持久的能力适配,将记忆从对历史上下文的被动记录重塑为维持与演化模型行为的主动基质。The routed-state interface unifies contextual memory and persistent capability adaptation within the model's native computation, reframing memory from a passive record of prior context into an active substrate for maintaining and evolving model behavior.
本文提出 DagEvo,通过利用 self-play 中求解器的失败历史来引导问题生成,无需外部任务资源,并表明混合生成、跨状态拼接的记忆状态更新以及双置信度过滤共同贡献了这些性能提升。DagEvo is introduced, which guides question generation using the solver's failure history from self-play, without external task resources, and shows that mixed generation, memory-state updates with cross-state stitching, and double-confidence filtering contribute to these gains.
本工作借鉴线性模型的 state space model(SSM)视角,提出三项核心方法改进并组合形成更具表达力的递推结构:源自 SSM 离散化的递推式、用于更丰富状态追踪的复数值状态更新规则,以及在不增加 decode 延迟前提下提升模型性能的多输入多输出(MIMO)建模。This work introduces three core methodological improvements inspired by the state space model (SSM) viewpoint of linear models that combine to form a more expressive recurrence derived from SSM discretization, a complex-valued state update rule that enables richer state tracking, and a multi-input, multi-output (MIMO) formulation for better model performance without increasing decode latency.
在持续多天、超过 70 轮迭代的部署中,HoH 自主开发出一款第一人称射击游戏,具备完整的主线剧情、完整实现的核心机制、可供人类游玩的体验、精美的画面与集成的音效。In a multi-day deployment with more than 70 iterations, HoH autonomously develops a first-person-shooter game, featuring a coherent storyline, fully implemented core mechanics, human-playable experience, polished visuals and integrated audio.
本文提出 ZimaBlue,一个可扩展的框架,用于从大规模视频中学习可泛化的 World Action Models (WAMs),并采用异步 Slow-Fast 双系统架构,使生成式 WAM 具备面向实时控制的实用性。This work introduces ZimaBlue, a scalable framework for learning generalizable World Action Models (WAMs) from large-scale video, and adopts an asynchronous Slow-Fast dual-system architecture to make generative WAMs practical for real-time control.
实验表明,该方法在大幅保留通用视觉-语言能力的同时,具备出色的 3D 感知与驾驶场景理解能力;在开环、伪闭环与闭环设定下的综合评估进一步显示其运动规划性能具有很强的竞争力。Experiments demonstrate strong 3D perception and driving scene understanding while largely preserving general vision-language capability and comprehensive evaluations across open-loop, pseudo-closed-loop, and closed-loop settings further show highly competitive motion-planning performance.
本文提出 UI-Venus-2,一个面向移动、Web 和桌面环境的通用 GUI 基础 Agent,采用统一的闭环推理-行动框架,并通过集成安全感知机制来确保关键操作的可控执行。UI-Venus-2 is presented, a general-purpose foundation GUI agent designed to operate across mobile, web, and desktop environments through a unified closed-loop reasoning-action framework that integrates safety-aware mechanisms to ensure controlled execution of consequential actions.
ACToR 在生成过程中识别关键 token,按需触发有针对性的检索,在这些决定性位置提供仓库上下文;并为稠密检索器设计了一种位置感知加权方法,以优先考虑对生成更具信息量的上下文。ACToR identifies critical tokens during generation and triggers targeted retrieval on demand to provide repository context at these decisive positions, and designs a position-aware weighting method for dense retrievers to prioritize context that is more informative for generation.