提出 EASEL benchmark,用于评估受控的灵巧视觉工具使用,以参考引导的视觉重建为主代理任务:agent 逐步在画布上绘制以匹配参考图像。EASEL is proposed, a benchmark evaluating a controlled instance of dexterous visual tool use that adopts reference-guided visual reconstruction as its primary proxy task: the agent incrementally paints a canvas to match a reference image.
论文
1753 张论文卡片
本工作将 embedding 缩放作为正交于稀疏度缩放的强有力维度加以探索,并推出 LongCat-Flash-Lite,一个从零训练的 68.5B 参数、约 30 亿激活参数的模型,不仅超越参数等量级的 MoE 基线,还对同规模现有模型展现出卓越竞争力。This work explores embedding scaling as a potent, orthogonal dimension for scaling sparsity and introduces LongCat-Flash-Lite, a 68.5B parameter model with ~3B activated trained from scratch that not only surpasses parameter-equivalent MoE baselines but also exhibits exceptional competitiveness against existing models of comparable scale.
提出 StepGuard,一种 step-level guard model,可对已完成的 agent trajectory 进行审计并在工具动作执行前进行检查;并引入 StepGen,一种自动数据引擎,能在风险步生成上下文相同但动作不同的安全与不安全 trajectory。This work proposes StepGuard, a step-level guard model that can audit completed agent trajectories and check tool actions before they are executed, and introduces StepGen, an automatic data engine that generates safe and unsafe trajectories with the same context but different actions at the risky step.
StarHarness 提供了一种实用方法,通过根据基线失败行为对任务进行分层、将 proposer 可见的搜索任务与 proposer 隐藏的选择任务分离,并为评估泛化能力保留 held-out 任务,从而缓解工具密集型企业任务中持续的 model-environment mismatch。StarHarness offers a practical way to reduce persistent model-environment mismatch in tool-rich enterprise tasks by stratifying tasks according to baseline failure behavior, separating proposer-visible search tasks from proposer-hidden selection tasks, and reserves held-out tasks for evaluating generalization.
提出 LoopArena,一个用于评估一个模型在长时任务中引导另一个独立 coding agent 能力的 benchmark,并在执行范围与成本各不相同的三个互补设置下评估该能力。This work introduces LoopArena, a benchmark for evaluating how well one model can guide a separate coding agent through a long-running task, and evaluates this ability in three complementary settings that differ in execution scope and cost.
本文综述了 agentic artifact creation,即有状态的构建过程,其中 AI 系统实体地构造或修改交付物,且中间观察结果会重定向后续工作;并提出了若干原则,使承诺与责任显式化、将反馈转化为针对性修复,以及在变更后重新验证受影响的状态。This survey examines agentic artifact creation, which is defined as stateful construction in which an AI system materially constructs or revises a deliverable and intermediate observations redirect later work, and formulate principles for keeping commitments and responsibility explicit, turning feedback into targeted repair, and revalidating affected state after change.
提出 Decoupled Block Attention,在保留共享 video-query 上下文访问的同时消除跨 box 依赖,并结合用于时间边界与空间几何的 localization-aware policy optimization;并行 tube generation 被证明是视频中 autoregressive 定位的一种高效替代方案。Decoupled Block Attention is introduced, which preserves access to shared video-query context while eliminating cross-box dependencies, together with localization-aware policy optimization for temporal boundaries and spatial geometry, and parallel tube generation is shown to be an efficient and effective alternative to autoregressive localization in videos.
该框架将循环序列模型中的 temporal alignment、plasticity、forgetting 与 bounded rehearsal 分离,并结合数值稳定的 positive-decay renormalization,使其在语言建模上保持竞争力,同时提升在可变位数加法任务上的长度外推能力。This framework separates temporal alignment, plasticity, forgetting, and bounded rehearsal in recurrent sequence models, together with numerically stable positive-decay renormalization, to remain competitive in language modeling and improve length extrapolation on variable-digit addition.
GGSS——Geodesic-Gated Spherical Steering——一种保范的干预方法,在单位超球面上发现反事实偏差子空间,沿 geodesic arc 引导视觉 token,并通过自适应门控聚焦于承载更强人口统计信号的 token。GGSS---Geodesic-Gated Spherical Steering---a norm-preserving intervention that discovers a counterfactual bias subspace on the unit hypersphere, steers visual tokens along geodesic arcs, and uses an adaptive gate to focus correction on tokens that carry stronger demographic signal.
本文提出 Cross-lingual Ranking Preference Optimization (CRPO),一种新框架,利用来自英语的鲁棒偏好知识来促进目标语言的偏好对齐,从而增强语言适应性与输出质量。This paper proposes Cross-lingual Ranking Preference Optimization~ (CRPO), a novel framework that leverages robust preference knowledge from English to facilitate preference alignment in the target language, thereby enhancing language adaptation and output quality.
提出 Intention Distillation (INDI),将行为级意图蒸馏到动作解码器中,并以目标依赖的方式组织下游预测;研究表明,动作解码器显式建模其生成行为的语义目标能够带来收益。Intention Distillation (INDI) is proposed, which distills behavior-level intent into the action decoder and organizes downstream predictions in an objective-dependent manner, and shows that action decoders benefit from explicitly modeling the semantic objective of the behavior they generate.
本文提出 DART-SD (Diamond-topology Aware Retrieval and Tuning for Self-Distillation),一种将范式从全局 forcing 转向拓扑引导的局部修正的新框架,相对传统 full-trajectory 基线取得显著提升。This work proposes DART-SD (Diamond-topology Aware Retrieval and Tuning for Self-Distillation), a novel framework that shifts the paradigm from global forcing to topology-guided localized correction, and significantly outperforms traditional full-trajectory baselines.
ViT-5 的设计与当代基础模型实践保持一致,可作为对 vanilla ViT 的直接替换升级方案,适用于 2020 年代中期的视觉骨干网络,并为生成建模提供更强大的骨干。With a design aligned with contemporary foundation-model practices, ViT-5 offers a simple drop-in upgrade over vanilla ViT for mid-2020s vision backbones and serves as a stronger backbone for generative modeling.
该研究将 358,896 条消息切分为 22,329 个时间连贯的块,并构建三种搜索表示:raw_text、生成的 summary,以及将 summary 与 raw_text 片段及其他固定文本相结合的 embedding_text。This study segmented 358,896 messages into 22,329 temporally coherent chunks and constructed three search representations: raw_text, a generated summary, and embedding_text, which combines a summary with a raw-text excerpt and other fixed text.
提出 CamoDocs,一种通过将对抗文档伪装在良性内容中来避免直接包含查询的投毒攻击,并表明 TrustRAG 等以擦除为主的聚类防御可降低 ASR,但会在 NeoQA 等依赖检索的基准上造成显著的效用下降。CamoDocs is proposed, a poisoning attack that avoids direct query inclusion by camouflaging adversarial documents among benign content, and shows that erasure-heavy clustering defenses such as TrustRAG can reduce ASR, but only with substantial utility drops on retrieval-dependent benchmarks such as NeoQA.
该工作表明带 sink 的 Sliding Window Attention(SWA)表现不逊于甚至优于后训练的 Linear Attention 模型,并建议改用 SWA 而非后训练线性模型。This work shows that Sliding Window Attention (SWA) with sinks performs as well or better than post-trained Linear Attention models, and recommends switching to SWA instead of post-training linear models.
结果表明,可靠的 Agent 自我进化需要协同设计验证、状态接地、见证语义和恢复语言表达能力,而非仅依赖迭代提示。The results indicate that reliable agent self-evolution requires co-designing verification, state grounding, witness semantics, and recovery-language expressivity rather than relying on iterative prompting alone.
结果表明,广泛的 SFT 带来模型大部分能力提升;当失败检测精准时,turn-local 监督可发挥作用,且观察到的迁移主要集中在同族模型之间。The results suggest that broad SFT brings most of the model's capability improvement; turn-local supervision can be effective when failure detection is precise, with observed transfer concentrated primarily within-family.
一种神经符号架构,将逻辑知识图谱(LKG)与动态求解器路由相结合,并引入基于本体的 LKG,将逻辑规则和约束视为一等拓扑节点,从而支持对从文本中抽取的依赖关系进行显式建模。A Neuro-Symbolic architecture that integrates a Logical Knowledge Graph (LKG) with dynamic solver routing, and introduces an ontology-based LKG that treats logical rules and constraints as first-class topological nodes, enabling explicit modeling of dependencies extracted from text.
该候选名单可在不牺牲检索质量的前提下短路全语料库加密搜索:200–500 个候选即可在 5 个零样本语料库(规模从 25K 到 5.4M 文档)中与全语料库检索效果接近匹配。This shortlist short-circuits full-corpus cryptographic search without sacrificing retrieval quality: with 200-500 candidates, it closely matches full-corpus retrieval across five zero-shot corpora spanning 25K to 5.4M documents.
提出 MnIST-PRO 基准,将 Agent 感知能力隔离测试:把 MNIST 数字识别转换为带回看约束的序列式 glimpse 搜索任务,并评估了十个多模态模型,结果表明仅获取视觉证据是不够的,Agent 还必须能够构建并更新可靠的感知状态。MnIST-PRO is addressed, a benchmark that isolates agentic perception by converting MNIST digit recognition into a sequential, glimpse-based search task with lookback constraints, and ten multimodal models are evaluated, showing that simply acquiring visual evidence is not enough and agents must also be able to build and update a reliable perceptual state.
Pera 描述了一种持久化 Agent,围绕感知和控制组件组织,持续从情景任务执行、上下文及周围环境变化中感知服务相关信号,并利用这些信号构建生命周期任务。Pera describes a persistent agent organized around perception and control components that continually perceive service-relevant signals from episodic task executions, internal context, and changes in the surrounding environment, and use these signals to construct lifecycle tasks.
跨数据集分析表明,语义切分在具有显式关系线索的抽取数据集(如 GM-CIHT 和 DDI)上表现更优,而固定切分在密集生化抽取和二分类场景(如 ChemProt 和 ADE)下仍具竞争力甚至更强。Cross-dataset analysis shows that semantic chunking improves extraction datasets with explicit relation cues, such as GM-CIHT and DDI, while fixed chunking remains competitive or stronger for dense biochemical extraction and binary classification settings such as ChemProt and ADE.
GLM-5 在真实编码任务中展现出前所未有的能力,在端到端软件工程挑战的处理上超越既有基线,并提出了新颖的异步 Agent RL 算法,进一步提升了 RL 质量。GLM-5 demonstrates unprecedented capability in real-world coding tasks, surpassing previous baselines in handling end-to-end software engineering challenges and proposing novel asynchronous agent RL algorithms that further improve RL quality.
结果揭示了集成组合对精度–召回权衡的直接影响:异构跨范式集成通常提升精度,而同构 LLM 集成更常取得更高的整体 F1。It is revealed that ensemble composition directly affects the precision-recall trade-off: heterogeneous cross-paradigm ensembles generally improve precision, whereas homogeneous LLM ensembles more often achieve higher overall F1-scores.
通过发布紧凑的 7B 生成器和 2K Refiner,该工作致力于使原生音视频生成平民化,并为统一音视频生成建模的未来研究提供可及的基础。By releasing the compact 7B generator and 2K Refiner, this work seeks to democratize native audio-video generation and provide an accessible foundation for future research in unified audio-video generative modeling.
提出 MTPaperBananaBench,一个面向多轮图表生成的基准,包含 292 张图像和 3,518 条用户需求标注,并引入 PaperBanana-Interact,一个通过内部 critique-and-refine 循环来优化图表的多智能体系统。MTPaperBananaBench, a benchmark for multi-turn diagram generation containing 292 images annotated with 3,518 user requirements, is presented and PaperBanana-Interact, a multi-agent system that refines diagrams via an internal critique-and-refine loop is introduced.
提出 Verification-Aware Training,一种插件式框架,在每一步训练中模拟验证并将产生的 accept 与 reject 模式转化为监督信号,在数学、代码和聊天基准上提升了平均接受长度和实际加速比。Verification-Aware Training is introduced, a plug-in framework that simulates verification at every training step and turns the resulting accept and reject patterns into supervision and improves average acceptance length and wall-clock speedup across math, code, and chat benchmarks.
大量实验表明,SOTA embedding 模型和静态 decompose-and-rerank 范式存在关系盲区,而 BundleWeaver 取得了显著性能提升,凸显了从原子打分转向动态关系组合的必要性。Extensive experiments demonstrate that while state-of-the-art embedding models and static decompose-and-rerank paradigms suffer from relational blindness, BundleWeaver achieves substantial performance gains, highlighting the necessity of shifting from atomic scoring to dynamic relational composition.
该工作提出 Super Library Agent 问题:Agent 顺序生成 N 个相关应用组成的组合,同时维护一个共享的 Super Library 用于跨应用可复用组件,并解决了基于候选引导的代码块摘要抽取、抽取前的代码库整合,以及利用抽取 trace 和调用图信息的上下文感知迁移。This work introduces the Super Library Agent problem, where an agent sequentially generates a portfolio of N related applications while maintaining a shared Super Library of reusable cross-application components, and addresses candidate-guided extraction over code chunk summaries, pre-extraction codebase consolidation, and context-aware migration using extraction traces and call-graph information.
结果表明,在所测试的单调用、预填工具场景下,当偏好信息通过工具传入或需从原始产物中推断时,CoT 监控的可靠性可能下降。The results suggest that CoT monitoring may be less reliable when preference information arrives through tools or must be inferred from raw artifacts, within the single-call, prefilled-tool setting tested here.
该论文提出 SafeAtlas-VL,一个包含 1.5M 训练实例的数据集,将图像、请求和响应级判断置于五级有序量表上,并通过 target-conditioned tuning 训练 SafeAtlas Guard 系列模型,用于多模态安全检测。This paper introduces SafeAtlas-VL, a dataset of 1.5M training instances that places image-, request-, and response-level judgments on a five-level ordered scale, and trains the SafeAtlas Guard series of models via target-conditioned tuning for multimodal safety detection.
提出 Hash-Atlas 网络,将 3D 场景编辑重新表述为对 2D atlas 图像的操作,从而实现 2D 编辑与 3D 重建流程的解耦。The Hash-Atlas network is proposed, which reformulates 3D scene editing as operations on 2D atlas images, thereby achieving a workflow decoupling of the 2D editing and 3D reconstruction processes.
提出 PRISK,一个具备自动化数据生成和定制化指标的动态评估框架,用于揭示当前 LLM 个性化中的系统性局限以及个性化信息如何塑造其响应。PRISK is proposed, a dynamic evaluation framework with automated data generation and tailored metrics that uncovers systematic limitations in current LLM personalization and how personalized information shapes its responses.
本文利用 LSTM 网络学习视频序列表示,并通过在 UCF-101 与 HMDB-51 数据集上对人体动作识别这一监督学习任务进行微调来评估所学表示。This work uses Long Short Term Memory networks to learn representations of video sequences and evaluates the representations by finetuning them for a supervised learning problem - human action recognition on the UCF-101 and HMDB-51 datasets.