本文提出 GameXpert-Bench,将游戏开发的三个生命周期阶段以 coding agent 操作化为三条互补的基准轨道,并发现现有 agent 在生成可玩基础框架和实现显式需求方面更可靠,而在发现缺陷、验证运行时行为以及跨变更保持功能一致性方面能力较弱。GameXpert-Bench is introduced, which operationalizes the three lifecycle stages of game development with a coding agent as three complementary benchmark tracks, and finds current agents are more reliable at producing playable foundations and implementing explicit requirements than at discovering defects, verifying runtime behavior, and preserving functionality across changes.
论文
1749 张论文卡片
本文提出 Block3D,一种块级扩散框架,将离散形状 token 序列划分为连续块,自回归地生成各块,并联合去噪当前块内的所有 token,同时引入置信度引导的块内修正机制,在每块定稿前对低置信度 token 进行修订。Block3D is proposed, a block-wise diffusion framework that partitions the discrete shape-token sequence into contiguous blocks, generates the blocks autoregressively, and jointly denoises all tokens within the current block and introduces confidence-guided intra-block correction, which revises low-confidence tokens before each block is finalized.
本文建立了一个包含超过 1,000 个细粒度编辑概念的综合性层次化分类体系,并提出一种密集监督训练策略,将多个互不干扰的概念合成到单个图像对中,显著提升了训练效率和模型整体性能。A comprehensive hierarchical taxonomy featuring over 1,000 fine-grained edit concepts is established and a dense supervision training strategy that synthesizes multiple non-interfering concepts into single image pairs is proposed that significantly enhances both training efficiency and overall model performance.
本文提出 ARC(Advantage Regularization via Conditioning),一种通过策略条件化 rollout 分组来恢复更公平的相对比较、并结合混合奖励与熵正则化的训练方法。The proposed ARC (Advantage Regularization via Conditioning), a training recipe that restores fairer relative comparison through strategy-conditioned rollout grouping, together with hybrid rewards and entropy regularization, is proposed.
本文提出 Knowledge Triage 框架,对 agent 知识库的每一行按类型分类,并为每种类型配置独立的保留策略,同时开源发布 AgentArtifactCorpus 数据集、分类器及参考实现。Knowledge Triage, a framework that classifies each line of an agent's knowledge base by type and routes each type through its own retention policy, is addressed, and AgentArtifactCorpus, the classifier, and the reference implementation are released.
本文介绍 TianoForge,这是面向 TianoCore 开源 UEFI 固件开发生态中 bug 分诊的集成方案,部署 AI(具体为机器学习)领域的 SOTA 方法以实现自动化 bug 分诊。This integrated approach to bug triage in the TianoCore open-source UEFI firmware development ecosystem, called TianoForge, deploys the state of the art in artificial intelligence, specifically machine learning, to enable automated bug triage.
本文通过全面超越传统开环基线,证明了当前主流的单体上下文扩展策略是一种因相关性衰减而受到惩罚的架构陷阱,并确立了以顺序、反馈驱动的编排作为生成式搜索的确定性范式。By dominating classical open-loop baselines, this work proves that the prevailing strategy of monolithic context widening is an architectural trap penalized by relevance decay and establishes sequential, feedback-driven orchestration as the definitive paradigm for generative search.
提出 TRUSTMARGIN,一种免训练、即插即用的仲裁层,利用模型自身的似然对两个候选进行打分,在不微调、无需外部评判或额外生成的情况下,在直接回答与 RAG 之间进行选择。TRUSTMARGIN is proposed, a training-free, plug-and-play arbitration layer that scores the two existing candidates with the model's own likelihoods and selects between Direct and RAG without fine-tuning, external judges, or additional generation.
本文提出 User Behavioral Densing Law,为大规模用户表示学习中的 tokenization 配置选择提供实用指导,并开发了 ALGN——一种自适应变长 tokenization 方法,可改善容量分配。The proposed User Behavioral Densing Law is proposed, providing practical guidance for tokenization configuration selection in large-scale user representation learning and ALGN, an adaptive variable-length tokenization method that improves capacity allocation, is developed.
EvoMap 的结果表明,经过验证的执行经验可以被保留并共享为可复用的外部资源,使模型能够提升长工作流完成度,而无需反复承担经验探索的全部成本。EvoMap results show that verified execution experience can be retained and shared as a reusable external resource, enabling models to improve long-workflow completion without repeatedly paying the full cost of experience discovery.
介绍 AutoResearch,一个连接 Idea Generation 与 Idea Execution 的两阶段系统,分别解决研究思路如何形成与如何通过实验可靠验证的问题,展示「实验前先夯实洞见、接受前先夯实结论」的研究流程。AutoResearch is introduced, a two-stage system that connects Idea Generation with Idea Execution to address both how research ideas are formed and how they are reliably established through experimentation to demonstrate a research process in which meaningful insight is grounded before experimentation and conclusions are grounded before acceptance.
目标是提供无需数周超参搜索即可部署的方案,直接从原始未压缩模型蒸馏 4-bit 学生模型,并以开源权重形式发布为 Hypernova-60B。The aim is a recipe deployable without a multi-week hyper-parameter search, which distills the 4-bit student directly from the original, uncompressed model, and is released open-weight as Hypernova-60B.
覆盖 4 个 FR backbone 与 2 个 VLM 编码器;EXPL-FR 无需访问模型架构,支持身份级、单图及差异式解释,并在三种监督设置(人工标注、VLM 伪标签、完全 prompt 驱动的审计)下针对真实核验行为进行属性级审计基准测试。This work covers four FR backbones and two VLM encoders, EXPL-FR needs no architecture access, and supports identity-level, per-image, and differential explanations, and benchmark attribute-level auditing under three supervision settings, human labels, VLM pseudo-labels, and the authors' fully prompt-driven audit, against real verification behavior.
介绍 ClawProBench:基于 OpenClaw(具备 workspace 工具及浏览、记忆、消息、调度、skill、subagent 等原生能力的实时 agent 运行时)实例化的 trace-aware、runtime-native agent 评估基准。ClawProBench is presented, a trace-aware benchmark for runtime-native agent evaluation instantiated on OpenClaw, a live agent runtime with workspace tools and native surfaces for browsing, memory, messaging, scheduling, skills, and subagents.
介绍 PinSieve——大规模内容质量流水线中的生产级案例:一个选择性 vision-language-model Serving Agent,仅处理轻量上游模型无法覆盖的 grey-zone 切片,在线暴露标量路由评分,并保留受控的人工升级通道。This work presents PinSieve, a production case study in a large-scale content-quality pipeline, a selective vision-language-model Serving Agent that operates only on the grey-zone slice left unresolved by lightweight upstream models, exposes a scalar routing score online, and preserves controlled human escalation.
AtlasNav 减少了 Evidence Blindness,更早实现完整证据,在 PhantomWiki 上对语料结构与规模变化保持鲁棒,并在异构企业数据上取得领先性能。AtlasNav reduces Evidence Blindness, reduces Evidence Blindness, realizes complete evidence earlier, remains robust to corpus-structure and scale shifts on PhantomWiki, and achieves leading performance on heterogeneous enterprise data.
介绍 WeMM-Embedding,一族通用多模态嵌入模型,支持文本、图像、视频、视觉文档及任意交错的多模态输入,输出维度灵活,在多个公开基准上取得 SOTA 表现。WeMM-Embedding is presented, a family of universal multimodal embedding models supporting text, images, videos, visual documents, and arbitrarily interleaved multimodal inputs with flexible output dimensions and achieves leading performance on multiple public benchmarks.
介绍 GameCleaner,一个无需 mask 的游戏 UI 移除模型,结合多模态语义理解与视频编辑能力,整体 VideoReward 较在带 UI 数据上训练的模型提升 6.83%。GameCleaner is proposed, a mask-free gameplay UI removal model that combines multimodal semantic understanding with video editing capabilities and improves overall VideoReward by 6.83% over those trained on UI-overlaid data.
本综述首次以统一框架系统研究智能眼镜,形式化第一人称数据流与受限任务效用,并提出覆盖采集、反应式感知、上下文辅助、持续状态、受控行动与具身耦合的 L0–L5 框架。This survey is the first to systematically study smart glasses through a unified framework, formalizing first-person data flow and constrained task utility, and introducing an L0-L5 framework spanning capture, reactive perception, contextual assistance, persistent state, governed action, and embodied coupling.
介绍 LAION-BVD,面向多模态学习的大规模开放视频数据集,包含从 CommonCrawl 收集的 1.3B 条平台特定视频 URL,并通过抽取场景切换帧,将视频帧作为图文数据的替代来源加以探索。This work presents LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl, and explores video frames as an alternative source of image-text data by extracting scene-changing frames.
介绍 Entropy-Valley (EV):一种无需训练的画布长度选择器,通过 all-mask 前向的预测平均熵对候选目标画布打分,并挑选出 backbone 最「准备好」填充的画布。Entropy-Valley (EV), a training-free length selector that scores candidate target canvases by mean predictive entropy from all-mask forward passes and selects the canvas the backbone is most prepared to fill, is introduced.
本文研究视频 MLLM 的 RL 后训练样本效率与可扩展性,并提出 OraRL——一种随模型规模与数据规模共同 scaling 的解耦 advantage estimator,在 0.8B 到 9B backbone 上均超越其基线,并在 100k prompts 规模下超越 GRPO。The sample efficiency and scalability of RL post-training for video MLLMs and introduces OraRL, a decoupled advantage estimator that scales with model size and data, surpassing its backbone from 0.8B to 9B and GRPO up to 100k prompts.
介绍 CyberFactory,一个统一的开源框架,贯通 PoC 生成、漏洞修补与 CyberQA 三大任务中的数据构建、轨迹合成与模型训练。CyberFactory is introduced, a unified open-source framework that connects data construction, trajectory synthesis, and model training across proof-of-concept (PoC) generation, vulnerability patching, and cybersecurity question answering (CyberQA).
介绍 DREAM,一种自主优化控制架构:在不替换现有流水线的前提下叠加感知可感知、可编排、可审计的策略层,支持将 agentic meta-control 作为工业推荐的一种可行范式。This work presents DREAM (Developing Recommender Engine with Agentic Methods), an autonomous optimization control architecture that adds a perception-aware, orchestrable, and auditable policy layer atop existing pipelines without replacing them, supporting agentic meta-control as a viable paradigm for industrial recommendation.
本文指出 LLM agent 中信息抽取与动作之间的 state-transmission 失效,并展示 handoff 变换如何在保留状态内容的同时削弱其对下游动作的约束。This work identifies a state-transmission failure between information extraction and action in large language model agents, and shows how handoff transformations can retain state content while weakening its constraints on downstream action.
本文提出 MoTE(Mixture of Task Experts),一种将大语言模型前馈网络转化为任务特定专家同时保持多模态 backbone 共享的 decoder 架构,并在五个 COIN 基准上使用显式任务路由进行评估。This work proposes MoTE (Mixture of Task Experts), a decoder architecture that converts large language model feed-forward networks into task-specific experts while keeping the multimodal backbone shared and evaluates it on five COIN benchmarks using explicit task routes.
AgentRoom 是一种面向并发编码 Agent 的实时协同编辑协议,通过在 CRDT 合并的共享文件系统上将文件级 claim、status 和 broadcast 暴露为 MCP 工具,且运行间的差异小于 CLI-stable 模型。AgentRoom is a realtime collaborative editing protocol for concurrent coding agents that exposes file-level claim, status, and broadcast as MCP tools on a CRDT-merged shared filesystem and has less run-to-run variation than CLI-stable models.
行为拓扑更多由部署 harness 决定而非 LLM 本身,为安全审计和运行时监控提供一种与模型无关的结构化 primitive,并同时满足两类预测目标。Behavioral topology is shaped more by the deployment harness than by the LLM, providing a model-agnostic structural primitive for safety auditing and runtime monitoring, and addresses both prediction goals.
本工作开发了一个 SOTA 的 paged attention kernel,完全基于领域特定即时编译语言 Triton 构建,在 NVIDIA 与 AMD GPU 上均达到 SOTA 性能。This work develops a state-of-the-art paged attention kernel that builds exclusively on the domain-specific just-in-time compiled language Triton to achieve state-of-the-art performance on both NVIDIA and AMD GPUs.
The Station 被评估为一个开放世界多 agent 环境,不同模型家族的 AI agent 在其中无需中心协调器或脚本化流程即可共同追求同一研究目标,并提供发现产生过程的透明记录。The Station is evaluated, an open-world multi-agent environment in which AI agents from different model families pursue a shared research goal without a central coordinator or scripted pipeline, providing a transparent record of how discoveries emerged.
本文提出 Secure On-Policy Distillation (SecOPD),提供 token 级反馈以指导防御性微调,并能泛化到训练中完全未见过的领域。This paper proposes Secure On-Policy Distillation (SecOPD) that provides token-level feedback to guide defensive fine-tuning, and generalizes to domains completely unseen in training.
本文提出 GigaBrain-0.7,一种跨多种机器人 embodiment 泛化能力显著增强的 embodied foundation model,并引入一阶段对齐训练,联合优化 vision-language 理解和多 embodiment 动作生成。This work presents GigaBrain-0.7, an embodied foundation model with substantially improved generalization across diverse robot embodiments, and introduces one-stage alignment training that jointly optimizes vision-language understanding and multi-embodiment action generation.
本文对 20 余个近期 MLLM 进行大规模评估,结果表明视频指令跟随对当前模型仍具挑战,尤其是在涉及多约束、语义约束或需要根据视频内容选择正确分支/路径的复杂条件结构时。This work conducts a large-scale evaluation of more than 20 recent MLLMs and shows that video instruction following remains challenging for current models, especially for instructions with many constraints, semantic constraints, or complex conditional structures that require selecting the correct branch or path based on video content.
评估表明,RubSE 在终轮和最佳轮设置下均显著优于朴素 self-evolution,refinement 轨迹更稳定,且轨迹级性能上限更高。Evaluations demonstrate that RubSE substantially outperforms na\"ive self-evolution in final-round and best-round settings, achieving more stable refinement trajectories and a higher trajectory-level performance ceiling.
本文变化 handoff 方向、时机与接口,对比保留仓库状态下的全轨迹传输、压缩与轨迹移除,发现偏好接口随方向反转:减少 LC-model 轨迹信息可提升 escalation 质量,而移除 HC-model 轨迹则会降低 downshift 质量。This work varies handoff direction, timing, and interface, comparing full-trajectory transfer, compaction, and trajectory removal while preserving the repository state, and finds that the preferred interface also reverses with direction: reducing LC-model trajectory information improves escalation quality, whereas removing the HC-model trajectory reduces downshift quality.