介绍 AutoResearch,一个连接 Idea Generation 与 Idea Execution 的两阶段系统,分别解决研究思路如何形成与如何通过实验可靠验证的问题,展示「实验前先夯实洞见、接受前先夯实结论」的研究流程。AutoResearch is introduced, a two-stage system that connects Idea Generation with Idea Execution to address both how research ideas are formed and how they are reliably established through experimentation to demonstrate a research process in which meaningful insight is grounded before experimentation and conclusions are grounded before acceptance.
论文
1640 张论文卡片 · OA 绿色
目标是提供无需数周超参搜索即可部署的方案,直接从原始未压缩模型蒸馏 4-bit 学生模型,并以开源权重形式发布为 Hypernova-60B。The aim is a recipe deployable without a multi-week hyper-parameter search, which distills the 4-bit student directly from the original, uncompressed model, and is released open-weight as Hypernova-60B.
覆盖 4 个 FR backbone 与 2 个 VLM 编码器;EXPL-FR 无需访问模型架构,支持身份级、单图及差异式解释,并在三种监督设置(人工标注、VLM 伪标签、完全 prompt 驱动的审计)下针对真实核验行为进行属性级审计基准测试。This work covers four FR backbones and two VLM encoders, EXPL-FR needs no architecture access, and supports identity-level, per-image, and differential explanations, and benchmark attribute-level auditing under three supervision settings, human labels, VLM pseudo-labels, and the authors' fully prompt-driven audit, against real verification behavior.
介绍 ClawProBench:基于 OpenClaw(具备 workspace 工具及浏览、记忆、消息、调度、skill、subagent 等原生能力的实时 agent 运行时)实例化的 trace-aware、runtime-native agent 评估基准。ClawProBench is presented, a trace-aware benchmark for runtime-native agent evaluation instantiated on OpenClaw, a live agent runtime with workspace tools and native surfaces for browsing, memory, messaging, scheduling, skills, and subagents.
介绍 PinSieve——大规模内容质量流水线中的生产级案例:一个选择性 vision-language-model Serving Agent,仅处理轻量上游模型无法覆盖的 grey-zone 切片,在线暴露标量路由评分,并保留受控的人工升级通道。This work presents PinSieve, a production case study in a large-scale content-quality pipeline, a selective vision-language-model Serving Agent that operates only on the grey-zone slice left unresolved by lightweight upstream models, exposes a scalar routing score online, and preserves controlled human escalation.
AtlasNav 减少了 Evidence Blindness,更早实现完整证据,在 PhantomWiki 上对语料结构与规模变化保持鲁棒,并在异构企业数据上取得领先性能。AtlasNav reduces Evidence Blindness, reduces Evidence Blindness, realizes complete evidence earlier, remains robust to corpus-structure and scale shifts on PhantomWiki, and achieves leading performance on heterogeneous enterprise data.
介绍 GameCleaner,一个无需 mask 的游戏 UI 移除模型,结合多模态语义理解与视频编辑能力,整体 VideoReward 较在带 UI 数据上训练的模型提升 6.83%。GameCleaner is proposed, a mask-free gameplay UI removal model that combines multimodal semantic understanding with video editing capabilities and improves overall VideoReward by 6.83% over those trained on UI-overlaid data.
本综述首次以统一框架系统研究智能眼镜,形式化第一人称数据流与受限任务效用,并提出覆盖采集、反应式感知、上下文辅助、持续状态、受控行动与具身耦合的 L0–L5 框架。This survey is the first to systematically study smart glasses through a unified framework, formalizing first-person data flow and constrained task utility, and introducing an L0-L5 framework spanning capture, reactive perception, contextual assistance, persistent state, governed action, and embodied coupling.
介绍 LAION-BVD,面向多模态学习的大规模开放视频数据集,包含从 CommonCrawl 收集的 1.3B 条平台特定视频 URL,并通过抽取场景切换帧,将视频帧作为图文数据的替代来源加以探索。This work presents LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl, and explores video frames as an alternative source of image-text data by extracting scene-changing frames.
介绍 Entropy-Valley (EV):一种无需训练的画布长度选择器,通过 all-mask 前向的预测平均熵对候选目标画布打分,并挑选出 backbone 最「准备好」填充的画布。Entropy-Valley (EV), a training-free length selector that scores candidate target canvases by mean predictive entropy from all-mask forward passes and selects the canvas the backbone is most prepared to fill, is introduced.
本文研究视频 MLLM 的 RL 后训练样本效率与可扩展性,并提出 OraRL——一种随模型规模与数据规模共同 scaling 的解耦 advantage estimator,在 0.8B 到 9B backbone 上均超越其基线,并在 100k prompts 规模下超越 GRPO。The sample efficiency and scalability of RL post-training for video MLLMs and introduces OraRL, a decoupled advantage estimator that scales with model size and data, surpassing its backbone from 0.8B to 9B and GRPO up to 100k prompts.
介绍 CyberFactory,一个统一的开源框架,贯通 PoC 生成、漏洞修补与 CyberQA 三大任务中的数据构建、轨迹合成与模型训练。CyberFactory is introduced, a unified open-source framework that connects data construction, trajectory synthesis, and model training across proof-of-concept (PoC) generation, vulnerability patching, and cybersecurity question answering (CyberQA).
介绍 DREAM,一种自主优化控制架构:在不替换现有流水线的前提下叠加感知可感知、可编排、可审计的策略层,支持将 agentic meta-control 作为工业推荐的一种可行范式。This work presents DREAM (Developing Recommender Engine with Agentic Methods), an autonomous optimization control architecture that adds a perception-aware, orchestrable, and auditable policy layer atop existing pipelines without replacing them, supporting agentic meta-control as a viable paradigm for industrial recommendation.
本文指出 LLM agent 中信息抽取与动作之间的 state-transmission 失效,并展示 handoff 变换如何在保留状态内容的同时削弱其对下游动作的约束。This work identifies a state-transmission failure between information extraction and action in large language model agents, and shows how handoff transformations can retain state content while weakening its constraints on downstream action.
本文提出 MoTE(Mixture of Task Experts),一种将大语言模型前馈网络转化为任务特定专家同时保持多模态 backbone 共享的 decoder 架构,并在五个 COIN 基准上使用显式任务路由进行评估。This work proposes MoTE (Mixture of Task Experts), a decoder architecture that converts large language model feed-forward networks into task-specific experts while keeping the multimodal backbone shared and evaluates it on five COIN benchmarks using explicit task routes.
AgentRoom 是一种面向并发编码 Agent 的实时协同编辑协议,通过在 CRDT 合并的共享文件系统上将文件级 claim、status 和 broadcast 暴露为 MCP 工具,且运行间的差异小于 CLI-stable 模型。AgentRoom is a realtime collaborative editing protocol for concurrent coding agents that exposes file-level claim, status, and broadcast as MCP tools on a CRDT-merged shared filesystem and has less run-to-run variation than CLI-stable models.
行为拓扑更多由部署 harness 决定而非 LLM 本身,为安全审计和运行时监控提供一种与模型无关的结构化 primitive,并同时满足两类预测目标。Behavioral topology is shaped more by the deployment harness than by the LLM, providing a model-agnostic structural primitive for safety auditing and runtime monitoring, and addresses both prediction goals.
本工作开发了一个 SOTA 的 paged attention kernel,完全基于领域特定即时编译语言 Triton 构建,在 NVIDIA 与 AMD GPU 上均达到 SOTA 性能。This work develops a state-of-the-art paged attention kernel that builds exclusively on the domain-specific just-in-time compiled language Triton to achieve state-of-the-art performance on both NVIDIA and AMD GPUs.
The Station 被评估为一个开放世界多 agent 环境,不同模型家族的 AI agent 在其中无需中心协调器或脚本化流程即可共同追求同一研究目标,并提供发现产生过程的透明记录。The Station is evaluated, an open-world multi-agent environment in which AI agents from different model families pursue a shared research goal without a central coordinator or scripted pipeline, providing a transparent record of how discoveries emerged.
本文提出 Secure On-Policy Distillation (SecOPD),提供 token 级反馈以指导防御性微调,并能泛化到训练中完全未见过的领域。This paper proposes Secure On-Policy Distillation (SecOPD) that provides token-level feedback to guide defensive fine-tuning, and generalizes to domains completely unseen in training.
本文提出 GigaBrain-0.7,一种跨多种机器人 embodiment 泛化能力显著增强的 embodied foundation model,并引入一阶段对齐训练,联合优化 vision-language 理解和多 embodiment 动作生成。This work presents GigaBrain-0.7, an embodied foundation model with substantially improved generalization across diverse robot embodiments, and introduces one-stage alignment training that jointly optimizes vision-language understanding and multi-embodiment action generation.
评估表明,RubSE 在终轮和最佳轮设置下均显著优于朴素 self-evolution,refinement 轨迹更稳定,且轨迹级性能上限更高。Evaluations demonstrate that RubSE substantially outperforms na\"ive self-evolution in final-round and best-round settings, achieving more stable refinement trajectories and a higher trajectory-level performance ceiling.
本文变化 handoff 方向、时机与接口,对比保留仓库状态下的全轨迹传输、压缩与轨迹移除,发现偏好接口随方向反转:减少 LC-model 轨迹信息可提升 escalation 质量,而移除 HC-model 轨迹则会降低 downshift 质量。This work varies handoff direction, timing, and interface, comparing full-trajectory transfer, compaction, and trajectory removal while preserving the repository state, and finds that the preferred interface also reverses with direction: reducing LC-model trajectory information improves escalation quality, whereas removing the HC-model trajectory reduces downshift quality.
结果表明,记忆、几何控制以及 rollout-aware 训练为生成连贯故事和持续演化的交互式世界提供了实用基础。Results indicate that memory, geometric control, and rollout-aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds.
Mixed SFT 是一种单阶段监督微调,联合在 no-CoT 和 long-CoT 数据上训练,相比 next-chunk reasoning RL 取得了明显更高的 RLVR 后性能上限,同时训练算力开销减少超过 60 倍。Mixed SFT, a single supervised fine-tuning stage that jointly trains on no-CoT and long-CoT data, achieves a clearly higher post-RLVR performance ceiling than next-chunk reasoning RL while requiring over 60 times less training compute.
本文提出 FHM-Video,一种基于 check-before-score 原则的、由 checklist 驱动的统一数据构建框架,在 FIRM-Video-Bench 上取得最佳综合 MAE,并在三种视频生成器的 Best-of-8 采样中始终获得最高的 VBench Total、Quality 和 Semantic Score。FHM-Video, a unified checklist-driven data construction framework based on a check-before-score principle, is introduced, which achieves the best overall MAE on FIRM-Video-Bench while consistently delivering the highest VBench Total, Quality, and Semantic Scores in Best-of-8 sampling across three video generators.
本文发布一个土耳其语多模态对话数据集,包含未脚本化的双人交互,并提供同步的前向视频、可归属到每位说话人的独立音频通道以及时间对齐的转写文本。A multimodal Turkish conversational dataset of unscripted dyadic interactions, comprising synchronized front-facing video, per-speaker audio channels that allow overlapping speech to be attributed to individual speakers, and time-aligned transcriptions is introduced.
本工作提出一种新颖的混合精度 PTQ 策略,直接最小化整个模型的全局误差传播,而非孤立地处理逐层误差,并开发了一种新颖的联合优化方法,在统一搜索空间中同时学习结构化剪枝决策与混合精度量化策略。This work proposes a novel mixed-precision PTQ strategy that directly minimizes global error propagation across the entire model, rather than isolating layer-wise errors, and develops a novel joint optimization approach that simultaneously learns structural pruning decisions and mixed-precision quantization policies within a unified search space.
本文用前馈 4D 重建奖励替代静态 critic,显式建模场景动态,使连贯运动获得高一致性奖励,并加入对自然 scene-flow 幅值进行奖励同时抑制抖动与非刚性伪影的运动先验。This work replaces the static critic with a feed-forward 4D reconstruction reward that explicitly models scene dynamics, allowing coherent motion to receive high consistency rewards, and adds a motion prior that rewards natural scene-flow magnitude while penalizing jitter and non-rigid artifacts.
PlanSightRAG 是一种 Visual-First 多模态 RAG 框架,直接对图纸图像建立索引并进行推理,集成了 ColNomic-3B 多向量检索、Agentic Planner-Retriever-Auditor-Synthesizer,并以 MaxSim 热力图作为证据链。A Visual-First Multimodal Retrieval-Augmented Generation (RAG) framework called PlanSightRAG, which indexes and reasons directly over plan imagery, integrates a ColNomic-3B multi-vector retrieval, an agentic Planner-Retriever-Auditor-Synthesizer, and MaxSim heatmaps as an evidence trail.
本文提出一种构建 Context-Enhanced MMKG (CEMMKG) 的新框架,能够有效利用上下文信息提升基于 MMKG 的 RAG 性能,并在不同基于 MMKG 的 RAG 方法上的有效性验证了其广泛适用性。A novel framework for constructing a Context-Enhanced MMKG (CEMMKG) is proposed, effective in leveraging contextual information to improve MMKG-based RAG performance and its effectiveness across different MMKG-based RAG methods demonstrates its broad applicability.
本文发布 LibriBrain100,一个面向语音解码的大规模 MEG 数据集,从设计上保证可复现、标准化评估,并展示了广泛多被试数据的价值:对预训练模型进行有监督微调可大幅弥补单被试数据不足。LibriBrain100 is introduced, a large-scale MEG dataset for speech decoding designed from the ground up for reproducible, standardised evaluation and the value of broad multi-subject data is demonstrated: supervised finetuning of a pre-trained model can substantially compensate for limited per-subject data.
本文发布 SWE Refactor Bench,一个包含 20 项全仓库迁移的基准,涵盖 4 类技术债务,作为开发面向可靠全仓库迁移的编码 Agent 的严格测试平台。SWE Refactor Bench is introduced, a benchmark comprising 20 whole-repository migrations, covering 4 kinds of technical debt, and SWE Refactor Bench is positioned as a rigorous testbed for developing coding agents for reliable whole-repository migrations.
本文通过多语言 self-play,正交于知识与综合基准性能对跨语言技能不一致性进行量化,表明技能差异是开发真正多语言模型过程中可衡量且主要的障碍。This work quantifies cross-lingual skill inconsistency orthogonally from knowledge and general benchmark performance via multilingual self-play, and shows that skill discrepancies are a measurable major roadblock in the development of truly multilingual models.
提出 Prefix Sliding,在推理过程中丢弃不属于前缀或最近几千 token 窗口的 token,从而实现高效的长时程 test-time scaling。Prefix Sliding is proposed, which discards tokens during reasoning that are not part of the prefix or the window of the last few thousand tokens, allowing for efficient long-horizon test-time scaling.