实验表明,OpenComputer 的硬编码验证器比 LLM-as-judge 评估更贴合人类裁定,尤其当任务成败取决于细粒度应用状态时。Experiments show that OpenComputer's hard-coded verifiers align more closely with human adjudication than LLM-as-judge evaluation, especially when success depends on fine-grained application state.
论文
144 张论文卡片 · 应用落地
RTP-LLM 是一个面向工业级 LLM 部署的高性能推理引擎,已在 Alibaba Group 成功部署,服务超过 1 亿用户,通过集成设计解决根本性瓶颈。RTP-LLM is presented, a high-performance inference engine for industrial-scale LLM deployment, successfully deployed across Alibaba Group serving over 100 million users, and addresses fundamental bottlenecks through integrated design.
总体而言,分阶段检索的影响往往超过"理想证据存在"本身;本文为专业科学场景下 RAG 系统的部署与诊断提供了实践指导,并为构建更可靠、可控的迭代式检索-推理框架奠定了基础。This is the first controlled, mechanism-level diagnostic evaluation of whether synchronized iterative retrieval and reasoning can surpass even an idealized static upper bound (Gold Context) RAG, and practical guidance for deploying and diagnosing RAG in specialized scientific settings.
本文提出三种协议级原语以填补Model Context Protocol的空白:身份传递、自适应工具预算与结构化错误语义,并提出Structured Error Recovery Framework (SERF),提供机器可读的失败语义以支持确定性的Agent自校正。Three protocol-level primitives are proposed to fill gaps in the Model Context Protocol: identity propagation, adaptive tool budgeting, and structured error semantics, and the Structured Error Recovery Framework (SERF), which provides machine-readable failure semantics that enable deterministic agent self-correction.
该工作提出 TAMP-Nav,一个用于高效具身导航的统一框架,可在关键节点动态触发 Chain-of-Thought 并仅在关键节点保留高保真记忆,将冗余轨迹压缩为轻量级时空指示器,从而保留关键历史信息并增强时空感知。This work proposes TAMP-Nav, a unified framework for efficient embodied navigation that dynamically triggers Chain-of-Thought and retains high-fidelity memory only at critical nodes, compressing redundant trajectories into lightweight Space-Time Indicators, thereby preserving critical historical information and enhancing spatio-temporal perception.
Agent Lightning v1.0 是一个轻量级的可控 Agentic RL 框架,约 3500 行代码实现,支持任意 Agent harness,并作为研究 retokenization、样本合并、优势计算、损失归一化与后端调度等挑战的实用测试平台。Agent Lightning v1.0 is presented, a lightweight framework for harnessed agentic RL implemented in approximately 3,500 lines of code that supports arbitrary agent harnesses and serves as a practical testbed for studying challenges in retokenization, sample merging, advantage calculation, loss normalization, and backend scheduling.
研究表明矛盾解析本质上是写入时并发控制,并将缺失的契约——一个在隔离性、模式与来源维度上被证明正确的写入时正确性规范——显式化,固定了每个生产启发式都默认假设、却没有任何已部署系统显式给出的保证。It is shown that contradiction resolution is write-time concurrency control and make the missing contract explicit, a write-time correctness specification, proved sound across isolation, schema, and provenance, pinning the guarantee every production heuristic assumes but no deployed system makes explicit.
提出首个 MuseCP 评估框架,涵盖四类音乐 facet,使用细粒度且量身定制的指标来捕捉音乐属性的细微变化,并希望为开发更有效、更可靠、具备强大 MuseCP 能力的音乐编辑策略提供实践指导。The first MuseCP evaluation framework is introduced that covers four categories of music facets with fine-grained and well-tailored metrics to capture nuanced changes in music attributes and hopes it can offer practical guidance for developing more effective and reliable music editing strategies with strong MuseCP capability.
本文引入一个均值修正项,有效抑制近似误差,即使在极端稀疏度下也能将性能下降保持在可控范围;并使用 PackGQA 内存访问、warp specialization 和 pingpong 流水线重新设计稀疏注意力算子。This paper introduces a mean correction term that effectively suppresses the approximation error, keeping performance degradation manageable even at extreme sparsity levels, and redesigns the sparse attention operator with PackGQA memory access, warp specialization, and pingpong pipelining.
结果表明,任务特定的 harness 进化是改进冻结 LLM Agent 的一条可行路径,但其效果存在明确的经验性边界。Results demonstrate task-specific harness evolution as a viable axis for improving frozen LLM agents under clear empirical limits.
本文介绍 TianoForge,这是面向 TianoCore 开源 UEFI 固件开发生态中 bug 分诊的集成方案,部署 AI(具体为机器学习)领域的 SOTA 方法以实现自动化 bug 分诊。This integrated approach to bug triage in the TianoCore open-source UEFI firmware development ecosystem, called TianoForge, deploys the state of the art in artificial intelligence, specifically machine learning, to enable automated bug triage.
本文通过全面超越传统开环基线,证明了当前主流的单体上下文扩展策略是一种因相关性衰减而受到惩罚的架构陷阱,并确立了以顺序、反馈驱动的编排作为生成式搜索的确定性范式。By dominating classical open-loop baselines, this work proves that the prevailing strategy of monolithic context widening is an architectural trap penalized by relevance decay and establishes sequential, feedback-driven orchestration as the definitive paradigm for generative search.
目标是提供无需数周超参搜索即可部署的方案,直接从原始未压缩模型蒸馏 4-bit 学生模型,并以开源权重形式发布为 Hypernova-60B。The aim is a recipe deployable without a multi-week hyper-parameter search, which distills the 4-bit student directly from the original, uncompressed model, and is released open-weight as Hypernova-60B.
介绍 PinSieve——大规模内容质量流水线中的生产级案例:一个选择性 vision-language-model Serving Agent,仅处理轻量上游模型无法覆盖的 grey-zone 切片,在线暴露标量路由评分,并保留受控的人工升级通道。This work presents PinSieve, a production case study in a large-scale content-quality pipeline, a selective vision-language-model Serving Agent that operates only on the grey-zone slice left unresolved by lightweight upstream models, exposes a scalar routing score online, and preserves controlled human escalation.
介绍 WeMM-Embedding,一族通用多模态嵌入模型,支持文本、图像、视频、视觉文档及任意交错的多模态输入,输出维度灵活,在多个公开基准上取得 SOTA 表现。WeMM-Embedding is presented, a family of universal multimodal embedding models supporting text, images, videos, visual documents, and arbitrarily interleaved multimodal inputs with flexible output dimensions and achieves leading performance on multiple public benchmarks.
行为拓扑更多由部署 harness 决定而非 LLM 本身,为安全审计和运行时监控提供一种与模型无关的结构化 primitive,并同时满足两类预测目标。Behavioral topology is shaped more by the deployment harness than by the LLM, providing a model-agnostic structural primitive for safety auditing and runtime monitoring, and addresses both prediction goals.
CrabOS 将人机交替主导复杂任务的支持从依赖桥接的应用层方案提升为原生操作系统能力,为开发与运行 AI agent 提供了新基础。CrabOS elevates support for complex tasks with alternating human and AI leadership from bridge-dependent application-level solutions to native operating-system capabilities, which provide a new foundation for developing and running AI agents.
结果揭示了集成组合对精度–召回权衡的直接影响:异构跨范式集成通常提升精度,而同构 LLM 集成更常取得更高的整体 F1。It is revealed that ensemble composition directly affects the precision-recall trade-off: heterogeneous cross-paradigm ensembles generally improve precision, whereas homogeneous LLM ensembles more often achieve higher overall F1-scores.
该工作提出 Super Library Agent 问题:Agent 顺序生成 N 个相关应用组成的组合,同时维护一个共享的 Super Library 用于跨应用可复用组件,并解决了基于候选引导的代码块摘要抽取、抽取前的代码库整合,以及利用抽取 trace 和调用图信息的上下文感知迁移。This work introduces the Super Library Agent problem, where an agent sequentially generates a portfolio of N related applications while maintaining a shared Super Library of reusable cross-application components, and addresses candidate-guided extraction over code chunk summaries, pre-extraction codebase consolidation, and context-aware migration using extraction traces and call-graph information.
本文提出 CASTER,一种无梯度方法:将源类统计量存储在判别子空间中,从目标批次矩估计一个类共享的仿射变换,并在分类前解析地将源类分布迁移到目标域,使其成为面向冻结模型部署的轻量适配机制。CASTER is introduced, a gradient-free method that stores source class statistics in a discriminative subspace, estimates a class-shared affine transformation from target-batch moments, and analytically transports the source class distributions before classification, which positions CASTER as a lightweight adaptation mechanism for frozen-model deployment.
为解决 LLM 与非语言 Agent 的协作问题,本文提出 latent state internalization,将子 Agent 的连续表示直接投射到 LLM 的 token 流中作为习得的状态 token,并随着动作推进环境状态而进行动态重编码。To solve LLM collaboration with non-language agents, latent state internalization is introduced, which projects the subagent's continuous representations directly into the LLM's token stream as learned state tokens, with dynamic re-encoding as actions advance the environment state.
本文提出 Debias-SparseGPT,一种 post-training 剪枝方法,通过在人口统计对比输入上定义的二阶项引入表示去偏好,在保持模型困惑度与零样本准确率的前提下,一致地降低剪枝带来的偏差,效果优于 SparseGPT。Debias-SparseGPT is introduced, a post-training pruning method incorporating representational debiasing using a second-order term defined over demographically contrasting inputs that consistently reduces pruning-induced bias compared to SparseGPT while preserving model perplexity and zero-shot accuracy.
提出 Temporal Context Routing,将脚本时序映射到视频与音频生成的共享时间轴上,并将每个 prompt 的引导路由到两种模态中的对应位置,同时保持与 baseline 相当的视觉质量与音视频同步性。Temporal Context Routing is introduced, which maps the script timing onto the shared temporal axis of video and audio generation and routes each prompt's guidance to the corresponding positions in both modalities, while maintaining visual quality and audio-visual synchronization comparable to those of the baselines.
ShallowStream 是一个利用 MLLM 浅层同时进行帧编码与检索索引构建的新框架,性能与当前最强流式方法相当,同时将单帧 prefill 延迟与 10 秒端到端延迟分别降低至多 52.1 倍和 11.9 倍。ShallowStream is a novel framework that leverages the shallow layers of an MLLM to simultaneously perform frame encoding and retrieval index building and achieves performance on par with the strongest existing streaming methods, while reducing per-frame prefill latency and 10-second end-to-end latency by up to 52.1x and 11.9x, respectively.
EmbodiedSkills 是一个统一框架,将每项技能决策视为执行提案:运行时先检查前置条件再执行,执行后再验证结果,为将底层 VLA 策略转化为闭环具身系统提供可训练且可检视的 Agent 层。EmbodiedSkills is a unified framework that treats each skill decision as an execution proposal: the runtime checks its prerequisites before execution and verifies the outcome afterward and provides a trainable and inspectable agent layer for turning low-level VLA policies into closed-loop embodied systems.
结果表明数据组成控制安全性与可用性的权衡,且安全对齐应在预期拒答边界的两侧进行评估。The results show that data composition controls the safety and usability trade-off, and that safety alignment should be evaluated on both sides of the intended refusal boundary.
本文提出通过让 evaluator 与解决方案协同进化来自动化 evaluator 的设计,并证明突破 evaluation 瓶颈可释放 ADRS 的潜力,为下一代数据系统生成高度优化、可部署的代码。This work proposes automating the design of evaluators by co-evolving them with the solutions, demonstrating that addressing the evaluation bottleneck unlocks the potential of ADRS to generate highly optimized, deployable code for next-generation data systems.
Marigold V2 在应用于其他稠密回归任务(如表面法向量估计与本征图像分解)时取得 SOTA 结果。Marigold V2 achieves state-of-the-art results when applied to other dense regression tasks, such as surface normals estimation and intrinsic image decomposition, and achieves state-of-the-art results when applied to other dense regression tasks, such as surface normals estimation and intrinsic image decomposition.
本文介绍 SCHEMEARENA,一个面向可扩展谋划行为压力测试的 400 场景基准,通过覆盖多种安全相关工具领域、工具性目标、监管条件与压力机制的因子化场景合成框架构建。This work introduces SCHEMEARENA, a 400-scenario benchmark for scalable scheming stress testing, constructed through a factorized scenario synthesis framework spanning diverse safety-relevant tool domains, instrumental goals, oversight conditions, and pressure mechanisms.
本文规定 EBL-Core,一个执行边界合规配置文件,用于判定一个规范且完全实例化的 AI 生成候选对象在明确条件下是否可获得操作范围的执行权限。EBL-Core, an execution-boundary conformance profile for deciding whether one canonical, fully materialized AI-generated candidate may receive action-scoped execution authority under explicit conditions, is specified.
EvoSafeHarness 是一个面向安全性的优化框架,针对目标领域中的冻结模型合成可部署的 harness,由模型行为、领域规范和新鲜上下文对抗审查共同引导,以拒绝基准特定规则。EvoSafeHarness is a safety-specific optimization framework that synthesizes a deployable harness for a frozen model in a target domain, guided by model behavior, domain specifications, and fresh-context adversarial review to reject benchmark-specific rules.
最终得到的 82M 参数模型 Wayu-Paxa-TTS-Edge 实现了无需参考音频的设备端泰语 TTS,并在三个系统中取得了最低的停顿位置错误率与词内停顿率。The resulting 82M-parameter model, Wayu-Paxa-TTS-Edge, enables on-device Thai TTS without reference audio and achieves the lowest pause-placement error and intra-word pause rates among the three systems.
本文探讨 multi-agent system,并指出当前尚未被充分解决的问题,同时探索了 multi-agent system 在区块链系统中的潜在应用,为其在真实分布式系统中的未来发展与落地提供启示。This paper explores multi-agent systems and identifies challenges that remain inadequately addressed, and explores potential applications of multi-agent systems in blockchain systems to shed light on their future development and application in real-world distributed systems.
提出 Continual Search,一个迭代框架:在多轮对话中持续推动判别器搜索尚未解决的诊断证据,在多个基准测试套件和模型系列上一致提升归因性能Continual Search is introduced, an iterative framework that nudges the judge, over successive turns, to keep searching for unresolved diagnostic evidence, which consistently improves attribution performance across multiple benchmark suites and model families.