本文为 LLM serving 中的 SD(投机解码)提出了一种简单且可解释的 latency 模型,能够准确刻画实际观测到的 latency,解释为何加速比常随服务器负载上升而下降,并系统刻画了 draft length、acceptance rate 以及 verifier 与 drafter 规模在不同 serving 条件下对 latency 的影响。A simple and interpretable latency model for SD in LLM serving is developed that accurately describes observed latency, explains why speedups often diminish as server load increases, and characterizes how draft length, acceptance rate, and verifier-drafter size shape latency across serving conditions are characterized.
论文
1640 张论文卡片 · OA 绿色
VDiff-Bench 提供一个针对性诊断基准,用于评估 MLLM 的比较视觉理解能力,揭示标准单图视觉-语言任务无法捕捉的失败模式。VDiff-Bench provides a targeted diagnostic for evaluating comparative visual understanding in MLLMs, exposing failures that are not captured by standard single-image vision-language tasks.
主张防御的运行单元应是可修订的协同 episodes,将观察到的迁移、任务权限与响应历史关联起来,并提出跨执行监控建议具有可测试性,但并不声称提出新的检测器或测得具体的遏制收益。It is argued that the operational unit of defence should be a revisable coordination episode linking observed transfers, task authority, and response history, and it makes the recommendation to monitor across executions testable without claiming a new detector or a measured containment benefit.
我们提出 RenderFormer-V2,一个统一的基于 transformer 的学习型神经渲染模型,可与现代基于物理的渲染系统互补,无需逐场景训练或专用代码,即可处理焦散、体积散射、环境光照、带纹理与置换的表面以及分布外材质等多种光传输效果。RenderFormer-V2 将全局光传输建模为序列到序列变换。沿袭前作,它仍采用两阶段流程:先是与视图无关的阶段,解析场景内基元到……We present 'RenderFormer-V2', a unified learned transformer-based neural rendering model, complementary to modern physics-based rendering systems, that can handle diverse light-transport effects such as caustics, volumetric scattering, environment lighting, textured and displaced surfaces and out-of-distribution materials without per-scene training or specialized code. RenderFormer-V2 models global light transport as a sequence-to-sequence transformation. Following its predecessor, RenderFormer-V2 also employs a two stage process: a view-independent stage that resolves intra-scene primitive to
提出 EvoHarnessBench,一个在工具、技能与 Agent 三个维度上对可控 harness 演化条件下的 Agent 进行评测的 benchmark,并将 harness 演化确立为一项独立挑战:Agent 需要在持续演化的 harness 下保持原有有效行为EvoHarnessBench is introduced, a benchmark for evaluating agents under controlled harness evolution across three axes (tools, skills, and agents), and establishes harness evolution as a distinct challenge for building agents that can keep pace with an evolving harness while preserving previously effective behavior.
提出 Graph Machine,一种保持 O(n) 规模状态并通过稀疏动态路由访问的架构,使用边——由类似指针追逐的引用机制以可微分方式更新的指针类对象。The Graph Machine is introduced, an architecture that maintains an O(n)-sized state and accesses it through sparse, dynamic routing and uses edges - pointer-like objects updated differentiably by a referral mechanism resembling pointer chasing.
我们研究一种受治理的企业分析方法:语言模型负责解读问题,确定性 policy 负责选取并运行预先批准的分析程序,返回结果与证据。我们证明,在限定的分析类内(包含关系运算,以及聚合、比较、窗口、排序和相似度),这种限制仍可保持表达力。固定的语义、policy、数据和执行规则也使结果可复现。在 440 次运行中,三个 8B 模型生成 SQL 并在运行时选取工具,而 Qwen3-8B 仅解读意图,policyWe study a governed approach to enterprise analytics: a language model interprets the question, while deterministic policy selects and runs a pre-approved analytical program that returns both results and evidence. We show that this restriction can remain expressive within a defined analytical class, using relational operations plus aggregation, comparison, windows, ranking, and similarity. Fixed meaning, policy, data, and execution rules also make results replayable. Across 440 runs, three 8B models generated SQL and selected tools at runtime, while Qwen3-8B interpreted intent only and policy
AgentAudit 沿能力、接地性、安全与行为四个方面的十个维度评估完整执行轨迹——即指令完整性、规划器、记忆、工具选择、工具调用、工具正确性、对齐、工具忠实性、安全性与执行完整性——以精确定位导致观察到的失败的具体阶段。AgentAudit evaluates the entire execution trace across ten capability, grounding, security and behavioural dimensions, namely instruction integrity, planner, memory, tool selection, tool invocation, tool correctness, alignment, tool faithfulness, security and execution integrity, to pinpoint the exact stage responsible for an observed failure.
Glyph 是一个生产系统,将列描述生成与列类型标注这两个耦合问题建模为协同工作的 LLM Agent,并以有状态图形式编排,使多 Agent LLM 目录编制可审计且可作为生产服务运行。Glyph, a production system that frames two coupled problems, column description generation and column type annotation for data classification, as cooperating LLM agents orchestrated as stateful graphs, makes multi-agent LLM cataloging auditable and operable as a production service.
评估 AI Agent 能否作为科学家利用 SAE 工具开展自主机理发现,旨在将实验性模型理解确立为可测量的能力,推动闭环自主 AI R&D。Whether AI agents can act as scientists utilizing SAE tools for autonomous mechanistic discovery is evaluated to establish experimental model understanding as a measurable capability for closed-loop autonomous AI R&D.
开发一条 on-policy 专家修正流水线,由元层级 MLE Agent 自动化,在弱模型自身的 rollout 中定位失败回合,并请专家仅重写该回合,从而保留模型的规划风格,融合 Harness 进化与模型适配带来的收益。An on-policy expert-correction pipeline is developed, automated by a meta-level MLE agent, that localizes the failing turn in the weaker model's own rollout and asks the expert to rewrite only that turn, which preserves the model's planning style and combines the gains of harness evolution and model adaptation.
Tutti 是一种高效的 SSD-backed KV caching 方案,将 CPU 从 HBM 与 SSD 之间的关键数据与 I/O 控制路径中彻底移除,在提供近乎无限容量的同时,实现了与 DRAM-backed LMCache 几乎相当的 inference 性能。Tutti is an efficient SSD-backed KV caching solution that eliminates CPU intervention from the critical data and I/O control paths between HBM and SSDs, and achieves nearly the same inference performance as DRAM-backed LMCache, while providing almost infinite capacity.
Discovery Certification Protocol(DCP)将结果声明转化为在已注册模型、信息边界与预算下的可执行审计,并由确定性验证器基于冻结记录复现本地决策The Discovery Certification Protocol (DCP) turns an outcome claim into an executable audit under a registered model, information boundary, and budget, and a deterministic verifier reproduces these local decisions from frozen records.
提出 DianShi-RxnDB,一个通过全自动抽取与归一化流水线(整合专利文本、图像和反应路线图)构建的大规模细粒度有机反应数据平台。D DianShi-RxnDB is presented, a large-scale, fine-grained organic reaction data platform built via a fully automated extraction and normalization pipeline integrating patent text, images, and reaction schemes integrating patent text, images, and reaction schemes.
识别出 diff 生成能够胜出的一种与架构无关的统一机制:它在短小且空间局部化的编辑上具有竞争力;其类别级优势恰好集中在作者数据集中平均编辑步数最低的两类任务——重构与错误处理/边界用例修复。A single, architecture-independent mechanism behind the conditions where diff-based generation does win is identified: it is competitive on short, spatially localized edits, and its category-level wins concentrate in exactly the two task categories - refactoring and error-handling/edge-case fixes - with the lowest mean edit-step count in the authors' dataset.
WearableQA 由 200 位真实用户的可穿戴时序数据、血液生物标志物和人口统计信息构建的 4,084 道十选一选择题组成,为评估 LLM 在真实可穿戴数据上的推理能力提供了现实且具诊断性的基准。WearableQA, a benchmark comprising 4,084 10-option multiple-choice questions constructed from the wearable time series, blood biomarkers, and demographics of 200 real users, provides a realistic and diagnostic benchmark for evaluating LLM reasoning over real-world wearable data.
揭示了复用造成的质量损失以及有效修复方法均取决于 LLM(即便在两个 8B 模型之间也是如此),并提供添加新方法的通用接口与交互式 leaderboardBoth the quality loss from reuse and which repairs help depend on the LLM, even between two 8B models are introduced, as well as a common interface for adding new methods and an interactive leaderboard.
本综述涵盖视觉与音视频 VideoLLM 的推理效率机制,这些机制报告了参数数量、每输入 FLOPs、延迟、内存或视觉/音频 token 数量的具体削减量,并按方法所作用的 pipeline 阶段加以组织。This survey covers inference-efficiency mechanisms for visual and audiovisual VideoLLMs that report concrete reductions in parameter count, FLOPs per input, latency, memory, or visual and audio token count, and organize methods by the pipeline stage at which they act.
本文介绍 Brain2Semantics2Text,一种通过中间语义嵌入空间重建文本的方法,并阐述了该方法的核心原理、实现方式以及缓解学习可靠神经-语义映射挑战的策略。This work introduces Brain2Semantics2Text, a method that reconstructs text through an intermediate semantic embedding space and describes the core principles of the approach, its implementation, and the strategies used to mitigate the challenges of learning a reliable neural-to-semantic mapping.
提出 AgentGrad,一种基于序贯干预与语义文本梯度抽象的多智能体系统 prompt 优化框架,在 5 个 MAS benchmark 上取得 SOTA 性能,同时降低优化耗时与成本AgentGrad is proposed, a prompt optimization framework for multi-agent systems based on sequential intervention and semantic textual gradient abstraction that achieves state-of-the-art performance across five MAS benchmarks while reducing wall-clock optimization time and optimization cost.
SWE-Bench Pro Verified 提供了一个更可信的基准来评估软件工程 Agent,其结合了反作弊保护(消除主要泄漏渠道而不干扰正常 Agent 功能)和任务精修(最小限度地修正有缺陷实例中的不一致性)。SWE-Bench Pro Verified offers a more trustworthy benchmark for assessing software engineering agents, which combines anti-hacking safeguards that eliminate major leakage channels without disrupting normal agent functionality, with task refinement that minimally corrects inconsistencies within flawed instances.
本文介绍 StochBench,一个基于 Lean 4 的基准,包含 450 道覆盖不同抽象层次的研究生随机过程问题,每道题配有其自然语言来源,更能代表领域特定的应用数学,同时对高级证明器仍具挑战。StochBench is introduced, a Lean 4 benchmark of 450 graduate stochastic-processes problems at varying abstraction levels, each paired with its natural-language source that better represents domain-specific applied mathematics while remaining challenging for advanced provers.
这些结果表明,Kueue、DAS 与 GAIE 等互补组件构成了一个高性能的协同平台,证明了 Kubernetes 能够作为承载高要求 GenAI 工作负载的统一底座。These findings illustrate that these complementary components (Kueue, DAS, and GAIE) form a cohesive, high-performance platform, proving Kubernetes' capability to serve as a unified foundation for demanding GenAI workloads.
本文介绍 SCHEMEARENA,一个面向可扩展谋划行为压力测试的 400 场景基准,通过覆盖多种安全相关工具领域、工具性目标、监管条件与压力机制的因子化场景合成框架构建。This work introduces SCHEMEARENA, a 400-scenario benchmark for scalable scheming stress testing, constructed through a factorized scenario synthesis framework spanning diverse safety-relevant tool domains, instrumental goals, oversight conditions, and pressure mechanisms.
本文介绍 PARSER,将阅读与推理解耦,对证据位置、顺序与距离的扰动具有鲁棒性——这些条件会导致序列方法产生大幅精度波动——同时将推理延迟降低多达 11 倍。PARSER, which decouples reading from reasoning, is introduced, which is robust to perturbations in evidence position, order, and distance, conditions that cause large accuracy swings in sequential methods, while reducing inference latency by up to 11x.
本文引入影响引导的响应改写方法,利用 IF 识别干预目标,在保持指令不变的情况下将其响应替换为行为对齐或行为对立的监督信号,以此推动对 TDA 方法的干预感知评估。Influence-guided response rewriting is introduced, which uses IF to identify intervention targets and replaces their responses with behavior-aligned or behavior-opposed supervision while keeping instructions fixed, motivating intervention-aware evaluation of TDA methods.
本文规定 EBL-Core,一个执行边界合规配置文件,用于判定一个规范且完全实例化的 AI 生成候选对象在明确条件下是否可获得操作范围的执行权限。EBL-Core, an execution-boundary conformance profile for deciding whether one canonical, fully materialized AI-generated candidate may receive action-scoped execution authority under explicit conditions, is specified.
本文介绍 AgentZip,第一个专为 AI Agent 沙箱设计的内存压缩系统,将压缩范围扩展到任何具有收益表示的页面,并将开销控制从压缩时页面选择转移到恢复时预取。AgentZip is presented, the first memory compression system designed specifically for AI-agent sandboxes, which broadens the compression scope to any page with a profitable representation and shifts overhead control from compression-time page selection to restore-time prefetching.
在真实的博卡拉湖畔地理环境中,由 100 个配备记忆机制的大语言模型 Agent 管理一个封闭且守恒的空间经济,并运行该多 Agent 模拟长达 26 个模拟周,远超典型 Agent 社会研究 1–2 周的时长。100 memory-equipped large language model agents in charge of a closed, money-conserving spatial economy on real Pokhara Lakeside geography and ran this multi-agent simulation for up to 26 simulated weeks, well past the 1-2 weeks typical of agent-society studies.
提出 X-AuT,一个渐进式框架,通过简短的行为探针选择层组合,并通过表征对齐、跨尺度蒸馏、调度式学生策略监督以及 LoRA 微调来恢复被剪枝的模型。X-AuT, a progressive framework that selects layer combinations through short behavioral probes and restores the pruned model through representation alignment, cross-scale distillation, scheduled student-policy supervision, and LoRA finetuning, is introduced.
Generative Late-Interaction Embeddings(GLIE):从归一化质心中学习每个页面 k<<N 个向量,既作为轻量级索引,也作为重建页面完整嵌入集的基础,解码器是其主要设计面。Generative Late-Interaction Embeddings (GLIE): k<<N vectors per page learned from the normalized centroids to serve as both a lightweight index and a basis for regenerating the page's full embedding set, with the decoder as its main design surface.
提出一个用于困难奥林匹克数学自然语言证明生成的开放模型测试时计算流水线,完全在自然语言中运行,无需形式化证明器、外部工具或互联网访问。An open-model test-time-compute pipeline for natural-language proof generation for hard olympiad mathematics that operates entirely in natural language, with no formal prover, external tools, or internet access is presented.
本文提出了一种简单的 flow-control 框架,通过控制 prompt 加入 LLM 活跃集合的速率,实现更高的 token 与 request 吞吐量、更低的平均与尾部 latency,以及更稳定的 KV cache 利用率。A simple flow-control framework is proposed that controls the rate at which prompts join the active set in large language models and achieves higher token and request throughput, lower average and tail latency, and more stable KV cache utilization.
引入 SpatialBlock-15k,一个包含 15,000 个堆块问题的合成数据集,涵盖 3D 到 2D 投影、视角变换和结构组合,并提出受人类认知发展启发的新范式:通过结构化堆块操作任务学习基础空间技能。This work introduces SpatialBlock-15k, a synthetic dataset of 15,000 block-stacking problems covering 3D-to-2D projection, viewpoint transformation, and structural combination and proposes a novel paradigm inspired by human cognitive development: learning foundational spatial skills through structured block-manipulation tasks.
EvoSafeHarness 是一个面向安全性的优化框架,针对目标领域中的冻结模型合成可部署的 harness,由模型行为、领域规范和新鲜上下文对抗审查共同引导,以拒绝基准特定规则。EvoSafeHarness is a safety-specific optimization framework that synthesizes a deployable harness for a frozen model in a target domain, guided by model behavior, domain specifications, and fresh-context adversarial review to reject benchmark-specific rules.
提出 MaP-WAM,一个将记忆作为规划的 Memory-as-Plans 框架,将依赖记忆的世界动作建模分解为基于记忆的规划和以规划为条件的执行,并以长期多模态情景上下文作为规划时证据,而非反复对执行器输入完整历史。MaP-WAM is introduced, a Memory-as-Plans framework that decomposes memory-dependent world-action modeling into memory-grounded planning and plan-conditioned execution, and uses long-term multimodal episodic context as planning-time evidence rather than repeatedly conditioning the executor on the full history.