识别出 diff 生成能够胜出的一种与架构无关的统一机制:它在短小且空间局部化的编辑上具有竞争力;其类别级优势恰好集中在作者数据集中平均编辑步数最低的两类任务——重构与错误处理/边界用例修复。A single, architecture-independent mechanism behind the conditions where diff-based generation does win is identified: it is competitive on short, spatially localized edits, and its category-level wins concentrate in exactly the two task categories - refactoring and error-handling/edge-case fixes - with the lowest mean edit-step count in the authors' dataset.
论文
1753 张论文卡片
WearableQA 由 200 位真实用户的可穿戴时序数据、血液生物标志物和人口统计信息构建的 4,084 道十选一选择题组成,为评估 LLM 在真实可穿戴数据上的推理能力提供了现实且具诊断性的基准。WearableQA, a benchmark comprising 4,084 10-option multiple-choice questions constructed from the wearable time series, blood biomarkers, and demographics of 200 real users, provides a realistic and diagnostic benchmark for evaluating LLM reasoning over real-world wearable data.
揭示了复用造成的质量损失以及有效修复方法均取决于 LLM(即便在两个 8B 模型之间也是如此),并提供添加新方法的通用接口与交互式 leaderboardBoth the quality loss from reuse and which repairs help depend on the LLM, even between two 8B models are introduced, as well as a common interface for adding new methods and an interactive leaderboard.
本综述涵盖视觉与音视频 VideoLLM 的推理效率机制,这些机制报告了参数数量、每输入 FLOPs、延迟、内存或视觉/音频 token 数量的具体削减量,并按方法所作用的 pipeline 阶段加以组织。This survey covers inference-efficiency mechanisms for visual and audiovisual VideoLLMs that report concrete reductions in parameter count, FLOPs per input, latency, memory, or visual and audio token count, and organize methods by the pipeline stage at which they act.
本文介绍 Brain2Semantics2Text,一种通过中间语义嵌入空间重建文本的方法,并阐述了该方法的核心原理、实现方式以及缓解学习可靠神经-语义映射挑战的策略。This work introduces Brain2Semantics2Text, a method that reconstructs text through an intermediate semantic embedding space and describes the core principles of the approach, its implementation, and the strategies used to mitigate the challenges of learning a reliable neural-to-semantic mapping.
提出 AgentGrad,一种基于序贯干预与语义文本梯度抽象的多智能体系统 prompt 优化框架,在 5 个 MAS benchmark 上取得 SOTA 性能,同时降低优化耗时与成本AgentGrad is proposed, a prompt optimization framework for multi-agent systems based on sequential intervention and semantic textual gradient abstraction that achieves state-of-the-art performance across five MAS benchmarks while reducing wall-clock optimization time and optimization cost.
SWE-Bench Pro Verified 提供了一个更可信的基准来评估软件工程 Agent,其结合了反作弊保护(消除主要泄漏渠道而不干扰正常 Agent 功能)和任务精修(最小限度地修正有缺陷实例中的不一致性)。SWE-Bench Pro Verified offers a more trustworthy benchmark for assessing software engineering agents, which combines anti-hacking safeguards that eliminate major leakage channels without disrupting normal agent functionality, with task refinement that minimally corrects inconsistencies within flawed instances.
本文介绍 StochBench,一个基于 Lean 4 的基准,包含 450 道覆盖不同抽象层次的研究生随机过程问题,每道题配有其自然语言来源,更能代表领域特定的应用数学,同时对高级证明器仍具挑战。StochBench is introduced, a Lean 4 benchmark of 450 graduate stochastic-processes problems at varying abstraction levels, each paired with its natural-language source that better represents domain-specific applied mathematics while remaining challenging for advanced provers.
这些结果表明,Kueue、DAS 与 GAIE 等互补组件构成了一个高性能的协同平台,证明了 Kubernetes 能够作为承载高要求 GenAI 工作负载的统一底座。These findings illustrate that these complementary components (Kueue, DAS, and GAIE) form a cohesive, high-performance platform, proving Kubernetes' capability to serve as a unified foundation for demanding GenAI workloads.
本文介绍 SCHEMEARENA,一个面向可扩展谋划行为压力测试的 400 场景基准,通过覆盖多种安全相关工具领域、工具性目标、监管条件与压力机制的因子化场景合成框架构建。This work introduces SCHEMEARENA, a 400-scenario benchmark for scalable scheming stress testing, constructed through a factorized scenario synthesis framework spanning diverse safety-relevant tool domains, instrumental goals, oversight conditions, and pressure mechanisms.
证明:对于任意固定的目标误差水平 $\delta$ 与任意松弛量 $\varepsilon>0$,数量级为 $p/\psi^2$ 的样本量足以在任意小的 $\psi$ 下完成支撑恢复。It is proved that, for every fixed target error level $\delta$ and every slack $\varepsilon>0$, a sample size of order $p/\psi^2$ is sufficient for support recovery for arbitrarily small $\psi$.
本文介绍 PARSER,将阅读与推理解耦,对证据位置、顺序与距离的扰动具有鲁棒性——这些条件会导致序列方法产生大幅精度波动——同时将推理延迟降低多达 11 倍。PARSER, which decouples reading from reasoning, is introduced, which is robust to perturbations in evidence position, order, and distance, conditions that cause large accuracy swings in sequential methods, while reducing inference latency by up to 11x.
本文引入影响引导的响应改写方法,利用 IF 识别干预目标,在保持指令不变的情况下将其响应替换为行为对齐或行为对立的监督信号,以此推动对 TDA 方法的干预感知评估。Influence-guided response rewriting is introduced, which uses IF to identify intervention targets and replaces their responses with behavior-aligned or behavior-opposed supervision while keeping instructions fixed, motivating intervention-aware evaluation of TDA methods.
本文规定 EBL-Core,一个执行边界合规配置文件,用于判定一个规范且完全实例化的 AI 生成候选对象在明确条件下是否可获得操作范围的执行权限。EBL-Core, an execution-boundary conformance profile for deciding whether one canonical, fully materialized AI-generated candidate may receive action-scoped execution authority under explicit conditions, is specified.
本文介绍 AgentZip,第一个专为 AI Agent 沙箱设计的内存压缩系统,将压缩范围扩展到任何具有收益表示的页面,并将开销控制从压缩时页面选择转移到恢复时预取。AgentZip is presented, the first memory compression system designed specifically for AI-agent sandboxes, which broadens the compression scope to any page with a profitable representation and shifts overhead control from compression-time page selection to restore-time prefetching.
在真实的博卡拉湖畔地理环境中,由 100 个配备记忆机制的大语言模型 Agent 管理一个封闭且守恒的空间经济,并运行该多 Agent 模拟长达 26 个模拟周,远超典型 Agent 社会研究 1–2 周的时长。100 memory-equipped large language model agents in charge of a closed, money-conserving spatial economy on real Pokhara Lakeside geography and ran this multi-agent simulation for up to 26 simulated weeks, well past the 1-2 weeks typical of agent-society studies.
提出 X-AuT,一个渐进式框架,通过简短的行为探针选择层组合,并通过表征对齐、跨尺度蒸馏、调度式学生策略监督以及 LoRA 微调来恢复被剪枝的模型。X-AuT, a progressive framework that selects layer combinations through short behavioral probes and restores the pruned model through representation alignment, cross-scale distillation, scheduled student-policy supervision, and LoRA finetuning, is introduced.
Generative Late-Interaction Embeddings(GLIE):从归一化质心中学习每个页面 k<<N 个向量,既作为轻量级索引,也作为重建页面完整嵌入集的基础,解码器是其主要设计面。Generative Late-Interaction Embeddings (GLIE): k<<N vectors per page learned from the normalized centroids to serve as both a lightweight index and a basis for regenerating the page's full embedding set, with the decoder as its main design surface.
提出一个用于困难奥林匹克数学自然语言证明生成的开放模型测试时计算流水线,完全在自然语言中运行,无需形式化证明器、外部工具或互联网访问。An open-model test-time-compute pipeline for natural-language proof generation for hard olympiad mathematics that operates entirely in natural language, with no formal prover, external tools, or internet access is presented.
本文提出了一种简单的 flow-control 框架,通过控制 prompt 加入 LLM 活跃集合的速率,实现更高的 token 与 request 吞吐量、更低的平均与尾部 latency,以及更稳定的 KV cache 利用率。A simple flow-control framework is proposed that controls the rate at which prompts join the active set in large language models and achieves higher token and request throughput, lower average and tail latency, and more stable KV cache utilization.
引入 SpatialBlock-15k,一个包含 15,000 个堆块问题的合成数据集,涵盖 3D 到 2D 投影、视角变换和结构组合,并提出受人类认知发展启发的新范式:通过结构化堆块操作任务学习基础空间技能。This work introduces SpatialBlock-15k, a synthetic dataset of 15,000 block-stacking problems covering 3D-to-2D projection, viewpoint transformation, and structural combination and proposes a novel paradigm inspired by human cognitive development: learning foundational spatial skills through structured block-manipulation tasks.
EvoSafeHarness 是一个面向安全性的优化框架,针对目标领域中的冻结模型合成可部署的 harness,由模型行为、领域规范和新鲜上下文对抗审查共同引导,以拒绝基准特定规则。EvoSafeHarness is a safety-specific optimization framework that synthesizes a deployable harness for a frozen model in a target domain, guided by model behavior, domain specifications, and fresh-context adversarial review to reject benchmark-specific rules.
提出 MaP-WAM,一个将记忆作为规划的 Memory-as-Plans 框架,将依赖记忆的世界动作建模分解为基于记忆的规划和以规划为条件的执行,并以长期多模态情景上下文作为规划时证据,而非反复对执行器输入完整历史。MaP-WAM is introduced, a Memory-as-Plans framework that decomposes memory-dependent world-action modeling into memory-grounded planning and plan-conditioned execution, and uses long-term multimodal episodic context as planning-time evidence rather than repeatedly conditioning the executor on the full history.
提出 Negative Self-Distillation(NSD),一个通过偏离有缺陷的推理而非模仿特权解来优化 LLM 的新框架,并一致优于 OPSD 及其他无标签、自举式强化学习(RL)基线。Negative Self-Distillation (NSD) is introduced, a new framework that optimizes LLMs by diverging from flawed reasoning rather than imitating privileged solutions, and consistently outperforms OPSD and other label-free, self-bootstrapping reinforcement learning (RL) baselines.
实验结果表明 DRG-MAPPO 达到了 87% 的 SOTA 胜率,表明该框架在合作空战中有效平衡了关系建模、可解释性和优化稳定性。Experimental results demonstrate that DRG-MAPPO achieves a state-of-the-art win rate of 87%, suggesting that the framework effectively balances relational modeling, interpretability, and optimization stability for cooperative air combat.
提出 UniH3,一个统一层次同质性与异质性的全新框架,用于一体化医学图像修复,在两个大规模基准上的大量实验表明 UniH3 在一体化和单任务医学图像修复上均达到 SOTA 性能。UniH3, a novel framework that Unifies Hierarchical Homogeneity and Heterogeneity and Heterogeneity for all-in-one medical image restoration, is proposed and extensive experiments on two large-scale benchmarks demonstrate that UniH3 achieves state-of-the-art performance on both all-in-one and single-task medical image restoration.
引入 MetroLLM-Bench,一个包含 955 个用例的基准,用于测试语言模型作为交通信息亭策略层的能力,并评估了来自六个厂商的 26 个模型,其中 23 个被排名。This work introduces MetroLLM-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk, and evaluates twenty-six models from six vendors, of which twenty-three are ranked.
Video-LLM 的时间推理基准常以语言为中介,留有来自选项措辞、答案相关性或语言先验带来的语言捷径空间,因此提出 TempCloze,一个用于评估 Video-LLM 视觉时间推理能力的视频完形填空基准。Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors, so TempCloze is introduced, a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs.
提出 Generative Verification(GenV),通过复用语言模型的原生词表空间,将离线 Z3 等价性预言机蒸馏为无参考、连续参考等价分数,并从理论上证明基于结构、仅判决的验证启发式在这些欺骗性合法轨迹上的检测能力在数学上有界于随机水平。Generative Verification (GenV) is introduced, which distills an offline Z3-equivalence oracle into a reference-free, continuous reference-equivalence score by repurposing the language model's native vocabulary space and theoretically proves that structural, verdict-only verification heuristics are mathematically bounded to chance-level detection on these deceptively valid traces.
一个简单、无训练的框架,其中具备推理能力的视觉语言模型在 Wikipedia 上迭代搜索与推理,动态收集证据,表明推理与检索在稀有实体上具有互补性。A simple, training-free framework in which a reasoning-capable vision-language model iteratively searches and reasons over Wikipedia, gathering evidence dynamically, shows that reasoning and retrieval are complementary on rare entities.
本文提出了 SeedRG,一个用于缓解 knowledge leakage 并应对 benchmark aging 问题的半合成 benchmark 生成 pipeline。SeedRG is introduced, a semi-synthetic benchmark generation pipeline that mitigates knowledge leakage and addresses the issue of benchmark aging.
表明将 L2 推理泛化到未见语言的关键路径在于更广泛的语言覆盖、现成可用的多语言非推理数据,以及足够强的英语推理骨干,表明推理是一种与语言无关的行为,可通过精心数据混合在类型多样的语言间迁移。It is shown the path to generalizing L2 reasoning to held-out languages goes through broader language coverage, readily available multilingual non-reasoning data, and a sufficient English reasoning backbone, indicating that reasoning is a language-agnostic behavior that can be transferred across typologically diverse languages through careful data mixing.
研究面向实现的研究方法规范的可编码化就绪度,定义为其是否为胜任的实现者或编码 Agent 提供足够的方法学信息,以在不引入未支持假设的情况下构建预期方法。This work studies the codification readiness of implementation-facing research-method specifications, defined by whether they provide sufficient methodological information for a competent implementer or coding agent to construct the intended method without unsupported assumptions.
构建一个受控的纯自回归测试床,在文本、图像、文生图(T2I)和图生文(I2T)预测的多模态持续预训练中跟踪任务特定验证损失,表明更好的重建并不一定带来更低的任务特定损失或更强的下游性能,且图像分词器的选择在联合优化下会影响文本建模。A controlled pure-autoregressive testbed is built and track task-specific validation losses during multimodal continual pretraining across text, image, text-to-image (T2I), and image-to-text (I2T) prediction, showing that better reconstruction does not necessarily yield lower task-specific losses or stronger downstream performance, and that image tokenizer choice can affect text modeling under joint optimization.
引入 ActReview,一个由反驳引导的后训练框架,将论文特定的诊断连接到具体、有依据的修订计划,并引入 ActReview-Bench,一个包含 1,000 个实例的人工整理基准,用于评估诊断质量和修订实用性。This work introduces ActReview, a rebuttal-guided post-training framework that connects paper-specific diagnoses to concrete, grounded revision plans, and introduces ActReview-Bench, a human-curated benchmark of 1,000 instances for evaluating diagnostic quality and revision usefulness.
结果表明,在评估 harness 中使用 Adaptive Bridge 可将关键的订阅者尾端 p95 延迟从最高 15 s 降至 1.55 ms,覆盖所有损伤严重度,同时保留发布者配置的吞吐量。The results show that using the Adaptive Bridge in the evaluation harness reduces the critical subscriber tail p95 latency from up to 15 s to 1.55 ms across all impairment severities while preserving the publisher's configured throughput.