Discovery Certification Protocol(DCP)将结果声明转化为在已注册模型、信息边界与预算下的可执行审计,并由确定性验证器基于冻结记录复现本地决策The Discovery Certification Protocol (DCP) turns an outcome claim into an executable audit under a registered model, information boundary, and budget, and a deterministic verifier reproduces these local decisions from frozen records.
论文
1173 张论文卡片 · 方法
提出 DianShi-RxnDB,一个通过全自动抽取与归一化流水线(整合专利文本、图像和反应路线图)构建的大规模细粒度有机反应数据平台。D DianShi-RxnDB is presented, a large-scale, fine-grained organic reaction data platform built via a fully automated extraction and normalization pipeline integrating patent text, images, and reaction schemes integrating patent text, images, and reaction schemes.
本文介绍 Brain2Semantics2Text,一种通过中间语义嵌入空间重建文本的方法,并阐述了该方法的核心原理、实现方式以及缓解学习可靠神经-语义映射挑战的策略。This work introduces Brain2Semantics2Text, a method that reconstructs text through an intermediate semantic embedding space and describes the core principles of the approach, its implementation, and the strategies used to mitigate the challenges of learning a reliable neural-to-semantic mapping.
提出 AgentGrad,一种基于序贯干预与语义文本梯度抽象的多智能体系统 prompt 优化框架,在 5 个 MAS benchmark 上取得 SOTA 性能,同时降低优化耗时与成本AgentGrad is proposed, a prompt optimization framework for multi-agent systems based on sequential intervention and semantic textual gradient abstraction that achieves state-of-the-art performance across five MAS benchmarks while reducing wall-clock optimization time and optimization cost.
这些结果表明,Kueue、DAS 与 GAIE 等互补组件构成了一个高性能的协同平台,证明了 Kubernetes 能够作为承载高要求 GenAI 工作负载的统一底座。These findings illustrate that these complementary components (Kueue, DAS, and GAIE) form a cohesive, high-performance platform, proving Kubernetes' capability to serve as a unified foundation for demanding GenAI workloads.
证明:对于任意固定的目标误差水平 $\delta$ 与任意松弛量 $\varepsilon>0$,数量级为 $p/\psi^2$ 的样本量足以在任意小的 $\psi$ 下完成支撑恢复。It is proved that, for every fixed target error level $\delta$ and every slack $\varepsilon>0$, a sample size of order $p/\psi^2$ is sufficient for support recovery for arbitrarily small $\psi$.
本文介绍 PARSER,将阅读与推理解耦,对证据位置、顺序与距离的扰动具有鲁棒性——这些条件会导致序列方法产生大幅精度波动——同时将推理延迟降低多达 11 倍。PARSER, which decouples reading from reasoning, is introduced, which is robust to perturbations in evidence position, order, and distance, conditions that cause large accuracy swings in sequential methods, while reducing inference latency by up to 11x.
本文引入影响引导的响应改写方法,利用 IF 识别干预目标,在保持指令不变的情况下将其响应替换为行为对齐或行为对立的监督信号,以此推动对 TDA 方法的干预感知评估。Influence-guided response rewriting is introduced, which uses IF to identify intervention targets and replaces their responses with behavior-aligned or behavior-opposed supervision while keeping instructions fixed, motivating intervention-aware evaluation of TDA methods.
本文介绍 AgentZip,第一个专为 AI Agent 沙箱设计的内存压缩系统,将压缩范围扩展到任何具有收益表示的页面,并将开销控制从压缩时页面选择转移到恢复时预取。AgentZip is presented, the first memory compression system designed specifically for AI-agent sandboxes, which broadens the compression scope to any page with a profitable representation and shifts overhead control from compression-time page selection to restore-time prefetching.
在真实的博卡拉湖畔地理环境中,由 100 个配备记忆机制的大语言模型 Agent 管理一个封闭且守恒的空间经济,并运行该多 Agent 模拟长达 26 个模拟周,远超典型 Agent 社会研究 1–2 周的时长。100 memory-equipped large language model agents in charge of a closed, money-conserving spatial economy on real Pokhara Lakeside geography and ran this multi-agent simulation for up to 26 simulated weeks, well past the 1-2 weeks typical of agent-society studies.
提出 X-AuT,一个渐进式框架,通过简短的行为探针选择层组合,并通过表征对齐、跨尺度蒸馏、调度式学生策略监督以及 LoRA 微调来恢复被剪枝的模型。X-AuT, a progressive framework that selects layer combinations through short behavioral probes and restores the pruned model through representation alignment, cross-scale distillation, scheduled student-policy supervision, and LoRA finetuning, is introduced.
Generative Late-Interaction Embeddings(GLIE):从归一化质心中学习每个页面 k<<N 个向量,既作为轻量级索引,也作为重建页面完整嵌入集的基础,解码器是其主要设计面。Generative Late-Interaction Embeddings (GLIE): k<<N vectors per page learned from the normalized centroids to serve as both a lightweight index and a basis for regenerating the page's full embedding set, with the decoder as its main design surface.
提出一个用于困难奥林匹克数学自然语言证明生成的开放模型测试时计算流水线,完全在自然语言中运行,无需形式化证明器、外部工具或互联网访问。An open-model test-time-compute pipeline for natural-language proof generation for hard olympiad mathematics that operates entirely in natural language, with no formal prover, external tools, or internet access is presented.
本文提出了一种简单的 flow-control 框架,通过控制 prompt 加入 LLM 活跃集合的速率,实现更高的 token 与 request 吞吐量、更低的平均与尾部 latency,以及更稳定的 KV cache 利用率。A simple flow-control framework is proposed that controls the rate at which prompts join the active set in large language models and achieves higher token and request throughput, lower average and tail latency, and more stable KV cache utilization.
引入 SpatialBlock-15k,一个包含 15,000 个堆块问题的合成数据集,涵盖 3D 到 2D 投影、视角变换和结构组合,并提出受人类认知发展启发的新范式:通过结构化堆块操作任务学习基础空间技能。This work introduces SpatialBlock-15k, a synthetic dataset of 15,000 block-stacking problems covering 3D-to-2D projection, viewpoint transformation, and structural combination and proposes a novel paradigm inspired by human cognitive development: learning foundational spatial skills through structured block-manipulation tasks.
提出 MaP-WAM,一个将记忆作为规划的 Memory-as-Plans 框架,将依赖记忆的世界动作建模分解为基于记忆的规划和以规划为条件的执行,并以长期多模态情景上下文作为规划时证据,而非反复对执行器输入完整历史。MaP-WAM is introduced, a Memory-as-Plans framework that decomposes memory-dependent world-action modeling into memory-grounded planning and plan-conditioned execution, and uses long-term multimodal episodic context as planning-time evidence rather than repeatedly conditioning the executor on the full history.
提出 Negative Self-Distillation(NSD),一个通过偏离有缺陷的推理而非模仿特权解来优化 LLM 的新框架,并一致优于 OPSD 及其他无标签、自举式强化学习(RL)基线。Negative Self-Distillation (NSD) is introduced, a new framework that optimizes LLMs by diverging from flawed reasoning rather than imitating privileged solutions, and consistently outperforms OPSD and other label-free, self-bootstrapping reinforcement learning (RL) baselines.
实验结果表明 DRG-MAPPO 达到了 87% 的 SOTA 胜率,表明该框架在合作空战中有效平衡了关系建模、可解释性和优化稳定性。Experimental results demonstrate that DRG-MAPPO achieves a state-of-the-art win rate of 87%, suggesting that the framework effectively balances relational modeling, interpretability, and optimization stability for cooperative air combat.
提出 UniH3,一个统一层次同质性与异质性的全新框架,用于一体化医学图像修复,在两个大规模基准上的大量实验表明 UniH3 在一体化和单任务医学图像修复上均达到 SOTA 性能。UniH3, a novel framework that Unifies Hierarchical Homogeneity and Heterogeneity and Heterogeneity for all-in-one medical image restoration, is proposed and extensive experiments on two large-scale benchmarks demonstrate that UniH3 achieves state-of-the-art performance on both all-in-one and single-task medical image restoration.
提出 Generative Verification(GenV),通过复用语言模型的原生词表空间,将离线 Z3 等价性预言机蒸馏为无参考、连续参考等价分数,并从理论上证明基于结构、仅判决的验证启发式在这些欺骗性合法轨迹上的检测能力在数学上有界于随机水平。Generative Verification (GenV) is introduced, which distills an offline Z3-equivalence oracle into a reference-free, continuous reference-equivalence score by repurposing the language model's native vocabulary space and theoretically proves that structural, verdict-only verification heuristics are mathematically bounded to chance-level detection on these deceptively valid traces.
一个简单、无训练的框架,其中具备推理能力的视觉语言模型在 Wikipedia 上迭代搜索与推理,动态收集证据,表明推理与检索在稀有实体上具有互补性。A simple, training-free framework in which a reasoning-capable vision-language model iteratively searches and reasons over Wikipedia, gathering evidence dynamically, shows that reasoning and retrieval are complementary on rare entities.
表明将 L2 推理泛化到未见语言的关键路径在于更广泛的语言覆盖、现成可用的多语言非推理数据,以及足够强的英语推理骨干,表明推理是一种与语言无关的行为,可通过精心数据混合在类型多样的语言间迁移。It is shown the path to generalizing L2 reasoning to held-out languages goes through broader language coverage, readily available multilingual non-reasoning data, and a sufficient English reasoning backbone, indicating that reasoning is a language-agnostic behavior that can be transferred across typologically diverse languages through careful data mixing.
构建一个受控的纯自回归测试床,在文本、图像、文生图(T2I)和图生文(I2T)预测的多模态持续预训练中跟踪任务特定验证损失,表明更好的重建并不一定带来更低的任务特定损失或更强的下游性能,且图像分词器的选择在联合优化下会影响文本建模。A controlled pure-autoregressive testbed is built and track task-specific validation losses during multimodal continual pretraining across text, image, text-to-image (T2I), and image-to-text (I2T) prediction, showing that better reconstruction does not necessarily yield lower task-specific losses or stronger downstream performance, and that image tokenizer choice can affect text modeling under joint optimization.
引入 ActReview,一个由反驳引导的后训练框架,将论文特定的诊断连接到具体、有依据的修订计划,并引入 ActReview-Bench,一个包含 1,000 个实例的人工整理基准,用于评估诊断质量和修订实用性。This work introduces ActReview, a rebuttal-guided post-training framework that connects paper-specific diagnoses to concrete, grounded revision plans, and introduces ActReview-Bench, a human-curated benchmark of 1,000 instances for evaluating diagnostic quality and revision usefulness.
结果表明,在评估 harness 中使用 Adaptive Bridge 可将关键的订阅者尾端 p95 延迟从最高 15 s 降至 1.55 ms,覆盖所有损伤严重度,同时保留发布者配置的吞吐量。The results show that using the Adaptive Bridge in the evaluation harness reduces the critical subscriber tail p95 latency from up to 15 s to 1.55 ms across all impairment severities while preserving the publisher's configured throughput.
LIT(Latent Interface Training)是一个与框架无关的两阶段策略:先在无图像条件下建立空间目标条件化的动作先验,再通过姿态监督的潜在接口约束视觉条件化,可在保持或提升 LIBERO 平均成功率的同时改善 LIBERO-Plus 综合成功率。Latent Interface Training (LIT), a framework-agnostic two-stage strategy that first establishes a spatial-goal-conditioned action prior without images, then constrains visual conditioning through a pose-supervised latent interface, improves overall LIBERO-Plus success while preserving or improving average LIBERO success.
在推理、长上下文理解和 Agent 任务中,SAS 在不同注意力预算下均优于可训练的稀疏注意力基线,在紧预算下增益尤为显著,表明其上下文排序对下游任务更有效。Across reasoning, long-context understanding, and agentic tasks, SAS consistently outperforms trainable sparse attention baselines across attention budgets, with especially large gains under tight budgets, demonstrating more effective context ranking for downstream tasks.
提出一种物理引导框架,用于提升单图像部件感知 3D 生成效果,使其几何结构物理相容且连接稳定,并在部件接触面引入参数化连接器。This work proposes a physics-guided framework for improving single-image part-aware 3D generation with physically compatible geometry and stable connections, and introduces parameterized connectors at their contact surfaces.
COBRA-Skills 是一个高效框架,将技能优化建模为在动态演化的候选空间上的预算式序贯优化,对 Agent harness 变化保持鲁棒,并在目标模型自身用于技能生成与优化时依然有效。COBRA-Skills is introduced, an efficient framework that formulates skill optimization as budgeted sequential optimization over a dynamically evolving candidate space and remains robust to changes in the agent harness and performs effectively when the target model itself is used for skill generation and refinement.
PLC-DPO 将每个偏好对的训练信号路由为 clean、flip 或 tie 三类,将噪声偏好学习从单纯过滤可疑样本重构为主动修正监督方向与强度。PLC-DPO is proposed to robustly optimize preferences by routing each pair's training signal as a clean, flip, or tie case, which reframes noisy preference learning as actively correcting supervision direction and strength rather than merely filtering suspicious examples.
本研究提出了 StepAudio 3 Gen,一个通用音频生成模型,在统一框架下支持零样本文本到语音合成(TTS)、声音设计、人声生成、音效、音乐、风格化语音以及多种音频类型的混合生成。This study introduces StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types within a unified framework.
这是首个端到端证明:一个七人独立团队可以训练出在 agentic 网络安全能力上领先的开源权重模型,且全部三个 checkpoint 在相近参数规模下均排名第一。This is the first end-to-end demonstration that a seven-person independent team can train open-weight models with leading agentic cyber capability and all three checkpoints rank 1st among models at comparable parameter scales.
本研究探索在反馈有限的在线场景下,将 prompt 自适应路由到大语言模型专家以最大化响应质量,并提出了策略性地选择和观察奖励以最小化遗憾的算法。This work studies adaptive routing of prompts to large language model experts to maximize response quality in an online setting with limited feedback and proposes algorithms that strategically select and observe rewards to minimize regret.
提出了 Diverse Skill Routing,一个具备多样性感知能力的重排序框架,使用 Determinantal Point Process 在相关性与非冗余性之间取得平衡,在强 pointwise 重排序基线之上提升了召回率与完整覆盖率,且在多技能 query 上增益更大。Diverse Skill Routing is proposed, a diversity-aware reranking framework that uses a Determinantal Point Process to balance relevance and non-redundancy and improves recall and full coverage over a strong pointwise reranking baseline, with larger gains on multi-skill queries.
研究发现,仅使用合成数据的训练即可具备竞争力;将 0.9B 参数的 PaddleOCR-VL-1.6 适配为 Wayu-Paxa-OCR-Zero,一个无需真实泰语文档页 OCR 标签即可适配的泰语 OCR 模型,表明仅合成数据训练即可具备竞争力。It is found that synthetic-only training can be competitive, and the 0.9B-parameter PaddleOCR-VL-1.6 is adapted into Wayu-Paxa-OCR-Zero, a Thai OCR model adapted without OCR labels from real Thai document pages, showing that synthetic-only training can be competitive.
提出面向检索增强生成(RAG)的通用微调方案——RAG 模型融合预训练参数化记忆与非参数化记忆进行语言生成;研究发现,相较 SOTA 的纯参数化 seq2seq 基线,RAG 模型生成的文本更具针对性、更多样且更符合事实。A general-purpose fine-tuning recipe for retrieval-augmented generation (RAG) -- models which combine pre-trained parametric and non-parametric memory for language generation, and finds that RAG models generate more specific, diverse and factual language than a state-of-the-art parametric-only seq2seq baseline.