研究库 论文知识库
Papers · organized/paper_cards

论文

1640 张论文卡片 · OA 绿色

开放获取 全部 绿色 · 1640
Cross-Model Memory Transfer via Target-Side Reader Adaptation
通过目标侧 Reader 适配的跨模型记忆迁移
arXiv:2608.17050 RAG 检索增强 方法 OA · 绿色 被引 3 · S2

结果表明,Engram 可充当可复用的外部知识工件,前提是目标侧具备兼容的 Reader 接口;当直接复用 Reader 效果不足时,目标侧适配可进一步改善对齐效果。The results suggest that Engram can serve as a reusable external knowledge artifact, provided that the target has access to a compatible reader interface and target-side adaptation can further improve alignment when direct reader reuse is insufficient.

Demystifying Agent Skills: Why They Work-Until They Don't
揭秘 Agent Skill:为何它们有效——直到失效
arXiv:2608.14036 Agent 智能体 评测集 OA · 绿色 被引 2 · S2

本文设计了一项对比研究,结合受控定量实验与配对轨迹分析,并将观察结果归纳为一个包含三个高层类别与十二种 Skill 使用模式的分类法,表明当噪声轨迹转化为稳定执行的过程性锚点时,Skill 便会发挥作用。This work designs a contrastive study that combines controlled quantitative experiments with paired trajectory analysis and consolidates observations into a taxonomy of three high-level categories and twelve skill-use modes, showing that skills work when noisy trajectories become procedural anchors that stabilize execution.

MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding
MoE-ViE:面向高效图像与视频理解的混合专家视觉编码器
arXiv:2608.17402 多模态 方法 OA · 绿色 被引 1 · S2

本文系统性地研究了视觉编码器扩展中的 MoE 设计,发现细粒度 MoE 拓扑相较于稠密与标准 MoE 基线均带来显著提升;提出了一种无辅助损失的均衡变体以改善专家利用率,并设计了专用 MoE kernel 以缓解推理时延开销。This work systematically study MoE designs for vision encoder scaling and finds that fine-grained MoE topologies yield substantial gains over both dense and standard MoE counterparts, and proposes an auxiliary-loss-free balancing variant for better expert utilization, and designs a specialized MoE kernel to mitigate inference latency overhead.

[TOKI] A Bitemporal Operator Algebra for Contradiction Resolution in LLM-Agent Persistent Memory
[TOKI] 面向LLM-Agent持久记忆中矛盾解析的双时态算子代数
arXiv:2606.06240 Agent 智能体 应用落地 OA · 绿色 被引 8 · S2

研究表明矛盾解析本质上是写入时并发控制,并将缺失的契约——一个在隔离性、模式与来源维度上被证明正确的写入时正确性规范——显式化,固定了每个生产启发式都默认假设、却没有任何已部署系统显式给出的保证。It is shown that contradiction resolution is write-time concurrency control and make the missing contract explicit, a write-time correctness specification, proved sound across isolation, schema, and provenance, pinning the guarantee every production heuristic assumes but no deployed system makes explicit.

LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents
LEGO-RL:面向编码 Agent 的 Harness 原生强化学习
arXiv:2608.17393 Agent 智能体 方法 OA · 绿色 被引 3 · S2

本文提出了 LIFE-RL 框架,在不修改内部控制流的前提下,将原生编码 Agent harness 与可扩展的策略梯度优化相连接,并通过 GSPO 在三个原生编码 Agent harness 上训练稀疏 MoE 模型 Qwen3.5-35B-A3B 对其进行了评估。LIFE-RL, a framework that bridges native coding-agent harnesses with scalable policy-gradient optimization without modifying their internal control flow, is presented and evaluated by training the sparse MoE model Qwen3.5-35B-A3B with GSPO across three native coding-agent harnesses.

PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX
PTXBench:面向 GPU kernel 优化与架构特定 PTX 的 LLM 基准与适配
arXiv:2608.17379 评测基准 评测集 OA · 绿色 被引 1 · S2

PTXBench 提供了一个可审计的测试平台,用于衡量并提升 LLM 利用持续演进 GPU 架构的能力,并表明各 LLM 在架构特定 PTX 能力上仍参差不齐。PTXBench provides an auditable testbed for measuring and improving LLMs'ability to exploit evolving GPU architectures, and shows that architecture-specific PTX capability remains uneven.

The Problem Is the Problem: Towards Scalable Mathematical Discovery
问题才是问题:迈向可扩展的数学发现
arXiv:2608.16977 RAG 检索增强 方法 OA · 绿色 被引 2 · S2

受搜索与推荐系统启发,本文构建了 Find、Attempt 与 Recommend(FAR),即一个从文献到综述的级联流程,可自动搜索合适的问题,并将人类注意力聚焦于经过多阶段筛选的成果上。Inspired by search and recommender systems, this work builds Find, Attempt, and Recommend (FAR), a literature-to-review cascade that automates the search for suitable problems and focuses human attention on artifacts that have passed several stages of filtering.

Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL
Co-RL:无监督推理在多智能体强化学习中从多样化群体中涌现
arXiv:2608.17253 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Co-RL,一种由多个解耦模型组成的框架,这些模型不共享参数,通过基于彼此输出奖励的强化学习同时进行优化,并表明无监督推理可以通过协作式多智能体训练涌现This work introduces Co-RL, a framework in which multiple decoupled models, sharing no parameters, are simultaneously optimized through RL using rewards derived from their peers, and shows that unsupervised reasoning can emerge through cooperative multi-agent training.

Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning
潜在世界模型中的决策度量对齐:用于 MPC 规划的诊断方法与动作条件目标
arXiv:2608.18746 安全与风险 方法 OA · 绿色 被引 7 · S2

动作条件目标改善了基于欧几里得代价与 CEM 的潜在 MPC 所使用的几何结构,DA-LeWM 在 LeWM 基础上增加了逆动力学和演示条件的目标-动作头,加速了收敛并取得比 LeWM 更高的在线成功率Action-conditioned objectives improve the geometry used by Euclidean-cost, CEM-based latent MPC, and DA-LeWM augments LeWM with inverse-dynamics and demonstration-conditioned goal-action heads, and accelerates convergence and achieves higher online success than LeWM.

SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation
SoftVTBench:面向可形变物体操作的形变感知视触觉数据集与基准
arXiv:2608.18701 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

实验表明,仅提供触觉本身并不能确保有效的多模态融合,SoftVTBench 为研究策略不仅能否成功,还在于其如何与可形变物体物理交互,以及触觉在何时改善这种交互,提供了统一的视触觉资源Results show that making touch available does not by itself ensure effective multimodal fusion, and SoftVTBench provides a common visuo-tactile resource for studying not only whether a policy succeeds, but how it physically interacts with deformable objects and when touch improves that interaction.

SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation
SemComp-Bench:视频生成中的语义任务完成度基准评测
arXiv:2608.17426 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

在代表性视频生成模型上的实验表明,在保持参考图像中任务相关语义一致性的同时实现预期结果仍然具有挑战性Experiments on representative video generation models show that achieving intended outcomes while maintaining task-relevant semantic grounding in reference images remains challenging.

Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence
Zetta ζ:面向自进化物理智能的高效闭环具身 Harness
arXiv:2608.16590 评测基准 方法 OA · 绿色 被引 9 · S2

本文提出 Zetta,一种闭环具身 harness,在保持基础策略冻结的同时在线演化基于代码的运行时评判器与恢复技能,表明闭环 harness 的自进化为可靠的物理智能开辟了一条可扩展的路径Zetta is presented, a closed-loop embodied harness that evolves code-based runtime critics and recovery skills online while keeping the base policy frozen, and shows that closed-loop harness self-evolution opens a scaling path for reliable physical intelligence.

Training Leaves Traces: Centered Residual Signatures for Language Model Lineage Verification
训练留痕:用于语言模型谱系验证的居中残差签名
arXiv:2608.14929 工程化 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

结果为兼容的开源权重语言模型检查点建立了一种被动、无需数据的溯源信号,且该投影配对信号出现在六个及更多语言模型系列中The results establish a passive, data-free provenance signal for compatible open-weight language-model checkpoints, and the projection-pairing signal appears across six language-model families and beyond.

Scaling Creative Writing Beyond Story-Centric Data with Attribute-Guided Genre Expansion
超越故事中心数据的可扩展创意写作:基于属性引导的题材扩展
arXiv:2608.13947 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

实验表明,在该数据上微调的模型不仅持续优于基座模型和面向写作的专门基线,还优于基于现有写作语料训练的模型,表明受控的题材扩展是稳健创意写作能力的关键驱动力Experiments demonstrate that models fine-tuned on the data consistently surpass not only base models and writing-specialized baselines, but also models trained on existing writing corpora, indicating that controlled genre expansion is a key driver of robust creative writing capability.

[Larch] Learned Query Optimization for Semantic Predicates
[Larch] 面向语义谓词的习得式查询优化
arXiv:2606.07923 数据与向量库 方法 OA · 绿色 被引 2 · S2

本文提出Larch,一个用于优化AI SQL查询中语义过滤器执行的框架,并给出其两种变体:Larch-A2C与Larch-Sel,二者在token使用量上均始终优于现有语义过滤器优化技术。This paper introduces Larch, a framework for optimizing the execution of semantic filters in AI SQL queries and presents two Larch variants: Larch-A2C and Larch-Sel, which always outperform existing semantic filter optimization techniques in terms of token usage.

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist
OmniScientist:全模态跨学科 AI 科学家
arXiv:2608.13558 多模态 方法 OA · 绿色 被引 1 · S2

本文提出 OmniScientist,一种端到端、全模态 AI 科学家,可直接基于异构原始证据开展跨学科研究,表明全生命周期感知对于基于证据的科学发现至关重要,并为构建广泛适用的 AI 科学家提供了一条切实可行的路径OmniScientist is introduced, an end-to-end, omni-modal AI scientist that conducts multidisciplinary research directly from heterogeneous raw evidence and demonstrates that lifecycle-wide perception is essential for evidence-grounded scientific discovery and provides a practical path toward broadly capable AI scientists.

CTIFoundry: An Agent-Native Corpus Scaffold for Cyber Threat Intelligence
CTIFoundry:面向网络威胁情报的 Agent 原生语料架构
arXiv:2608.18613 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

论文认为,agentic CTI 调查的瓶颈在于该 substrate 而非模型能力,并提出了面向 Agent 的语料库脚手架 CTIFoundry。It is argued that this substrate, not model capability, is the bottleneck on agentic CTI investigation, and CTIFoundry, an agent-native corpus scaffold, is presented.

SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents
SkillGate:在长视野 Agent 中训练策略内技能选择
arXiv:2608.18852 Agent 智能体 方法 OA · 绿色 被引 1 · S2

SkillGate 将 9B 策略的成功率从 40.8% 提升至 53.2%,显著优于将相同预算仅用于 outcome reward 的方案,同时将误导性候选的暴露减少三分之二,并读取更少的 skill。SkillGate lifts a 9B policy from 40.8% to 53.2% trial success, well ahead of the identical budget spent on outcome reward alone, while cutting exposure to misleading candidates by two thirds and reading fewer skills.

The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning
越流行越难遗忘:面向 LLM 遗忘的自适应流行度方法
arXiv:2608.14229 工程化 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 AdaPop(自适应流行度)方法,将局部 token 置信度与源自外部代理的逐事实流行度相关指数相结合,并通过双上升控制器在每个 epoch 调整 retain 惩罚来自动平衡遗忘与保留。The AdaPop (Adaptive Popularity) method is proposed, which combines local token confidence with a per-fact popularity-dependent exponent derived from an external proxy, and automates the forget-retain balance via a dual-ascent controller that adjusts the retain penalty each epoch.

MissDiag: Diagnostic Evaluation of Incomplete-Knowledge Robustness in KGQA and KG-RAG
MissDiag:KGQA 与 KG-RAG 中不完整知识鲁棒性的诊断式评估
arXiv:2608.18489 RAG 检索增强 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

MissDiag 将聚合鲁棒性度量转化为类型化的诊断归因,为在不完备知识下比较、诊断和压力测试 KGQA 与 KG-RAG 系统提供了更具可解释性的基础。By transforming aggregate robustness measurement into typed diagnostic attribution, MissDiag provides a more interpretable basis for comparing, diagnosing, and stress-testing KGQA and KG-RAG systems under incomplete knowledge.

VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation
VA-Judger:基于人类偏好反馈的联合音视频生成奖励建模
arXiv:2608.18607 多模态 方法 OA · 绿色 被引 4 · S2

提出 VA-Judger,一种链式思考的通用奖励模型:从质量差距明显的样本对中学习以建立结构化输出和粗粒度偏好判别,再通过拒绝采样(以人类标注验证)蒸馏出可靠的偏好解释用于更难的质量相近样本比较,最后执行维度级强化学习,将人类反馈分解到各独立质量维度以获得更稠密的奖励信号。VA-Judger is proposed, a chain-of-thought omni-reward model that learns from pairs with clear quality gaps to establish structured output and coarse preference discrimination, then distills reliable preference explanations for harder near-quality comparisons via rejection sampling verified against human annotations, and finally performs dimension-wise reinforcement learning that decomposes human feedback into individual quality dimensions for denser reward signals.

Evaluating Music Context Preservation: A Multi-facet Framework for Music Editing Systems
音乐上下文保留评估:面向音乐编辑系统的多面框架
arXiv:2512.14629 评测基准 应用落地 OA · 绿色 被引 1 · S2

提出首个 MuseCP 评估框架,涵盖四类音乐 facet,使用细粒度且量身定制的指标来捕捉音乐属性的细微变化,并希望为开发更有效、更可靠、具备强大 MuseCP 能力的音乐编辑策略提供实践指导。The first MuseCP evaluation framework is introduced that covers four categories of music facets with fine-grained and well-tailored metrics to capture nuanced changes in music attributes and hopes it can offer practical guidance for developing more effective and reliable music editing strategies with strong MuseCP capability.

Towards Real-Time and Adaptable LiDAR Scene Completion
迈向实时且自适应的 LiDAR 场景补全
arXiv:2608.16490 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 RapidLiDAR,一种将初始化本身视为可学习、数据驱动组件的 LiDAR 场景补全方法,在与 SOTA 相当的补全性能下,0.1 秒完成整个场景,比此前最快方法快 2.3 倍。RapidLiDAR is presented, a LiDAR scene completion method that treats the initialization itself as a learned, data-driven component and achieves completion performance on par with the state of the art while completing a full scene in 0.1 seconds, which is 2.3 times faster than the fastest prior method.

Bounded Agents: Delegation Security for Multi-Agent AI Systems
Bounded Agents:多 Agent AI 系统的委派安全
arXiv:2608.15888 Agent 智能体 方法 OA · 绿色 被引 6 · S2

受损模型评估测试在第一个合法 tool call 之后插入 ground-truth 攻击调用,从而独立于模型行为测试 APC,证明了 APC 实现的 Blast Radius 单调性与组合可靠性。The compromised-model evaluation tests APC independently of model behavior by inserting the ground-truth attack call after the first legitimate tool call, which proves Blast Radius Monotonicity and Composition Soundness for APC implementations and proves Blast Radius Monotonicity and Composition Soundness for APC implementations.

FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving
FlashPrefill V2:面向长上下文 LLM 服务的块稀疏 Prefill 注意力
arXiv:2608.19758 LLM 基础设施 应用落地 OA · 绿色 被引 2 · S2

本文引入一个均值修正项,有效抑制近似误差,即使在极端稀疏度下也能将性能下降保持在可控范围;并使用 PackGQA 内存访问、warp specialization 和 pingpong 流水线重新设计稀疏注意力算子。This paper introduces a mean correction term that effectively suppresses the approximation error, keeping performance degradation manageable even at extreme sparsity levels, and redesigns the sparse attention operator with PackGQA memory access, warp specialization, and pingpong pipelining.

[DataEvolver] Automatic Data Preparation for Large Language Models through Multi-Level Self-Evolving
DataEvolver:基于多层级自演化的 LLM 自动化数据准备
arXiv:2606.07001 工程化 方法 OA · 绿色 被引 3 · S2

实验表明,DataEvolver 显著提升了数据质量,相比在原始数据上训练,下游 LLM 性能平均提升 10%,凸显了 LLM 与数据迭代协同演化的新机遇。Experiments show that DataEvolver substantially improves data quality and achieves an average 10\% gain in downstream LLM performance compared with training on original data, highlighting new opportunities for the iterative co-evolution of LLMs and data.

4DAnyone: Create Anyone in 4D from a Casual Monocular Video
4DAnyone:从随手单目视频创建 4D 人物
arXiv:2608.20335 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

4DAnyone 在新视角视频质量和下游 4DGS 重建上均优于先前方法,并具有稳健的野外泛化能力。4DAnyone outperforms prior methods in both novel-view video quality and downstream 4DGS reconstruction, with robust in-the-wild generalization.

Repo0: Design-Driven Zero-to-All Code Generation
Repo0:设计驱动的零到全代码生成
arXiv:2608.19854 Agent 智能体 方法 OA · 绿色 被引 2 · S2

提出 Repo0,一个面向零到全代码生成的持续结构演化框架,维护显式的架构状态,实例化为双有向无环图(Dual-DAG),由需求级 DAG、组件级 DAG 及其对齐关系组成。Repo0 is presented, a continuous structural evolution framework for zero-to-all code generation that maintains an explicit architectural state instantiated as a Dual-Directed-Acyclic-Graph (Dual-DAG), consisting of a requirement-level DAG, a component-level DAG, and their alignment relation.

Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization
Inject, Align, Recover:面向免检索文档知识内化的分阶段后训练
arXiv:2608.20281 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 IAR(Inject, Align, and Recover),一个三阶段后训练框架,将结构化文档知识注入、问答行为对齐与通用能力恢复解耦,提升面向无检索文档内化的"领域主—领域通"前沿。This work proposes IAR (Inject, Align, and Recover), a three-stage post-training framework that separates structured document knowledge injection, QA behavior alignment, and general ability recovery and improves the domain-primary domain-general frontier for retrieval-free document internalization.

Towards Quantifying Benchmark Optimization in ASR Models
迈向 ASR 模型中基准过拟合的量化
arXiv:2608.19936 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一种量化基准优化的方法论,聚焦于音频对参考转写不充分确定的情形,指出高性能模型会表现出基准条件化行为,从而虚高基准得分,却未必反映通用转写能力的真正提升。This work presents a methodology for quantifying benchmark optimization, focusing on cases where the audio underdetermines the reference transcript, and indicates that high-performing models exhibit benchmark-conditioned behaviors that can inflate benchmark performance without reflecting improved general-purpose transcription ability.

SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?
SWE-bench Science:编码 Agent 能解决科学领域的工程任务吗?
arXiv:2608.19799 Agent 智能体 评测集 OA · 绿色 被引 4 · S2

一项配对消融实验在保留仓库与可执行工程上下文的同时移除显式科学指导,表明科学知识并非一律有益:可靠信息能约束修复、提升平均表现与 token 效率,而错位指导则会诱发锚定,不必然提升精确修复成功率。A paired ablation that removes explicit scientific guidance while preserving the repository and executable engineering context shows that scientific knowledge is not uniformly beneficial: well-grounded information can constrain repair and improve average performance and token efficiency, whereas poorly aligned guidance can induce anchoring and does not necessarily improve exact repair success.

Chain-of-Experience for Continual LLM Improvement
Chain-of-Experience:基于经验链的 LLM 持续改进
arXiv:2608.18027 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文研究 LLM 在测试时从迭代经验中学习的机制,称之为 Chain-of-Experience (CoE):模型通过与自身或环境反馈的迭代交互积累经验痕迹,形成超越零样本推理的持续改进循环。This study studies how LLMs learn from iterative experience at test time, a setting the authors refer to as Chain-of-Experience (CoE), where models accumulate experiential traces through iterative interactions with self or environmental feedback to form a continual improvement loop beyond zero-shot inference.

SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback
SkillEvo:从多轮交互反馈中自我更新的演化梯度
arXiv:2608.13120 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文指出,持续技能演化的关键瓶颈既非编辑能力亦非迭代轮数,而在于评估反馈能否持续提供可信的演化梯度;据此提出 SkillEvo,由可信反馈生成梯度,由可控治理约束方向。This work argues that the binding constraint on sustained skill evolution is neither editing capability nor the number of iterations, but whether the evaluation feedback keeps supplying trustworthy evolution gradients, and introduces SkillEvo, in which trustworthy feedback generates the gradient and controllable governance constrains its direction.

Inadvertent Context Leakage in Language Models
语言模型中的非故意上下文泄露
arXiv:2608.19857 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

泄露带来两类现实攻击:一个训练好的分类器可从常规自然语言输出中推断用户记忆的语义谓词;一个由 RL 训练的对抗者可从生产级风格的 Agent 中完整提取社会安全号码。Leakage enables two practical attacks: a trained classifier that infers semantic predicates about user memories from routine natural-language outputs, and an RL-trained adversary that extracts full Social Security Numbers from a production-style agent.

Listening Forward: Next Patch Embedding Prediction Enables Scalable Audio Learners
Listening Forward:下一 patch 嵌入预测实现可扩展的音频学习
arXiv:2608.19863 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 NAPE(Next-Audio-Patch-Embedding prediction),一个自监督框架:因果 Transformer 仅依据因果掩码与 stop-gradient,从先前 patch 嵌入预测对数梅尔频谱图的下一 patch 嵌入。NAPE (Next-Audio-Patch-Embedding prediction), a self-supervised framework in which a causal Transformer predicts each next patch embedding of a log-mel spectrogram from the previous ones, using causal masking and stop-gradient as its sole training signal is introduced.

Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See
低资源语言下的思考:SFT 构建什么,RL 修复什么,准确率看不到什么
arXiv:2608.17744 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

选取三个前沿混合专家模型在低资源语言上进行推理微调,提出六个可度量的行为维度,且每维度均设门拒绝任何与输出长度相关的指标,并报告其自家评测工具为何失效。Take three frontier mixture-of-experts models and fine-tune them to reason in a low-resource language and propose six behavioural dimensions that make changes measurable, each gated to reject any metric that correlates with output length, and report how their own instruments lied.