研究库 论文知识库
Papers · organized/paper_cards

论文

1173 张论文卡片 · 方法

开放获取 全部 绿色 · 1640
BoundaryMORPH: Budgeted Reranking via Active Set Selection for Diffuse Retrieval
BoundaryMORPH:通过主动集合选择实现预算化重排序的扩散检索
arXiv:2609.27213 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

BoundaryMORPH 是一种新算法,专门为 LLM 的上下文容量 k 分配 CE 预算,在多个模型与数据集上针对开放性查询取得了 SOTA 集合检索质量。BoundaryMORPH is introduced, a novel algorithm that allocates CE budget specifically for the LLM's context capacity $k$ and achieves state-of-the-art set retrieval quality across multiple models and datasets with open-ended queries.

SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance
SAGE:通过拓扑引导缓解长程推理偏置
arXiv:2609.30192 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 SAGE(Structural Admissibility-Guided Exploration),一个通过注入结构引导来缓解长程推理中探索偏差与累积偏差的统一框架,在 Andrews-Curtis 问题上取得了最高 8 倍的提升。This work proposes SAGE (Structural Admissibility-Guided Exploration), a unified framework that injects structural guidance to alleviate exploration bias and compounding bias in long-horizon reasoning and achieves up to an 8-fold improvement on the Andrews-Curtis problem.

LastOPD: Taming Collapse in Latent On-Policy Distillation
LastOPD:驯服潜在 On-Policy 蒸馏中的坍缩问题
arXiv:2609.28845 多模态 方法 OA · 绿色 被引 2 · S2

LastOPD 仅在最后一层状态(LM head 共同读取的接口)施加潜在信号,且仅在进入 token 级 OPD 前的 10 步交叉淡入阶段施加,从而保留潜在信号中的有用部分,并在坍缩发生前将 student 交由 token 级监督。LastOPD is proposed, which applies the latent signal only at the last-layer state, the common interface both LM heads read, and only during a 10-step crossfade into token-level OPD, which keeps the useful part of the latent signal and hands the student to token-level supervision before the collapse sets in.

TimeEvo: Failure-Driven Self-Evolution of a Time Series Agent
TimeEvo:时序 Agent 的失败驱动自演化
arXiv:2609.27277 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

TimeEvo 将 Agent 诊断出的失败聚类为能力缺口,为每个缺口规划测量,合成只产出证据的工具来填补缺口,并仅通过配对的准入门控接纳候选工具库。TimeEvo is proposed, which clusters an agent's diagnosed failures into capability gaps, plans a measurement for each, synthesizes evidence-only tools that fill them, and admits the candidate library only through a paired admission gate.

Disaggregated Quantization: Specializing LLM Prefill and Decode
解耦量化:针对 LLM Prefill 与 Decode 的专用量化
arXiv:2609.26333 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

评估 vLLM 在分离式 serving 下的精度,并通过训练后量化在最高 2.8T 参数的模型上进一步验证共享权重格式的分离式 serving。Assessment of accuracy under disaggregated serving in vLLM and further validate shared-weight format disaggregation through post-training quantization on models up to 2.8T parameters are evaluated.

D-JEPA: A Decision-Aligned Latent World Model
D-JEPA:一种决策对齐的潜在世界模型
arXiv:2609.24749 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

D-JEPA 是一种决策对齐的潜在世界模型,从已执行结果中学习候选未来之间的决策相关关系,将决策相关关系结构确立为预测世界建模与有效控制之间的直接桥梁。D-JEPA is introduced, a decision-aligned latent world model that learns decision-relevant relations among candidate futures from executed outcomes, establishing decision-relevant relational structure as a direct bridge between predictive world modeling and effective control.

Just-In-Time Agent Memory with Runtime Agentic Research
基于运行时 Agentic 研究的 Just-In-Time Agent 记忆。
arXiv:2609.34385 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Just-In-Time Agent Memory (JAM) 是一个可训练框架,在运行时进行查询条件下的上下文构建,其任务性能优于 AOT 式记忆系统,同时比先前的可训练 Agent 记忆方法显著更高效。Just-In-Time Agent Memory (JAM), a trainable framework for query-conditioned context construction at runtime, is proposed, where it achieves stronger task performance than AOT-style memory systems while remaining substantially more efficient than prior trained agentic memory approaches.

Stashbird: Efficient Speaker-Indexed Memory for Conversational Agents
Stashbird:面向对话 Agent 的高效说话人索引记忆。
arXiv:2609.34242 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Stashbird 是一种 Agent 记忆系统,通过显式来源溯源将源事件链接到派生记忆状态,在 LongMemEval-S 和 GroupMemBench 上准确率高于 Hindsight,在 EverMemBench 上与之相当。Stashbird is presented, an agent memory system that links source episodes to derived memory state through explicit provenance and achieves higher accuracy than Hindsight on LongMemEval-S and GroupMemBench and comparable accuracy on EverMemBench.

InfiniHand: Streaming World-Space Hand Motion Estimation from Egocentric Video
InfiniHand:基于第一人称视频的流式世界空间手部运动估计。
arXiv:2609.35743 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

InfiniHand 是一个端到端流式前馈框架,联合估计 MANO 参数、相机轨迹与手部位置,直接从未标定的自我中心视频出发,相比 ViDiHand 在 ARCTIC PA-p 上降低 21%,并显著缓解世界空间漂移。InfiniHand is an end-to-end streaming feed-forward framework that jointly estimates MANO parameters, camera trajectories, and hand locations directly from uncalibrated egocentric video, achieving a 21% reduction in ARCTIC PA-p compared to ViDiHand while substantially mitigating world-space drift.

GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space
GeoVerse: 几何潜空间中的世界一致新视角合成
arXiv:2609.35734 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

GeoVerse 是一个在世界一致条件下合成新视角的框架,在预训练 3D 基础模型的几何潜在空间内进行生成,并通过 ControlNet 风格 adapter 注入视频生成模型的外观先验。GeoVerse is a framework that synthesizes world-consistent novel views by performing generation within the geometric latent space of a pretrained 3D foundation model and injecting appearance priors from a video generative model via a ControlNet-style adapter.

Systems 补充候选
arXiv:2606.03910 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

NetKV,一种使用该 oracle 信息的 O(|D|) 每请求贪心策略,其层级排序被证明对过时遥测数据具有鲁棒性;并证明随着上下文长度增长,忽略网络项会使仅缓存感知的调度任意次优。NetKV, the O(|D|) per-request greedy that consumes this oracle, has tier rankings that are provably robust to stale telemetry, and it is proved that ignoring the network term renders cache-aware-only scheduling arbitrarily suboptimal as context length grows.

Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning
通过交错强化学习在统一模型中学习原生反思
arXiv:2609.35767 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

UMM-Reflection 在统一模型内利用强化学习(RL)完成完整反思轨迹:兄弟轨迹共享同一初始图像,因此组相对优势可比较不同反思策略,而单条轨迹级优势同时更新反思 token 与基于 flow 的修订,避免了逐轮信用分配的组合爆炸。UMM-Reflection is introduced, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and the flow-based revisions, avoiding the combinatorial blow-up of per-round credit assignment.

When Do Model Internals Help? Exploring the Role of Representation Engineering in LLM Safety
模型内部何时有用?探究表征工程在 LLM 安全中的作用
arXiv:2609.34771 安全与风险 方法 OA · 绿色 被引 0 · S2 + OpenAlex

总体而言,representation engineering 并不能普遍替代行为安全护栏,但在特定条件下具有实际优势,并可带来互补的安全收益。Overall, representation engineering does not generally replace behavioral safeguards, but offers practical advantages under specific conditions and can provide complementary safety benefits.

WideSWE: Can Coding Agents Coordinate Changes Across Repositories?
WideSWE: 编码 Agent 能否跨代码仓库协调变更?
arXiv:2609.33382 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 WideSWE,用于在跨仓库任务上评估编码 Agent,并在相同 prompt 下将其与联合执行进行对比,以考察逐仓库工作是否能缓解相关困难。This work introduces WideSWE to evaluate coding agents on cross-repository tasks, and compares it with joint execution under identical prompts to examine whether working on one repository at a time can alleviate difficulties.

Residual Transferability in Neural Image Watermarking
神经图像水印中的残差可迁移性
arXiv:2609.32241 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文识别出两种强化水印证据对载体图像依赖性、抑制残余可迁移性的机制,并提出 CoverLock,一种即插即用策略,可在不重新设计架构的前提下强化现有水印系统的图像依赖性。This work identifies two mechanisms that strengthen the dependence of watermark evidence on the cover image, thereby suppressing the residual transferability, and introduces CoverLock, a plug-and-play strategy for existing watermarking systems that strengthens such image dependence without architectural redesign.

FlowTool: Controlling Tool Parameter in Image Retouching via Flow Matching
FlowTool: 基于流匹配控制图像修图中的工具参数
arXiv:2609.35673 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 FlowTool,一个通过 conditional rectified flow 直接建模以输入图像和用户指令为条件的高质量工具参数分布的框架,并显著提升了推理效率。This work introduces FlowTool, a framework that directly models the distribution of high-quality tool parameters conditioned on the input image and user instruction using conditional rectified flow, and significantly improves inference efficiency.

Nereus: Adaptive Parallelism for LLM Post-Training
Nereus: 面向 LLM 后训练的自适应并行
arXiv:2609.34645 工程化 方法 OA · 绿色 被引 1 · S2

Nereus 是一种成本感知的运行时,将 RL 后训练任务适配为高效执行计划,并基于内存可行的全局计划执行状态转移,使用与运行任务校准的成本模型来接纳转移。Nereus is a cost-aware runtime that adapts RL post-training jobs into efficient execution plans and executes a transition using a memory-feasible global plan and admits the transition using a cost model calibrated against the running job.

DISCO: Distributed Long Context Scaling with Grounding-Reasoning Disaggregation
DISCO: 基于接地-推理解耦的分布式长上下文扩展
arXiv:2609.33485 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

受 Apache Spark 等分布式计算框架启发,DISCO 将长上下文切分到一组专职 Worker LLM 上,进行并行、局部的 grounding,建立了鲁棒长上下文推理的高效范式。Inspired by distributed computing frameworks like Apache Spark, DISCO partitions long context across a fleet of Worker LLMs dedicated exclusively to parallel, localized grounding, establishing a highly efficient paradigm for robust long-context inference.

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
面向代码 Agent 强化学习的分组智能评分与优势再分配
arXiv:2609.32577 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 GAGAR,一个面向代码 agent RL 的质量感知信用再分配框架;结果表明,将基于测试的验证与成组的 agentic 评分相结合,可提升代码 agent RL 的质量与稳定性。GAGAR, a framework for quality-aware credit redistribution in code agent RL, is introduced and results support combining test-based verification with groupwise agentic grading to improve the quality and stability of code agent RL.

Relic: From Multi-Agent Collaboration to Persistent Organizational Capability
Relic:从多 Agent 协作到持久化的组织能力
arXiv:2609.32965 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 Relic,它将反复出现的协作失败转化为组织拥有的、可执行的协议,使协作经验成为超越其创建者仍持续有用的组织级持久状态。The introduction of Relic, which turns recurring collaboration failures into organization-owned, executable protocols, and how collaboration experience can become persistent organizational state that remains useful beyond the members who created it are shown.

Systems 补充候选
arXiv:2510.09665 LLM 基础设施 方法 OA · 绿色 被引 174 · S2

本工作提出 LMCACHE,首个也是目前最高效的开源 KV 缓存方案,可将现代 LLM 引擎生成的 KV 缓存从 GPU 显存中提取并存储,并跨引擎和查询共享。This work presents LMCACHE, the first and so far the most efficient open-source KV caching solution, which extracts and stores KV caches generated by modern LLM engines out of the GPU memory and shares them across engines and queries.

G^2PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation
G^2PTQ:通过广义梯度补偿改进 LLM 训练后量化
arXiv:2609.31009 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 G²PTQ,一个具有广义梯度补偿的统一 PTQ 框架,在全局监督的分块优化目标下融合一阶与二阶信息,能更好地对齐全精度模型,性能优于 SOTA 基线。G$^2$PTQ is presented, a unified PTQ framework with Generalized Gradient Compensation that integrates both first- and second-order information under a globally supervised, block-wise optimization objective and enables better alignment with the full-precision model, outperforming state-of-the-art baselines.

SANTA++: Sampling Attention through Representative Keys
SANTA++: 通过代表性 Key 进行采样的注意力机制
arXiv:2609.35629 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

注意力往往集中在上下文中一小部分 token 上,但每个 query 关注的关键子集各不相同。为利用这种动态结构,我们提出 SANTA++,一种免训练的随机注意力方法,通过代表性 key 进行内存高效的选择,无需扫描整个 KV cache。缓存的 key 被组织成若干 team,query 对每个 team 中的代表性 key 打分以决定采样哪些 team。我们在采样得到的 team 内计算精确的注意力分数,并通过其采样概率的倒数对各 team 的贡献进行重新加权。Attention often concentrates on a small subset of tokens in the context, but which subset matters changes from one query to the next. To exploit this changing structure, we introduce SANTA++, a training-free stochastic attention method that uses representative keys for memory-efficient selection without scanning the entire key-value (KV) cache. Cached keys are organized into teams, and the query scores one representative from each team to decide which teams to sample. We compute exact attention scores within the sampled teams and reweight each team's contribution by the inverse of its inclusio

Learning from Teacher Continuations at Student States
在学生状态处从教师续写中学习
arXiv:2609.36246 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

OLIVE 通过从持续演化的 student 生成前缀,在离线蒸馏进入平台期后继续提升,同时更好地保持了 student 的通用能力与可塑性。By regenerating prefixes from the evolving student, OLIVE continues improving after offline distillation plateaus while better preserving the general capabilities and plasticity of the student.

Can We Trust the Teacher? Decoupled Credit Direction-Magnitude for Self-Distillation
我们能信任教师吗?面向自蒸馏的解耦信用方向-幅度
arXiv:2609.34848 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出解耦信用自蒸馏,理论上将信用方向与幅度解耦为两个可靠信号,并据此校准特权教师监督,从而完成策略优化的 step-to-token 信用分配。Decoupled Credit Self-Distillation is introduced, which theoretically decouples credit direction and magnitude into two reliable signals and uses them to calibrate privileged teacher supervision, enabling step-to-token credit assignment for policy optimization.

Selecting Diverse SFT Traces Improves Post-RL Generalization
选择多样化的 SFT 轨迹可提升 RL 后的泛化能力
arXiv:2609.33780 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,推理路径多样性可作为筛选 SFT 数据的实用准则,能更好地为 RL 准备模型,并据此提出一种轻量级、基于规则的指纹方法用于筛选。These results identify reasoning-route diversity as a practical criterion for selecting SFT data that better prepares models for RL, and propose a lightweight, rule-based fingerprint to select for it.

NVAlign: Direct-Gradient Optimization for Non-Verbal Control in Continuous Autoregressive Flow Matching Text-to-Speech
NVAlign:面向连续自回归流匹配文本到语音系统中非言语控制的直接梯度优化
arXiv:2609.31892 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

NVV-SuperBench 与人工听力评测的结果显示,NVAlign 在标签跟随准确率上优于 SFT 与 Flow-GRPO 基线,证明直接对奖励梯度进行优化可提升连续自回归流匹配 TTS 中的非语言控制能力。Results from NVV-SuperBench and human listening evaluations show that NVAlign improves tag-following accuracy over SFT and Flow-GRPO baselines, demonstrating that direct reward-gradient optimization can improve non-verbal control in continuous autoregressive flow-matching TTS.

Towards Semi-Automatically Comparing Keyword-Based and Semantic Search Accuracy
Towards Semi-Automatically Comparing Keyword-Based and Semantic Search Accuracy
arXiv:2609.37749 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一个新颖的初步框架,可定量评估输出格式不同(如列表与消息)的检索系统的 IR 准确度,为客观评估基于关键词与基于语义的对话式检索方法奠定坚实基础。This work introduces a novel, preliminary framework to quantitatively assess IR accuracy of search systems that produce different output formats, such as lists and messages, and provides a strong foundation for objectively assessing keyword-based and semantic chat-based search methods.

Context Language Models
Context Language Models
arXiv:2609.37725 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Context Language Models,能原生管理自身上下文,并可通过经标准 skill 优化循环演化的自然语言指令进行引导,在上下文管理任务上将未见数据的准确率最高提升 35.9 分,同时降低计算开销。Context Language Models are introduced, language models that natively manage their own context and can be steered with natural-language instructions evolved through a standard skill-optimization loop, improving held-out accuracy by up to 35.9 points on a context-management task while reducing compute.

Multimodal 补充候选
arXiv:2606.13578 多模态 方法 OA · 绿色 被引 4 · S2

构建了 RoboGenesis,一个基于仿真的工作流与数据引擎,可从原子技能组合配置好的实验工作流,对 rollout 进行验证与过滤,并跨支持的机器人配置导出结构化演示数据。RoboGenesis is built, a simulation-based workflow and data engine that composes configured laboratory workflows from atomic skills, validates and filters rollouts, and exports structured demonstrations across supported robot profiles.

Chinese-Jev: Bringing System One Model to Chinese-Language Tasks
Chinese-Jev: Bringing System One Model to Chinese-Language Tasks
arXiv:2609.36965 工程化 方法 被引 0 · S2

提出 Chinese-Jev,一个 System One 模型,通过统一的数据处理与训练流水线弥合预训练分布与下游中文场景间的差距,在各专业领域上达到 Jev 平均准确率的 92%。Chinese-Jev is introduced, a System One model that addresses the gap between the pre-training distribution and downstream Chinese-language scenarios through a unified data processing and training pipeline, and achieves 92% of Jev's average accuracy across specialized domains.

Follow the Entities: A Corpus Map for Agentic Search
[标题中文] 跟随实体:面向 Agentic 搜索的语料库地图
arXiv:2609.37226 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 CorpusMap,一个以语料中重复出现的实体为核心的导航层;这些实体可从文档自身识别,并能在不同来源间把单个文档与众多其他文档相连,表明实体可作为大型文档集合导航的有效锚点。CorpusMap is introduced, a navigation layer that organizes the corpus around its recurring entities, which are identifiable from the documents themselves and can link a single document to many others across sources, suggesting that entities serve as effective anchors for navigating large document collections.

Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection
[标题中文] 相同字节,不同权威:Chat 模板 Prompt 注入中的保留 token 表示
arXiv:2609.35932 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

在所测试的每一对基座与指令微调模型中,指令微调都强化了模型对保留标记的偏好,且该差距在该通道上持续存在。In every base and instruction-tuned pair the authors test, instruction tuning strengthens the model's preference for reserved markers, and the gap persists on that channel.

Pretraining Transformers with Quantized Softmax in Attention
[标题中文] 使用量化 Softmax 的注意力机制预训练 Transformer
arXiv:2609.33591 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文推导了对应的反向传播规则(包括校准项的导数),并在模型、数据、优化器均一致的预训练实验中对不同选择进行了比较。This work derives the corresponding backward rules, including calibration derivatives, and compares these choices in pretraining experiments matched on model, data, and optimizer, and compares these choices in pretraining experiments matched on model, data, and optimizer.

Omni-IO Skills: Harnessing Your Agent Omni-Native
[标题中文] Omni-IO Skills:让你的 Agent 原生支持全模态
arXiv:2609.31847 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,harness 层面的能力组合是构建广泛、可演进 Omni 系统的实用路径,且无需改变宿主 agent 的推理核心。The results establish harness-level capability composition as a practical route to broad, evolvable Omni systems without changing the host agent's reasoning core.

Omni-Decision: Evidence-Ledger Planning for Omni-Modal Agents
[标题中文] Omni-Decision:面向全模态 Agent 的证据账本式规划
arXiv:2607.11433 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Omni-Decision,一个基于证据账本规划的 omni-modal agent:用显式的证据账本替代不断膨胀的对话历史,记录仍缺失的证据、已确认的内容以及记录间的冲突。Omni-Decision is presented, an omni-modal agent built on evidence-ledger planning: it replaces the growing dialogue history with an explicit evidence ledger that records what evidence is still missing, what has been confirmed, and where records conflict.