研究库 论文知识库
Papers · organized/paper_cards

论文

1094 张论文卡片 · 方法 · OA 绿色

开放获取 全部 绿色 · 1640
FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
FoMo:生成轨迹中的分叉时刻作为感知距离
arXiv:2609.25716 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一条全自动的数据生成流程,可在无需任何人工标注的情况下生成图像对之间的逐点感知距离标签,并证明 diffusion trajectory 与人类视觉系统高度一致。This paper proposes a fully automated data generation pipeline that generates pointwise perceptual distance labels between image pairs without any human annotation, and demonstrates that the diffusion trajectory aligns well with the human visual system.

Do Implicit Personalization and Explicit Styles Conflict? PsPLUG: A Lightweight Plug-in for Balancing Personalization and Style in Customized LLMs
[标题中文] 隐式个性化与显式风格是否冲突?PsPLUG:用于在定制化 LLM 中平衡个性化与风格的轻量级插件
arXiv:2601.06362 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

本文提出 PsPLUG,一个轻量级插件式模块,在刻画目标风格后学习用户专属的残差,能更好地保留用户偏好,并对个性化与风格遵循之间的平衡提供更精确的控制。PsPLUG is proposed, a lightweight plug-in that learns a user-specific residual after accounting for the requested style, and better preserves user preferences while providing precise control over the balance between personalization and style adherence.

Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching
[标题中文] 通过连续深度批处理实现循环语言模型的深度自适应推理
arXiv:2608.09444 LLM 基础设施 方法 OA · 绿色 被引 5 · S2

本文提出首个通过 continuous depth batching (CDB) 实现 depth-adaptive looped LM 的高效方法:在 loop 步骤之间重组 batch,动态调度架构中的 looped 与非 looped 部分,管理 looped KV-caching,并提前预测将退出 loop 的 token,以便异步准备 batch。This work introduces the first efficient method for depth-adaptive looped LMs via continuous depth batching (CDB), which forms new batches between loop steps, and dynamically schedules looped and non-looped parts of the architecture, manages looped KV-caching, and predicts which tokens will exit the loop in advance so it can prepare batches asynchronously.

Systems 补充候选
arXiv:2511.02230 Agent 智能体 方法 Open MIND OA · 绿色 被引 68 · S2

Continnum,一种通过为 KV cache 保留引入 TTL 机制来优化多轮 Agent 工作负载任务完成时间的服务系统,与程序级 FCFS 结合时可保持多轮连续性,并降低 Agent 工作流的延迟。Continnum, a serving system to optimize job completion time for multi-turn agent workloads by introducing time-to-live mechanism for KV cache retention, and when combined with program-level first-come-first-serve preserves multi-turn continuity, and reduces delay for agentic workflows.

Paragraph Boundaries Are Not White Space:Compression Depth as the Signature of Hierarchical Structure
[标题中文] 段落边界并非空白:压缩深度作为层级结构的特征
arXiv:2609.23551 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

使用一种分层旋转位置编码(hRoPE),将段落、句子和 token 索引表示为独立通道,保持 token 序列固定,对段落坐标 $p_1$ 进行干预,并使用 token 距离精确估计器测量跨段落注意力。A hierarchical rotary positional encoding (hRoPE) that represents paragraph, sentence, and token indices as separate channels, hold the token sequence fixed, intervene on the paragraph coordinate $p_1$, and measure cross-paragraph attention with a token-distance-exact estimator is used.

BoundInk: Boundary-Aware Online Handwriting Generation
[标题中文] BoundInk:面向边界感知的在线手写生成
arXiv:2604.02103 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

BoundInk 是一种以书写者为条件的框架,将字符间的边界视为显式生成单元,既能保留书写者特有的字形外观,又能在完整文本行内改善连接性与字距。BoundInk is introduced, a writer-conditioned framework that treats inter-character boundaries as explicit generation units and preserves writer-specific glyph appearance while improving connectivity and spacing across complete text lines.

VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models
[标题中文] VLA-Precision:面向视觉-语言-动作模型高效真实世界在线强化学习的非对称协同自举
arXiv:2609.04355 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 VLA-Precision,一个面向真实世界的高效在线 RL 框架,包含 Asymmetric Co-Bootstrapping (ACoB) 算法与 ACoB-Stream 架构;其中 ACoB-Stream 以不变状态解耦与按需流式传输为设计原则,构建了经验-策略的闭环架构,可实现最高 10.9% 的吞吐与计算效率提升。VLA-Precision, an efficient real-world online RL framework featuring the Asymmetric Co-Bootstrapping (ACoB) algorithm and the ACoB-Stream architecture, and ACoB-Stream, a closed-loop experience--policy architecture that establishes invariant-state decoupling and on-demand streaming as design principles, deliver up to 10.9% improvements in throughput and computational efficiency.

QReason: Query-Focused Decoupled Chain-of-Thought for Efficient Passage Reranking
QReason:面向高效段落重排序的查询聚焦解耦思维链。
arXiv:2609.30904 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

QReason 是一种解耦框架,将面向 query 的推理与针对窗口的 passage 相关性评估分离,显著减少冗余推理,在取得与 reasoning-based reranker 相当乃至更优的排序性能的同时,超越了现有的 query rewriting 模型。QReason is a decoupled framework that separates query-focused reasoning from window-specific passage relevance assessment, and significantly reduces redundant reasoning, achieves ranking performance comparable to or better than strong reasoning-based rerankers, and outperforms existing query rewriting models.

Intent2Tc: Automated Intent-to-Traffic Control Translation with Language Models
Intent2Tc:基于语言模型实现意图到流量控制的自动化翻译。
arXiv:2609.31397 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

Intent2Tc 是一个由语言模型驱动的闭环框架,将业务级流量整形意图转换为声明式子意图,进而生成经过验证的可执行 Linux traffic control (tc) 配置,并展示了该框架的实际适用性。Intent2Tc is presented, a closed-loop language-model-driven framework that translates business-level traffic-shaping intents into declarative sub-intents and subsequently into validated, executable Linux traffic control (tc) configurations and demonstrates the practical applicability of the proposed framework.

Softmax Reparameterization for Output-Head Quantization
面向输出头量化的 Softmax 重参数化。
arXiv:2609.31291 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 softmax reparameterization,一种训练后方法,在量化前搜索功能等价的输出头,并展示了在总体 logit 误差增大的情况下保真度仍可提升。This work introduces softmax reparameterization, a post-training method that searches over functionally equivalent output heads before quantization and shows how fidelity can improve despite greater total logit error.

LightMIS: Ultra-Lightweight Medical Image Segmentation Without a Stage-Wise Decoder
LightMIS:无逐级解码器的超轻量医学图像分割。
arXiv:2609.28327 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

LightMIS 通过 Scale-Aligned Projection 块将五级编码器的输出对齐至统一分辨率,仅聚合一次,并利用结合 Adaptive Kernel Fusion 与所提 Progressive Receptive Fusion 模块的 Adaptive Fusion Cascade 精炼融合表示。LightMIS aligns the outputs of a five-level encoder to a common resolution using Scale-Aligned Projection blocks, aggregates them once, and refines the fused representation with an Adaptive Fusion Cascade, which combines Adaptive Kernel Fusion with the proposed Progressive Receptive Fusion module.

Systems 补充候选
arXiv:2606.01751 RAG 检索增强 方法 OA · 绿色 被引 3 · S2

SarseX 模型无关、无需训练,并与 Prefix Cache 兼容,可为多轮对话、检索增强生成 (RAG) 和 Agent 工作流等常见在线服务场景提供统一支持。SarseX is model-agnostic, training-free, and compatible with Prefix Cache, and it provides unified support for common online serving scenarios including multi-round chat, retrieval-augmented generation (RAG), and agent workflows.

BoundaryMORPH: Budgeted Reranking via Active Set Selection for Diffuse Retrieval
BoundaryMORPH:通过主动集合选择实现预算化重排序的扩散检索
arXiv:2609.27213 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

BoundaryMORPH 是一种新算法,专门为 LLM 的上下文容量 k 分配 CE 预算,在多个模型与数据集上针对开放性查询取得了 SOTA 集合检索质量。BoundaryMORPH is introduced, a novel algorithm that allocates CE budget specifically for the LLM's context capacity $k$ and achieves state-of-the-art set retrieval quality across multiple models and datasets with open-ended queries.

SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance
SAGE:通过拓扑引导缓解长程推理偏置
arXiv:2609.30192 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 SAGE(Structural Admissibility-Guided Exploration),一个通过注入结构引导来缓解长程推理中探索偏差与累积偏差的统一框架,在 Andrews-Curtis 问题上取得了最高 8 倍的提升。This work proposes SAGE (Structural Admissibility-Guided Exploration), a unified framework that injects structural guidance to alleviate exploration bias and compounding bias in long-horizon reasoning and achieves up to an 8-fold improvement on the Andrews-Curtis problem.

LastOPD: Taming Collapse in Latent On-Policy Distillation
LastOPD:驯服潜在 On-Policy 蒸馏中的坍缩问题
arXiv:2609.28845 多模态 方法 OA · 绿色 被引 2 · S2

LastOPD 仅在最后一层状态(LM head 共同读取的接口)施加潜在信号,且仅在进入 token 级 OPD 前的 10 步交叉淡入阶段施加,从而保留潜在信号中的有用部分,并在坍缩发生前将 student 交由 token 级监督。LastOPD is proposed, which applies the latent signal only at the last-layer state, the common interface both LM heads read, and only during a 10-step crossfade into token-level OPD, which keeps the useful part of the latent signal and hands the student to token-level supervision before the collapse sets in.

TimeEvo: Failure-Driven Self-Evolution of a Time Series Agent
TimeEvo:时序 Agent 的失败驱动自演化
arXiv:2609.27277 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

TimeEvo 将 Agent 诊断出的失败聚类为能力缺口,为每个缺口规划测量,合成只产出证据的工具来填补缺口,并仅通过配对的准入门控接纳候选工具库。TimeEvo is proposed, which clusters an agent's diagnosed failures into capability gaps, plans a measurement for each, synthesizes evidence-only tools that fill them, and admits the candidate library only through a paired admission gate.

Disaggregated Quantization: Specializing LLM Prefill and Decode
解耦量化:针对 LLM Prefill 与 Decode 的专用量化
arXiv:2609.26333 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

评估 vLLM 在分离式 serving 下的精度,并通过训练后量化在最高 2.8T 参数的模型上进一步验证共享权重格式的分离式 serving。Assessment of accuracy under disaggregated serving in vLLM and further validate shared-weight format disaggregation through post-training quantization on models up to 2.8T parameters are evaluated.

D-JEPA: A Decision-Aligned Latent World Model
D-JEPA:一种决策对齐的潜在世界模型
arXiv:2609.24749 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

D-JEPA 是一种决策对齐的潜在世界模型,从已执行结果中学习候选未来之间的决策相关关系,将决策相关关系结构确立为预测世界建模与有效控制之间的直接桥梁。D-JEPA is introduced, a decision-aligned latent world model that learns decision-relevant relations among candidate futures from executed outcomes, establishing decision-relevant relational structure as a direct bridge between predictive world modeling and effective control.

Just-In-Time Agent Memory with Runtime Agentic Research
基于运行时 Agentic 研究的 Just-In-Time Agent 记忆。
arXiv:2609.34385 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Just-In-Time Agent Memory (JAM) 是一个可训练框架,在运行时进行查询条件下的上下文构建,其任务性能优于 AOT 式记忆系统,同时比先前的可训练 Agent 记忆方法显著更高效。Just-In-Time Agent Memory (JAM), a trainable framework for query-conditioned context construction at runtime, is proposed, where it achieves stronger task performance than AOT-style memory systems while remaining substantially more efficient than prior trained agentic memory approaches.

Stashbird: Efficient Speaker-Indexed Memory for Conversational Agents
Stashbird:面向对话 Agent 的高效说话人索引记忆。
arXiv:2609.34242 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Stashbird 是一种 Agent 记忆系统,通过显式来源溯源将源事件链接到派生记忆状态,在 LongMemEval-S 和 GroupMemBench 上准确率高于 Hindsight,在 EverMemBench 上与之相当。Stashbird is presented, an agent memory system that links source episodes to derived memory state through explicit provenance and achieves higher accuracy than Hindsight on LongMemEval-S and GroupMemBench and comparable accuracy on EverMemBench.

InfiniHand: Streaming World-Space Hand Motion Estimation from Egocentric Video
InfiniHand:基于第一人称视频的流式世界空间手部运动估计。
arXiv:2609.35743 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

InfiniHand 是一个端到端流式前馈框架,联合估计 MANO 参数、相机轨迹与手部位置,直接从未标定的自我中心视频出发,相比 ViDiHand 在 ARCTIC PA-p 上降低 21%,并显著缓解世界空间漂移。InfiniHand is an end-to-end streaming feed-forward framework that jointly estimates MANO parameters, camera trajectories, and hand locations directly from uncalibrated egocentric video, achieving a 21% reduction in ARCTIC PA-p compared to ViDiHand while substantially mitigating world-space drift.

GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space
GeoVerse: 几何潜空间中的世界一致新视角合成
arXiv:2609.35734 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

GeoVerse 是一个在世界一致条件下合成新视角的框架,在预训练 3D 基础模型的几何潜在空间内进行生成,并通过 ControlNet 风格 adapter 注入视频生成模型的外观先验。GeoVerse is a framework that synthesizes world-consistent novel views by performing generation within the geometric latent space of a pretrained 3D foundation model and injecting appearance priors from a video generative model via a ControlNet-style adapter.

Systems 补充候选
arXiv:2606.03910 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

NetKV,一种使用该 oracle 信息的 O(|D|) 每请求贪心策略,其层级排序被证明对过时遥测数据具有鲁棒性;并证明随着上下文长度增长,忽略网络项会使仅缓存感知的调度任意次优。NetKV, the O(|D|) per-request greedy that consumes this oracle, has tier rankings that are provably robust to stale telemetry, and it is proved that ignoring the network term renders cache-aware-only scheduling arbitrarily suboptimal as context length grows.

Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning
通过交错强化学习在统一模型中学习原生反思
arXiv:2609.35767 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

UMM-Reflection 在统一模型内利用强化学习(RL)完成完整反思轨迹:兄弟轨迹共享同一初始图像,因此组相对优势可比较不同反思策略,而单条轨迹级优势同时更新反思 token 与基于 flow 的修订,避免了逐轮信用分配的组合爆炸。UMM-Reflection is introduced, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and the flow-based revisions, avoiding the combinatorial blow-up of per-round credit assignment.

When Do Model Internals Help? Exploring the Role of Representation Engineering in LLM Safety
模型内部何时有用?探究表征工程在 LLM 安全中的作用
arXiv:2609.34771 安全与风险 方法 OA · 绿色 被引 0 · S2 + OpenAlex

总体而言,representation engineering 并不能普遍替代行为安全护栏,但在特定条件下具有实际优势,并可带来互补的安全收益。Overall, representation engineering does not generally replace behavioral safeguards, but offers practical advantages under specific conditions and can provide complementary safety benefits.

WideSWE: Can Coding Agents Coordinate Changes Across Repositories?
WideSWE: 编码 Agent 能否跨代码仓库协调变更?
arXiv:2609.33382 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 WideSWE,用于在跨仓库任务上评估编码 Agent,并在相同 prompt 下将其与联合执行进行对比,以考察逐仓库工作是否能缓解相关困难。This work introduces WideSWE to evaluate coding agents on cross-repository tasks, and compares it with joint execution under identical prompts to examine whether working on one repository at a time can alleviate difficulties.

Residual Transferability in Neural Image Watermarking
神经图像水印中的残差可迁移性
arXiv:2609.32241 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文识别出两种强化水印证据对载体图像依赖性、抑制残余可迁移性的机制,并提出 CoverLock,一种即插即用策略,可在不重新设计架构的前提下强化现有水印系统的图像依赖性。This work identifies two mechanisms that strengthen the dependence of watermark evidence on the cover image, thereby suppressing the residual transferability, and introduces CoverLock, a plug-and-play strategy for existing watermarking systems that strengthens such image dependence without architectural redesign.

FlowTool: Controlling Tool Parameter in Image Retouching via Flow Matching
FlowTool: 基于流匹配控制图像修图中的工具参数
arXiv:2609.35673 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 FlowTool,一个通过 conditional rectified flow 直接建模以输入图像和用户指令为条件的高质量工具参数分布的框架,并显著提升了推理效率。This work introduces FlowTool, a framework that directly models the distribution of high-quality tool parameters conditioned on the input image and user instruction using conditional rectified flow, and significantly improves inference efficiency.

Nereus: Adaptive Parallelism for LLM Post-Training
Nereus: 面向 LLM 后训练的自适应并行
arXiv:2609.34645 工程化 方法 OA · 绿色 被引 1 · S2

Nereus 是一种成本感知的运行时,将 RL 后训练任务适配为高效执行计划,并基于内存可行的全局计划执行状态转移,使用与运行任务校准的成本模型来接纳转移。Nereus is a cost-aware runtime that adapts RL post-training jobs into efficient execution plans and executes a transition using a memory-feasible global plan and admits the transition using a cost model calibrated against the running job.

DISCO: Distributed Long Context Scaling with Grounding-Reasoning Disaggregation
DISCO: 基于接地-推理解耦的分布式长上下文扩展
arXiv:2609.33485 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

受 Apache Spark 等分布式计算框架启发,DISCO 将长上下文切分到一组专职 Worker LLM 上,进行并行、局部的 grounding,建立了鲁棒长上下文推理的高效范式。Inspired by distributed computing frameworks like Apache Spark, DISCO partitions long context across a fleet of Worker LLMs dedicated exclusively to parallel, localized grounding, establishing a highly efficient paradigm for robust long-context inference.

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
面向代码 Agent 强化学习的分组智能评分与优势再分配
arXiv:2609.32577 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 GAGAR,一个面向代码 agent RL 的质量感知信用再分配框架;结果表明,将基于测试的验证与成组的 agentic 评分相结合,可提升代码 agent RL 的质量与稳定性。GAGAR, a framework for quality-aware credit redistribution in code agent RL, is introduced and results support combining test-based verification with groupwise agentic grading to improve the quality and stability of code agent RL.

Relic: From Multi-Agent Collaboration to Persistent Organizational Capability
Relic:从多 Agent 协作到持久化的组织能力
arXiv:2609.32965 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 Relic,它将反复出现的协作失败转化为组织拥有的、可执行的协议,使协作经验成为超越其创建者仍持续有用的组织级持久状态。The introduction of Relic, which turns recurring collaboration failures into organization-owned, executable protocols, and how collaboration experience can become persistent organizational state that remains useful beyond the members who created it are shown.

Systems 补充候选
arXiv:2510.09665 LLM 基础设施 方法 OA · 绿色 被引 174 · S2

本工作提出 LMCACHE,首个也是目前最高效的开源 KV 缓存方案,可将现代 LLM 引擎生成的 KV 缓存从 GPU 显存中提取并存储,并跨引擎和查询共享。This work presents LMCACHE, the first and so far the most efficient open-source KV caching solution, which extracts and stores KV caches generated by modern LLM engines out of the GPU memory and shares them across engines and queries.

G^2PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation
G^2PTQ:通过广义梯度补偿改进 LLM 训练后量化
arXiv:2609.31009 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 G²PTQ,一个具有广义梯度补偿的统一 PTQ 框架,在全局监督的分块优化目标下融合一阶与二阶信息,能更好地对齐全精度模型,性能优于 SOTA 基线。G$^2$PTQ is presented, a unified PTQ framework with Generalized Gradient Compensation that integrates both first- and second-order information under a globally supervised, block-wise optimization objective and enables better alignment with the full-precision model, outperforming state-of-the-art baselines.

SANTA++: Sampling Attention through Representative Keys
SANTA++: 通过代表性 Key 进行采样的注意力机制
arXiv:2609.35629 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

注意力往往集中在上下文中一小部分 token 上,但每个 query 关注的关键子集各不相同。为利用这种动态结构,我们提出 SANTA++,一种免训练的随机注意力方法,通过代表性 key 进行内存高效的选择,无需扫描整个 KV cache。缓存的 key 被组织成若干 team,query 对每个 team 中的代表性 key 打分以决定采样哪些 team。我们在采样得到的 team 内计算精确的注意力分数,并通过其采样概率的倒数对各 team 的贡献进行重新加权。Attention often concentrates on a small subset of tokens in the context, but which subset matters changes from one query to the next. To exploit this changing structure, we introduce SANTA++, a training-free stochastic attention method that uses representative keys for memory-efficient selection without scanning the entire key-value (KV) cache. Cached keys are organized into teams, and the query scores one representative from each team to decide which teams to sample. We compute exact attention scores within the sampled teams and reweight each team's contribution by the inverse of its inclusio

Learning from Teacher Continuations at Student States
在学生状态处从教师续写中学习
arXiv:2609.36246 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

OLIVE 通过从持续演化的 student 生成前缀,在离线蒸馏进入平台期后继续提升,同时更好地保持了 student 的通用能力与可塑性。By regenerating prefixes from the evolving student, OLIVE continues improving after offline distillation plateaus while better preserving the general capabilities and plasticity of the student.