研究库 论文知识库
Papers · organized/paper_cards

论文

10 张论文卡片 · LLM 基础设施 · 观点 · OA 绿色

开放获取 全部 绿色 · 1640
5️⃣ Multi-Segment Attention · 分块位置感知KV驱逐 — arXiv:2606.02964(⭐⭐⭐ 新鲜 arXiv)
arXiv:2606.02964 LLM 基础设施 观点 OA · 绿色 被引 2 · S2

AsymCache 是一个面向 LLM 推理的计算-延迟感知 KV cache 管理系统,将 cache 驻留决策与 GPU attention kernel 性能显式对齐,包含三个关键组件:用于高效处理非连续 KV 上下文的多段注意力(MSA)、联合优化命中率与位置感知重计算代价的 cache 淘汰策略,以及面向高硬件利用率的自适应分片调度器。AsymCache is proposed, a computation-latency-aware KV cache management system for LLM inference that explicitly aligns cache residency decisions with GPU attention kernel performance, including three key components: Multi-Segment Attention (MSA) for efficient non-contiguous KV context processing, a cache eviction policy that jointly optimizes hit rate and position-aware recomputation cost, and an adaptive chunking scheduler for high hardware utilization.

7️⃣ arXiv · Position Paper:LLM Serving 需要数学优化,而非仅靠启发式 ⭐⭐⭐⭐⭐ 学术前沿
arXiv:2605.01280 LLM 基础设施 观点 OA · 绿色 被引 1 · S2

这篇立场论文认为,LLM 推理 serving 已超越通用启发式方法,如今需要数学优化与算法基础,并呼吁社区将 LLM serving 的算法设计视为一个新的研究前沿。This position paper argues that LLM inference serving has outgrown generic heuristics and now demands mathematical optimization and algorithmic foundations, and calls on the community to recognize algorithmic design for LLM serving as a research frontier.

What Does an Evaluation License? A Commit-Bound Census of Claim-Relative Inference in Inspect Evals
评估授权了什么?Inspect Evals 中声明相对推理的提交绑定普查
arXiv:2608.19269 LLM 基础设施 观点 OA · 绿色 被引 0 · S2 + OpenAlex

冻结一个大规模评测集合,并尝试对其历史断言进行 replay,以使这一原本隐式的推理步骤变得显式且可执行。A large evaluation collection is frozen and an attempt to replay its historical claims is made to make this otherwise implicit inference step explicit and executable.

KVShareArena: KV-Cache Reuse Across Contexts and Model Checkpoints
KVShareArena:跨上下文与模型 checkpoint 的 KV cache 复用
arXiv:2609.10266 LLM 基础设施 观点 OA · 绿色 被引 0 · S2 + OpenAlex

揭示了复用造成的质量损失以及有效修复方法均取决于 LLM(即便在两个 8B 模型之间也是如此),并提供添加新方法的通用接口与交互式 leaderboardBoth the quality loss from reuse and which repairs help depend on the LLM, even between two 8B models are introduced, as well as a common interface for adding new methods and an interactive leaderboard.

Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction
Grouped Value Attention:通过按需 Key 重建实现高效 KV 缓存
arXiv:2609.13285 LLM 基础设施 观点 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出了分组值注意力(Grouped Value Attention, GVA),通过存储分组值并在推理时利用可学习的线性映射重构内容键,该映射可被吸收进 query 中,从而在预期解码路径中无需显式物化内容键。This work introduces Grouped Value Attention (GVA), which stores grouped values and reconstructs content keys with a learned linear map at inference, which can be absorbed into the query, eliminating the need to materialize content keys in the intended decode path.

The Linear Representation Hypothesis Needs a Group Action
线性表示假说需要一个群作用
arXiv:2609.27158 LLM 基础设施 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文论证线性表征假设不是单一假设,而是一组通过表征等价性加以区分的命题族;该框架厘清了假设如何随度量、读取点与分析阶段变化,并被用于审计常见表征量与近期可解释性分析。It is argued that the Linear Representation Hypothesis is not one hypothesis but a family of claims distinguished by representation equivalence, and this framework clarifies how assumptions can change across metrics, reading points, and analysis stages, and is used to audit common representation quantities and recent interpretability analyses.

Evaluating the accuracy of KV cache reuse techniques
评估 KV cache 复用技术的准确性。
arXiv:2609.31415 LLM 基础设施 观点 OA · 绿色 被引 0 · S2 + OpenAlex

Boxoffice 是一个程序化生成评测数据集的工具,用以覆盖具有挑战性的 KV cache 复用模式;研究表明现有数据集并未展现出充分评估此类技术所需的复用动态特性。Boxoffice, a tool that programmatically generates evaluation datasets that exercise challenging KV cache reuse patterns, is introduced and it is shown that existing datasets do not exhibit the reuse dynamics needed to thoroughly evaluate such techniques.

Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression
周期性弱点:分块 KV-Cache 压缩带来的相位敏感性
arXiv:2609.36322 LLM 基础设施 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文揭示使用 chunked KV-cache 压缩的模型中存在一种系统性不对称:相同信息在一个阶段易于检索,在另一阶段却难以获取,暴露出平均 benchmark 分数可能掩盖的周期性薄弱点。A systematic asymmetry in models using chunked KV-cache compression is uncovered: the same information can be easy to retrieve at one phase and difficult at another, revealing periodic weak spots that average benchmark scores can conceal.

LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling
LoopCoder-v2:仅循环一次以实现高效测试时计算扩展
arXiv:2606.18023 LLM 基础设施 观点 OA · 绿色 被引 4 · S2

本文通过收益–成本视角研究 PLT 循环次数选择:额外循环可精炼表示,但 CLP 也会在每次循环边界引入位置错配,由此解释 PLT 在两次循环时趋于饱和的现象,并为循环次数选择提供诊断依据。This study studies PLT loop-count selection through a gain--cost view: an extra loop may refine representations, but CLP also introduces a positional mismatch at each loop boundary, explaining PLT's saturation at two loops and providing diagnostics for loop-count selection.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management
HiSparse:通过分层 KV 缓存管理扩展稀疏注意力解码
arXiv:2608.07009 LLM 基础设施 观点 OA · 绿色 被引 4 · S2

HiSparse 已合入上游 SGLang,并在 H200、B200 与 GH200 平台上针对 DSA、NSA、Quest 三类稀疏注意力家族进行评估:在长上下文负载下峰值生成吞吐提升最高达 4.7×,同时保持可比的 per-token 延迟,并降低高负载下的 time-to-first-token。HiSparse is merged into upstream SGLang and evaluated across three sparse-attention families (DSA, NSA, and Quest) on H200, B200, and GH200 platforms: it improves peak generation throughput by up to 4.7x on long-context workloads while preserving comparable per-token latency and reducing time-to-first-token at high load.