DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

  • 类型:arxiv
  • 标识:2609.19969
  • 链接:https://arxiv.org/abs/2609.19969
  • 主分类:llm-infra
  • 形态:application
  • TLDR:The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transfer bandwidth. Together, these compute, storage, and bandwidth demands constitute the primary bottleneck to further lowering deployment costs. To address this challenge, we introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for context
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-18-agent-rag-longcontext-candidates.json