DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
- 类型:arxiv
- 标识:2609.19969
- 链接:https://arxiv.org/abs/2609.19969
- 主分类:llm-infra
- 形态:application
- TLDR:The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transfer bandwidth. Together, these compute, storage, and bandwidth demands constitute the primary bottleneck to further lowering deployment costs. To address this challenge, we introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for context
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-09-18-agent-rag-longcontext-candidates.json