intel/neural-compressor

  • 类型:github
  • 标识:intel/neural-compressor
  • 链接:https://github.com/intel/neural-compressor
  • 主题:llm-infra
  • 主分类:llm-infra
  • 形态:model
  • 分类:ai
  • Stars:2697
  • 周增:+0
  • 语言:Python
  • 许可:Apache-2.0
  • 最近提交:2026-08-11
  • 简介:SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity; leading model compression techniques on PyTorch, TensorFlow, and ONNX Runtime
  • 上次采集:2026-08-11
  • 首次采集:2026-08-11
  • 待LLM分类:否
  • 成熟度:research
  • 简介中文:SOTA 低比特 LLM 量化(INT8/FP8/MXFP8/INT4/MXFP4/NVFP4)与稀疏化方案;面向 PyTorch、TensorFlow 与 ONNX Runtime 的领先模型压缩技术
  • 来源文件
  • [GitHub Search]