一个 2.78 万亿参数的 Kimi K3,在仅 8.24 GB 内存的单颗 CPU 上运行推理。可移植的 C99:无需 BLAS,无需框架,无需 GPU。A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
仓库/Skill 库
263 个 · LLM 基础设施
精选的 LLM、AI 绘画等领域的教程和资源。Curated tutorials and resources for Large Language Models, AI Painting, and more.
高质量且极速的预训练深度学习模型与 demo。Pre-trained Deep Learning models and demos (high quality and extremely fast)
针对加速基础设施和微服务架构优化的生成式 AI 参考工作流。Generative AI reference workflows optimized for accelerated infrastructure and microservice architecture.
Fengshenbang-LM(封神榜大模型)是IDEA研究院认知计算与自然语言研究中心主导的大模型开源体系,成为中文AIGC和认知智能的基础设施。
ChatGPT、Claude 等 LLM 的所有前端 GUI 客户端汇总。Every front-end GUI client for ChatGPT, Claude, and other LLMs
任意模型,任意硬件,零妥协。基于 @ziglang / @openxla / MLIR / @bazelbuild 构建。Any model. Any hardware. Zero compromise. Built with @ziglang / @openxla / MLIR / @bazelbuild
🚀 通过易用的硬件优化工具,加速 🤗 Transformers、Diffusers、TIMM 和 Sentence Transformers 的推理与训练。🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools
Apple Silicon 上最快的本地 AI 引擎。比 Ollama 快 4.2 倍,缓存 TTFT 仅 0.08s,工具调用支持率 100%。内置 17 种工具解析器、prompt cache、推理分离、云端路由。可作为 OpenAI 的即插即用替代,兼容 Claude Code、Cursor、Aider。The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
在任意消费级硬件上运行 Qwen3.8-Flash-Next:Windows / Linux 一键安装。Strata 推理引擎,本地 localhost 提供 OpenAI / Anthropic API,支持可选图像输入。Qwen3.8-Flash-Next on any consumer hardware: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input.
📚 OpenVINO™ 的 Jupyter notebook 教程。📚 Jupyter notebook tutorials for OpenVINO™
AgentENV (AENV) 是一个用于大规模运行 Agent 环境的分布式平台。AgentENV (AENV) is a distributed platform for running agent environments at scale.
PyTorch 原生的量化和稀疏化方案,支持训练与推理PyTorch native quantization and sparsity for training and inference
SOTA 低比特 LLM 量化(INT8/FP8/MXFP8/INT4/MXFP4/NVFP4)与稀疏化方案;面向 PyTorch、TensorFlow 与 ONNX Runtime 的领先模型压缩技术SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity; leading model compression techniques on PyTorch, TensorFlow, and ONNX Runtime
为 vLLM 在 Ascend 上的社区维护硬件插件Community maintained hardware plugin for vLLM on Ascend
基于 Rust 的 ONNX 模型快速 ML 推理与训练Fast ML inference & training for ONNX models in Rust
一个简洁、高性能且可扩展的 Jax LLM!A simple, performant, and scalable Jax LLM!
高效的浮点神经网络推理算子,覆盖移动端、服务端和 WebHigh-efficiency floating-point neural network inference operators for mobile, server, and Web
将任意电脑或边缘设备打造为计算机视觉项目的指挥中心Turn any computer or edge device into a command center for your computer vision projects.
Bionic 是 ChatGPT 的本地化部署替代方案,在保持严格数据机密性的同时提供生成式 AI 的能力Bionic is an on-premise replacement for ChatGPT, offering the advantages of Generative AI while maintaining strict data confidentiality
🎨 NeMo Data Designer:从零或从种子数据生成高质量合成数据。🎨 NeMo Data Designer: Generate high-quality synthetic data from scratch or from seed data.
通过统一接口对接 LLM 服务商Communicate with an LLM provider using a single interface
基于 Envoy Gateway 构建,提供生成式 AI 服务的统一接入管理Manages Unified Access to Generative AI Services built on Envoy Gateway
超快速 serverless GPU 推理、沙箱和后台任务Ultrafast serverless GPU inference, sandboxes, and background jobs
可扩展的 LLM 数据预处理与清洗工具集Scalable data pre processing and curation toolkit for LLMs
在家用 3090 上由开源模型生成语义化 if。独立项目,与 Jev 或 TypeSafe 无关。Semantic ifs from open models, on a 3090 at home. Independent; not affiliated with Jev or TypeSafe.
面向 AI 辅助草稿的可读性与自然节奏改进的开源 pipeline 与参考实现。Open-source pipeline and reference implementations for improving the readability and natural cadence of AI-assisted drafts.
基于 nvidia-smi 二进制工具的 Nvidia GPU Prometheus exporterNvidia GPU exporter for prometheus using nvidia-smi binary