研究库 开源仓库
Repositories · organized/repo_cards

仓库/Skill 库

263 个 · LLM 基础设施

排序 Stars 周增
FareedKhan-dev/kimi-k3-in-c
C · 2026-08-07 LLM 基础设施 框架 研究原型 Stars 4709 周增 +889

一个 2.78 万亿参数的 Kimi K3,在仅 8.24 GB 内存的单颗 CPU 上运行推理。可移植的 C99:无需 BLAS,无需框架,无需 GPU。A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.

llm-infra
OpenNMT/CTranslate2
C++ · 2026-08-05 LLM 基础设施 应用 研究原型 Stars 4617 周增 +7

Transformer 模型的高速推理引擎。Fast inference engine for Transformer models

llm-infra
MiniMax-AI/MiniMax-H3
Python · 2026-08-10 LLM 基础设施 模型 研究原型 Stars 4611 周增 +4445
luban-agi/Awesome-AIGC-Tutorials
未知语言 · 2024-03-31 LLM 基础设施 收藏榜 生产可用 Stars 4531 周增 -7

精选的 LLM、AI 绘画等领域的教程和资源。Curated tutorials and resources for Large Language Models, AI Painting, and more.

multimodalllm-infra
openvinotoolkit/open_model_zoo
Python · 2026-07-09 LLM 基础设施 模型 研究原型 Stars 4415 周增 +7

高质量且极速的预训练深度学习模型与 demo。Pre-trained Deep Learning models and demos (high quality and extremely fast)

llm-infra
algorithmicsuperintelligence/optillm
Python · 2026-07-18 LLM 基础设施 应用 研究原型 Stars 4236 周增 +7

为 LLM 优化推理代理Optimizing inference proxy for LLMs

agentllm-infra
NVIDIA/GenerativeAIExamples
Jupyter Notebook · 2026-08-05 LLM 基础设施 框架 生产可用 Stars 4144 周增 +14

针对加速基础设施和微服务架构优化的生成式 AI 参考工作流。Generative AI reference workflows optimized for accelerated infrastructure and microservice architecture.

ragllm-infra
IDEA-CCNL/Fengshenbang-LM
Python · 2026-06-08 LLM 基础设施 模型 生产可用 Stars 4126 周增 +7

Fengshenbang-LM(封神榜大模型)是IDEA研究院认知计算与自然语言研究中心主导的大模型开源体系,成为中文AIGC和认知智能的基础设施。

multimodal
billmei/every-chatgpt-gui
未知语言 · 2026-07-23 LLM 基础设施 收藏榜 生产可用 Stars 3997 周增 +0

ChatGPT、Claude 等 LLM 的所有前端 GUI 客户端汇总。Every front-end GUI client for ChatGPT, Claude, and other LLMs

llm-infra
zml/zml
Zig · 2026-08-03 LLM 基础设施 框架 研究原型 Stars 3972 周增 +7

任意模型,任意硬件,零妥协。基于 @ziglang / @openxla / MLIR / @bazelbuild 构建。Any model. Any hardware. Zero compromise. Built with @ziglang / @openxla / MLIR / @bazelbuild

llm-infra
huggingface/optimum
Python · 2026-08-10 LLM 基础设施 工具 研究原型 Stars 3457 周增 +7

🚀 通过易用的硬件优化工具,加速 🤗 Transformers、Diffusers、TIMM 和 Sentence Transformers 的推理与训练。🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools

llm-infra
raullenchai/Rapid-MLX
Python · 2026-08-11 LLM 基础设施 框架 研究原型 Stars 3433 周增 +21

Apple Silicon 上最快的本地 AI 引擎。比 Ollama 快 4.2 倍,缓存 TTFT 仅 0.08s,工具调用支持率 100%。内置 17 种工具解析器、prompt cache、推理分离、云端路由。可作为 OpenAI 的即插即用替代,兼容 Claude Code、Cursor、Aider。The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.

llm-infra
DSXiangLi/DecryptPrompt
未知语言 · 2026-05-06 LLM 基础设施 收藏榜 生产可用 Stars 3432 周增 +7

总结Prompt&LLM论文,开源数据&模型,AIGC应用

agentllm-infra
Niko1221/Strata
C++ · 2026-10-01 LLM 基础设施 应用 研究原型 Stars 3310 周增 +4375

在任意消费级硬件上运行 Qwen3.8-Flash-Next:Windows / Linux 一键安装。Strata 推理引擎,本地 localhost 提供 OpenAI / Anthropic API,支持可选图像输入。Qwen3.8-Flash-Next on any consumer hardware: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input.

multimodalllm-infra
openvinotoolkit/openvino_notebooks
Jupyter Notebook · 2026-08-10 LLM 基础设施 教程 研究原型 Stars 3195 周增 +0

📚 OpenVINO™ 的 Jupyter notebook 教程。📚 Jupyter notebook tutorials for OpenVINO™

multimodalllm-infra
kvcache-ai/AgentENV
Rust · 2026-08-11 LLM 基础设施 框架 研究原型 Stars 3140 周增 +126

AgentENV (AENV) 是一个用于大规模运行 Agent 环境的分布式平台。AgentENV (AENV) is a distributed platform for running agent environments at scale.

agent
pytorch/ao
Python · 2026-08-10 LLM 基础设施 应用 研究原型 Stars 2936 周增 +14

PyTorch 原生的量化和稀疏化方案,支持训练与推理PyTorch native quantization and sparsity for training and inference

llm-infra
intel/neural-compressor
Python · 2026-08-11 LLM 基础设施 模型 研究原型 Stars 2697 周增 +0

SOTA 低比特 LLM 量化(INT8/FP8/MXFP8/INT4/MXFP4/NVFP4)与稀疏化方案;面向 PyTorch、TensorFlow 与 ONNX Runtime 的领先模型压缩技术SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity; leading model compression techniques on PyTorch, TensorFlow, and ONNX Runtime

llm-infra
vllm-project/vllm-ascend
C++ · 2026-08-11 LLM 基础设施 应用 研究原型 Stars 2600 周增 +49

为 vLLM 在 Ascend 上的社区维护硬件插件Community maintained hardware plugin for vLLM on Ascend

llm-infraengineering
huggingface/huggingface.js
TypeScript · 2026-08-10 LLM 基础设施 库 研究原型 Stars 2487 周增 -7

在 JavaScript 中使用 Hugging FaceUse Hugging Face with JavaScript

llm-infra
pykeio/ort
Rust · 2026-08-08 LLM 基础设施 应用 研究原型 Stars 2443 周增 +14

基于 Rust 的 ONNX 模型快速 ML 推理与训练Fast ML inference & training for ONNX models in Rust

llm-infra
AI-Hypercomputer/maxtext
Python · 2026-10-07 LLM 基础设施 框架 研究原型 Stars 2439 周增 +7

一个简洁、高性能且可扩展的 Jax LLM!A simple, performant, and scalable Jax LLM!

llm-infra
google/XNNPACK
C · 2026-08-10 LLM 基础设施 应用 研究原型 Stars 2421 周增 -7

高效的浮点神经网络推理算子,覆盖移动端、服务端和 WebHigh-efficiency floating-point neural network inference operators for mobile, server, and Web

llm-infra
roboflow/inference
Python · 2026-08-10 LLM 基础设施 应用 研究原型 Stars 2411 周增 +7

将任意电脑或边缘设备打造为计算机视觉项目的指挥中心Turn any computer or edge device into a command center for your computer vision projects.

agentmultimodalllm-infraengineering
bionic-gpt/bionic-gpt
Rust · 2026-08-03 LLM 基础设施 应用 研究原型 Stars 2349 周增 +7

Bionic 是 ChatGPT 的本地化部署替代方案,在保持严格数据机密性的同时提供生成式 AI 的能力Bionic is an on-premise replacement for ChatGPT, offering the advantages of Generative AI while maintaining strict data confidentiality

llm-infra
NVIDIA-NeMo/DataDesigner
Python · 2026-09-21 LLM 基础设施 工具 生产可用 Stars 2264 周增 +0

🎨 NeMo Data Designer:从零或从种子数据生成高质量合成数据。🎨 NeMo Data Designer: Generate high-quality synthetic data from scratch or from seed data.

agentmultimodalllm-infra
mozilla-ai/any-llm
Python · 2026-08-10 LLM 基础设施 库 研究原型 Stars 2150 周增 +0

通过统一接口对接 LLM 服务商Communicate with an LLM provider using a single interface

llm-infra
envoyproxy/ai-gateway
Go · 2026-08-10 LLM 基础设施 框架 研究原型 Stars 1919 周增 +14

基于 Envoy Gateway 构建,提供生成式 AI 服务的统一接入管理Manages Unified Access to Generative AI Services built on Envoy Gateway

llm-infra
mlc-ai/xgrammar
C++ · 2026-09-20 LLM 基础设施 库 生产可用 Stars 1898 周增 +0

快速、灵活、可移植的结构化生成。Fast, Flexible and Portable Structured Generation

llm-infra
beam-cloud/beta9
Go · 2026-09-21 LLM 基础设施 应用 研究原型 Stars 1785 周增 +8

超快速 serverless GPU 推理、沙箱和后台任务Ultrafast serverless GPU inference, sandboxes, and background jobs

llm-infra
NVIDIA-NeMo/Curator
Python · 2026-09-28 LLM 基础设施 工具 研究原型 Stars 1784 周增 +12

可扩展的 LLM 数据预处理与清洗工具集Scalable data pre processing and curation toolkit for LLMs

engineeringllm-infra
trymirai/uzu
Rust · 2026-08-10 LLM 基础设施 应用 研究原型 Stars 1670 周增 +0

高性能 AI 模型推理引擎A high-performance inference engine for AI models

llm-infra
TheoLeeCJ/SemIf
Python · 2026-09-19 LLM 基础设施 工具 实验 Stars 1648 周增 +0

在家用 3090 上由开源模型生成语义化 if。独立项目,与 Jev 或 TypeSafe 无关。Semantic ifs from open models, on a 3090 at home. Independent; not affiliated with Jev or TypeSafe.

Edge0-AI/Edge0
Python · 2026-09-14 LLM 基础设施 库 实验 Stars 1627 周增 +1930
lynote-ai/humanize-text
Python · 2026-08-05 LLM 基础设施 工具 研究原型 Stars 1558 周增 +21

面向 AI 辅助草稿的可读性与自然节奏改进的开源 pipeline 与参考实现。Open-source pipeline and reference implementations for improving the readability and natural cadence of AI-assisted drafts.

engineering
utkuozdemir/nvidia_gpu_exporter
Go · 2026-08-07 LLM 基础设施 工具 生产可用 Stars 1527 周增 +14

基于 nvidia-smi 二进制工具的 Nvidia GPU Prometheus exporterNvidia GPU exporter for prometheus using nvidia-smi binary

llm-infra