Repositories · organized/repo_cards

仓库/Skill 库

27 个 · LLM 基础设施 · 快速增长

排序 Stars 周增
open-webui/open-webui
Python · 2026-08-11 LLM 基础设施 应用 生产可用 Stars 148436 周增 +294

友好的 AI 交互界面(支持 Ollama、OpenAI API 等)。User-friendly AI Interface (Supports Ollama, OpenAI API, ...)

ragllm-infra
microsoft/generative-ai-for-beginners
Jupyter Notebook · 2026-08-06 LLM 基础设施 教程 生产可用 Stars 117401 周增 +651

21 节课程,开启生成式 AI 应用开发之旅21 Lessons, Get Started Building with Generative AI

llm-infra
vllm-project/vllm
Python · 2026-08-11 LLM 基础设施 应用 生产可用 Stars 88728 周增 +259

面向 LLM 的高吞吐、内存高效的推理与 serving 引擎。A high-throughput and memory-efficient inference and serving engine for LLMs

llm-infra
BerriAI/litellm
Python · 2026-08-11 LLM 基础设施 生产可用 Stars 56075 周增 +259

Python SDK、代理服务器(AI Gateway),以 OpenAI(或原生)格式调用 100+ LLM API,支持成本追踪、guardrails、负载均衡和日志记录。[Bedrock, Azure, OpenAI, VertexAI, Cohere, Anthropic, Sagemaker, HuggingFace, VLLM, NVIDIA NIM]The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthropic, OpenAI, VertexAI, vLLM, Nvidia NIM]

llm-infra
Wei-Shaw/sub2api
Go · 2026-08-11 LLM 基础设施 工具 生产可用 Stars 36568 周增 +287

Sub2API 是一款开源中继平台,将 Claude、OpenAI、Gemini 与 Antigravity 订阅统一为单一端点,支持账号共享与费用分摊,兼容原生工具调用。Sub2API 一站式开源中转服务,让 Claude、Openai 、Gemini、Grok订阅统一接入,支持拼车共享,更高效分摊成本,原生工具无缝使用。

sgl-project/sglang
Python · 2026-08-11 LLM 基础设施 框架 生产可用 Stars 31657 周增 +203

SGLang 是一个面向大语言模型和多模态模型的高性能 serving 框架。SGLang is a high-performance serving framework for large language models and multimodal models.

multimodalllm-infra
lyogavin/airllm
Jupyter Notebook · 2026-08-10 LLM 基础设施 应用 生产可用 Stars 30619 周增 +574

使用单卡 4GB GPU 推理 AirLLM 70BAirLLM 70B inference with single 4GB GPU

llm-infra
cheahjs/free-llm-api-resources
Python · 2026-08-04 LLM 基础设施 应用 生产可用 Stars 29364 周增 +259

通过 API 访问的免费 LLM 推理资源列表。A list of free LLM inference resources accessible via API.

llm-infra
JustVugg/colibri
C · 2026-08-10 LLM 基础设施 工具 生产可用 Stars 23805 周增 +1169

在现有硬件上运行前沿 MoE 模型——纯 C 实现、零依赖、专家权重从磁盘流式加载。轻量引擎,海量模型。🐦Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦

antirez/ds4
C · 2026-07-03 LLM 基础设施 应用 生产可用 Stars 17465 周增 +315

面向 Metal、CUDA 和 ROCm 的 DeepSeek 4 Flash 与 PRO 本地推理引擎。DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm

llm-infra
lidge-jun/opencodex
TypeScript · 2026-08-11 LLM 基础设施 工具 研究原型 Stars 9000 周增 +658

OpenAI Codex 与 Claude Code 的通用 provider 代理 —— 可在 Codex CLI、App、SDK 及 Claude Code 中使用任意 LLM(Claude、Gemini、Grok、DeepSeek、Ollama…)。Universal provider proxy for OpenAI Codex & Claude Code — use any LLM (Claude, Gemini, Grok, DeepSeek, Ollama…) with Codex CLI, App, SDK, and Claude Code

llm-infra
MoonshotAI/Kimi-K3
未知语言 · 2026-08-06 LLM 基础设施 模型 生产可用 Stars 8340 周增 +161

开放式前沿智能Open Frontier Intelligence

drumih/turbo-fieldfare
Swift · 2026-08-10 LLM 基础设施 应用 研究原型 Stars 5668 周增 +448

Gemma 4 26B-A4B 在任意 M 系列 MacBook 上以约 2 GB 内存进行推理。Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook

llm-infra
FareedKhan-dev/kimi-k3-in-c
C · 2026-08-07 LLM 基础设施 框架 研究原型 Stars 4709 周增 +889

一个 2.78 万亿参数的 Kimi K3,在仅 8.24 GB 内存的单颗 CPU 上运行推理。可移植的 C99:无需 BLAS,无需框架,无需 GPU。A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.

llm-infra
MiniMax-AI/MiniMax-H3
Python · 2026-08-10 LLM 基础设施 模型 研究原型 Stars 4611 周增 +4445
kvcache-ai/AgentENV
Rust · 2026-08-11 LLM 基础设施 框架 研究原型 Stars 3140 周增 +126

AgentENV (AENV) 是一个用于大规模运行 Agent 环境的分布式平台。AgentENV (AENV) is a distributed platform for running agent environments at scale.

agent
yjh051108/dsh-routing-suite
PowerShell · 2026-08-15 LLM 基础设施 工具 实验 Stars 1504 周增 +0

dsh-routing-suite——注入器 + 路由标准套件:先安装运行时注入器,再安装任务感知的推理模式路由预设(已实测 P1-P23)。dsh-routing-suite — injector + router-standard kit: install the runtime injector first, then the task-aware reasoning-mode router preset (measured P1-P23).

antirez/h3.c
C · 2026-08-11 LLM 基础设施 应用 研究原型 Stars 980 周增 +0

MiniMax H3 Mac 推理引擎。MiniMax H3 inference engine for Mac computers

llm-infra
syv-ai/qwen38-27b-rtx3090
Python · 2026-08-21 LLM 基础设施 评测集 研究原型 Stars 345 周增 +0

Qwen3.8-27B 在单卡 RTX 3090 上使用 vLLM 部署:64 并发下约 1,000 tok/s(int8 张量核心 GEMM、fp16 DeltaNet 状态),默认采样下单用户约 114 tok/s/贪心约 124 tok/s(MTP 草稿、自输出草稿词表、校准 int4 lm_head、split-KV 校验注意力),150k–262k 上下文;附带补丁、重新量化脚本与基准测试Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks

llm-infraevaluation
gvzdv/claudish-to-english
Shell · 2026-08-11 LLM 基础设施 工具 实验 Stars 341 周增 +0
pingmike2/freebuff2api-wokers
JavaScript · 2026-08-11 LLM 基础设施 工具 实验 Stars 236 周增 +0
MiaAI-Lab/DeepSeek-v4-Flash-One-DGX-Spark
Python · 2026-08-24 LLM 基础设施 模型 实验 Stars 215 周增 +0

DeepSeek v4 Flash EXL3 运行于单台 DGX SparkDeepSeek v4 Flash EXL3 on one DGX Spark

maliubiao/dgx-spark-2-deepseek-flash-0731
Shell · 2026-08-11 LLM 基础设施 教程 实验 Stars 169 周增 +0

在两台 dgx-spark 上从零部署 deepseek-flash-0731 的设置指南。setup guide for deepseek-flash-0731 on two dgx-spark from scratch

MiaAI-Lab/Qwen3.8-27B-SGLang-DGX-Spark
Python · 2026-08-19 LLM 基础设施 应用 实验 Stars 168 周增 +0

Qwen3.8 27B 在 SGLang 上运行于 DGX SparkQwen3.8 27B on SGLang for DGX Spark

llm-infra
trefeon/freebuff-proxy
Go · 2026-08-16 LLM 基础设施 工具 实验 Stars 85 周增 +112

FreeBuff 编码模型的 OpenAI 兼容网关。Token 池、会话生命周期、TLS stealth、嵌入式管理后台。无广告、无 CLI,只有 /v1/chat/completions。OpenAI-compatible gateway for FreeBuff coding models. Token pool, session lifecycle, TLS stealth, embedded admin dashboard. No ads, no CLI, just /v1/chat/completions.

agentllm-infra
meta-models/meta-oss-cookbook
Python · 2026-08-12 LLM 基础设施 教程 生产可用 Stars 63 周增 +0

Meta Inc. 所有 oss 模型的相关 recipes。All recipes for oss models from Meta Inc.

mattn/tensai
Go · 2026-08-25 LLM 基础设施 框架 实验 Stars 57 周增 +0

纯 Go 编写的轻量神经网络框架,使用 AVX2 SIMD 内核(GOEXPERIMENT=simd)。A tiny neural-network framework in pure Go with AVX2 SIMD kernels (GOEXPERIMENT=simd)