研究库 开源仓库
Repositories · organized/repo_cards

仓库/Skill 库

14 个 · LLM 基础设施 · 应用 · 快速增长

排序 Stars 周增
open-webui/open-webui
Python · 2026-08-11 LLM 基础设施 应用 生产可用 Stars 148436 周增 +294

友好的 AI 交互界面(支持 Ollama、OpenAI API 等)。User-friendly AI Interface (Supports Ollama, OpenAI API, ...)

ragllm-infra
vllm-project/vllm
Python · 2026-08-11 LLM 基础设施 应用 生产可用 Stars 88728 周增 +259

面向 LLM 的高吞吐、内存高效的推理与 serving 引擎。A high-throughput and memory-efficient inference and serving engine for LLMs

llm-infra
lyogavin/airllm
Jupyter Notebook · 2026-08-10 LLM 基础设施 应用 生产可用 Stars 30619 周增 +574

使用单卡 4GB GPU 推理 AirLLM 70BAirLLM 70B inference with single 4GB GPU

llm-infra
cheahjs/free-llm-api-resources
Python · 2026-08-04 LLM 基础设施 应用 生产可用 Stars 29364 周增 +259

通过 API 访问的免费 LLM 推理资源列表。A list of free LLM inference resources accessible via API.

llm-infra
antirez/ds4
C · 2026-07-03 LLM 基础设施 应用 生产可用 Stars 17465 周增 +315

面向 Metal、CUDA 和 ROCm 的 DeepSeek 4 Flash 与 PRO 本地推理引擎。DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm

llm-infra
drumih/turbo-fieldfare
Swift · 2026-08-10 LLM 基础设施 应用 研究原型 Stars 5668 周增 +448

Gemma 4 26B-A4B 在任意 M 系列 MacBook 上以约 2 GB 内存进行推理。Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook

llm-infra
Niko1221/Strata
C++ · 2026-10-01 LLM 基础设施 应用 研究原型 Stars 3310 周增 +4375

在任意消费级硬件上运行 Qwen3.8-Flash-Next:Windows / Linux 一键安装。Strata 推理引擎,本地 localhost 提供 OpenAI / Anthropic API,支持可选图像输入。Qwen3.8-Flash-Next on any consumer hardware: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input.

multimodalllm-infra
antirez/h3.c
C · 2026-08-11 LLM 基础设施 应用 研究原型 Stars 980 周增 +0

MiniMax H3 Mac 推理引擎。MiniMax H3 inference engine for Mac computers

llm-infra
jev-chat/jev-chat-jarvis-mac
Python · 2026-09-24 LLM 基础设施 应用 研究原型 Stars 356 周增 +364

聊天悬浮窗助手(macOS):屏幕感知 + 本地小模型判断意图与风险,按话术生成回复候选。纯只读。

llm-infra
MiaAI-Lab/Qwen3.8-27B-SGLang-DGX-Spark
Python · 2026-08-19 LLM 基础设施 应用 实验 Stars 168 周增 +0

Qwen3.8 27B 在 SGLang 上运行于 DGX SparkQwen3.8 27B on SGLang for DGX Spark

llm-infra
kunpengtalk/OmniStudio
TypeScript · 2026-09-16 LLM 基础设施 应用 实验 Stars 108 周增 +114

OmniStudio 是一个本地大模型一体化桌面工作台,集模型市集下载、llama.cpp/vLLM/SGLang 三引擎推理管理,以及对话、语音合成、ASR语音识别、图片生成、视频生成、OCR 等多种大模型应用于一体,全程本地优先。

multimodalllm-infra
tonyd2wild/GLM-5.3-Flash-NVFP4-DFlash2-2x-DGX-Spark
Python · 2026-08-28 LLM 基础设施 应用 实验 Stars 69 周增 +0

2× NVIDIA DGX Spark 上的 GLM-5.3-Flash (NVFP4)——vLLM TP2,262K 上下文,MTP。全球首发部署方案:发现并修复 7 个 day-0 bug,附 sm121 镜像补丁、探针与完整报告。GLM-5.3-Flash (NVFP4) on 2x NVIDIA DGX Spark - vLLM TP2, 262K context, MTP. World-first deploy recipe: 7 day-0 bugs found and fixed, patched sm121 image, probes and full report.

multimodalllm-infra
ekzhang/openjev-sglang
Python · 2026-09-18 LLM 基础设施 应用 实验 Stars 69 周增 +0

基于开源模型的 Jev 兼容 API 端点(仅 prefill)Jev-compatible API endpoint based on open models (prefill-only)

llm-infra
Wallawalla47/Infernix
C++ · 2026-10-09 LLM 基础设施 应用 实验 Stars 54 周增 +0

面向 Qwen3.8(含支持专家卸载的 Qwen3.8-Flash-Next)的 C++/CUDA 单 GPU 推理引擎,运行于 RTX 5090;源自 NInfer。C++/CUDA single-GPU inference engine for Qwen3.8 (incl. Qwen3.8-Flash-Next with offloaded experts) on the RTX 5090; grew from NInfer

llm-infra