研究库 开源仓库
Repositories · organized/repo_cards

仓库/Skill 库

263 个 · LLM 基础设施

排序 Stars 周增
theopenco/llmgateway
TypeScript · 2026-08-11 LLM 基础设施 框架 研究原型 Stars 1520 周增 +7

通过统一 API 接口路由、管理和分析跨多家服务商的 LLM 请求Route, manage, and analyze your LLM requests across multiple providers with a unified API interface.

llm-infra
wangxuqi/Prompt-Engineering-Guide-Chinese
MDX · 2024-09-14 LLM 基础设施 教程 生产可用 Stars 1393 周增 +7

Prompt工程师指南,源自英文版,但增加了AIGC的prompt部分,为了降低同学们的学习门槛,翻译更新

alibaba/rtp-llm
Cuda · 2026-08-05 LLM 基础设施 应用 研究原型 Stars 1296 周增 +0

RTP-LLM:阿里巴巴面向多样化应用的高性能 LLM 推理引擎。RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.

llm-infra
StarlightSearch/EmbedAnything
Rust · 2026-08-11 LLM 基础设施 应用 研究原型 Stars 1294 周增 +5

基于 Rust 🦀 构建的高性能、模块化、内存安全、生产可用的推理、数据接入与索引系统Highly Performant, Modular, Memory Safe and Production-ready Inference, Ingestion and Indexing built in Rust 🦀

ragllm-infraengineeringdatabase
yibie/awesome-jev
Python · 2026-09-22 LLM 基础设施 收藏榜 实验 Stars 1257 周增 +1848

一个精选列表,收录基于 Jev(TypeSafe AI 的 System One 模型,用于类型化决策)构建的公开项目、集成与讨论。A curated list of public projects, integrations, and discussions built on Jev — TypeSafe AI's System One model for typed decisions.

llm-infra
NVIDIA/kvpress
Python · 2026-10-01 LLM 基础设施 应用 研究原型 Stars 1222 周增 +0

让 LLM KV cache 压缩变得更简单LLM KV cache compression made easy

llm-infra
Arize-ai/openinference
Python · 2026-09-04 LLM 基础设施 库 生产可用 Stars 1195 周增 +0

AI 可观测性的 OpenTelemetry 插桩OpenTelemetry Instrumentation for AI Observability

agentllm-infra
vndee/llm-sandbox
Python · 2026-08-23 LLM 基础设施 库 生产可用 Stars 1110 周增 +0

轻量、可移植的 LLM 沙箱运行时(代码解释器)Python 库。Lightweight and portable LLM sandbox runtime (code interpreter) Python library.

llm-infra
MAC-AutoML/MindPipe
Python · 2026-08-11 LLM 基础设施 框架 研究原型 Stars 1013 周增 +0

面向 LLM 和 LVLM 的强大模型压缩框架,适配 NVIDIA GPU 和华为昇腾 NPUA powerful model compression framework for LLMs and LVLMs, adapted for NVIDIA GPUs and Huawei Ascend NPUs.

multimodalllm-infraevaluationengineering
antirez/h3.c
C · 2026-08-11 LLM 基础设施 应用 研究原型 Stars 980 周增 +0

MiniMax H3 Mac 推理引擎。MiniMax H3 inference engine for Mac computers

llm-infra
TheoLeeCJ/openjev
Python · 2026-09-18 LLM 基础设施 工具 实验 Stars 964 周增 +0

我们能在家里用 3090 跑类似 Jev 的东西吗?Can we run something like Jev on a 3090 at home?

nick7nlp/Awesome-LLM-On-Policy-Distillation
Python · 2026-10-05 LLM 基础设施 收藏榜 研究原型 Stars 568 周增 +7

关于大语言模型 On-Policy Distillation 的精选论文与资源合集A curated collection of papers and resources on On-Policy Distillation for Large Language Models.

llm-infra
Mirrowel/LLM-API-Key-Proxy
Python · 2026-09-11 LLM 基础设施 工具 生产可用 Stars 550 周增 +4

通用 LLM 网关:一个 API 对接所有 LLM。提供兼容 OpenAI/Anthropic 的端点,支持多 provider 转换与智能负载均衡。Universal LLM Gateway: One API, every LLM. OpenAI/Anthropic-compatible endpoints with multi-provider translation and intelligent load-balancing.

llm-infra
zouyuxuan122/dsh-our-free-model
JavaScript · 2026-10-01 LLM 基础设施 工具 实验 Stars 532 周增 +508

在 dsh 里装上这个插件即可,无需登录、注册或填 API Key,就能使用包括 Muse Spark 1.3、MiMo V2.6 在内的前沿模型——完全免费,不限量。 All you do is install this plugin in dsh: no login, no sign-up, no API key — the frontier models are just there, Muse Spark 1.3 and MiMo V2.6 among them. Completely free, with no usage cap.

agentllm-infra
ArasTey/lunel
Python · 2026-09-13 LLM 基础设施 工具 研究原型 Stars 532 周增 +0
ArronAI007/Awesome-AGI
Jupyter Notebook · 2026-05-15 LLM 基础设施 收藏榜 研究原型 Stars 513 周增 +0

AGI资料汇总学习(主要包括LLM和AIGC),持续更新......

llm-infra
brontoguana/krasis
C++ · 2026-08-20 LLM 基础设施 工具 研究原型 Stars 512 周增 +3

Krasis 是一个混合 LLM 运行时,专注于在消费级 VRAM 受限硬件上高效运行大模型Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware

llm-infra
ollaya-dev/ollaya
Rust · 2026-09-27 LLM 基础设施 工具 实验 Stars 475 周增 +0

本地运行开源决策模型:在兼容 TypeSafe 的 API 后面拉取并服务 Laya、decider、NLI 和 GLiClass。基于 Ollama 的决策模型。Run open decision models locally: pull and serve Laya, decider, NLI and GLiClass behind a TypeSafe-compatible API. Ollama for decision models.

llm-infra
atria-asi/Atria-Dawn-Preview
未知语言 · 2026-09-17 LLM 基础设施 模型 实验 Stars 461 周增 +924
agentllm-infra
instavm/open-skills
Python · 2026-01-23 LLM 基础设施 框架 实验 Stars 445 周增 +0

OpenSkills:使用任意 LLM 在本地运行 Claude Skills。OpenSkills: Run Claude Skills Locally using any LLM

llm-infra
AbdelStark/awesome-typesafe
CSS · 2026-09-20 LLM 基础设施 收藏榜 实验 Stars 384 周增 +462

一个精选列表,收录 TypeSafe、System One 模型及 Jev 的官方资源与社区项目。A curated list of official resources and community projects for TypeSafe, System One models, and Jev.

agentllm-infra
jev-chat/jev-chat-jarvis-mac
Python · 2026-09-24 LLM 基础设施 应用 研究原型 Stars 356 周增 +364

聊天悬浮窗助手(macOS):屏幕感知 + 本地小模型判断意图与风险,按话术生成回复候选。纯只读。

llm-infra
syv-ai/qwen38-27b-rtx3090
Python · 2026-08-21 LLM 基础设施 评测集 研究原型 Stars 345 周增 +0

Qwen3.8-27B 在单卡 RTX 3090 上使用 vLLM 部署:64 并发下约 1,000 tok/s(int8 张量核心 GEMM、fp16 DeltaNet 状态),默认采样下单用户约 114 tok/s/贪心约 124 tok/s(MTP 草稿、自输出草稿词表、校准 int4 lm_head、split-KV 校验注意力),150k–262k 上下文;附带补丁、重新量化脚本与基准测试Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks

llm-infraevaluation
gvzdv/claudish-to-english
Shell · 2026-08-11 LLM 基础设施 工具 实验 Stars 341 周增 +0
incoai/splash
Python · 2026-09-20 LLM 基础设施 模型 研究原型 Stars 339 周增 +0

一款围绕该模型构建的、面向 Apple silicon 的本地推理引擎。A local inference engine for Apple silicon, built around the model.

agentllm-infra
Yu-Yang-Li/StarWhisper
Python · 2026-08-17 LLM 基础设施 应用 研究原型 Stars 325 周增 +0

StarWhisper 天文 LLMs、StarWhisper Telescope、Virtual-GOTTA,以及面向 embodied observing workflow 的天文定制研究 skills。StarWhisper astronomy LLMs, StarWhisper Telescope, Virtual-GOTTA, and astronomy-adapted research skills for embodied observing workflows

agentllm-infra
linguo2625469/workbuddy2api-panel
Go · 2026-09-16 LLM 基础设施 工具 实验 Stars 317 周增 +637

把腾讯WorkBuddy账号变成 OpenAI 兼容 API 的多账号网关,同时自动完成任务中心全部任务,附 Web 管理面板(账号池可视化 / 积分任务 / 配置热更新)。基于 Sliverkiss/workbuddy2api 的增强分支

kvmem/kvmem-llama.cpp
C++ · 2026-09-20 LLM 基础设施 库 实验 Stars 314 周增 +0
amitshekhariitbhu/llm-inference-engineering
Markdown · 2026-10-05 LLM 基础设施 应用 实验 Stars 294 周增 +12

逐步学习 LLM 推理工程——从 KV cache、PagedAttention 和连续批处理,到 vLLM、SGLang 和 GPU。Learn LLM Inference Engineering step by step - from KV cache, PagedAttention, and continuous batching to vLLM, SGLang, and GPUs.

llm-infra
MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-Spark
Python · 2026-09-09 LLM 基础设施 工具 实验 Stars 291 周增 +274

Qwen3.8-Flash-Next 部署于单台 DGX Spark(TP=1)Qwen3.8-Flash-Next on ONE DGX Spark (TP=1)

MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks-TensorFold
Shell · 2026-10-05 LLM 基础设施 模型 实验 Stars 251 周增 +0

基于 2x DGX Spark 与 TensorFold 的 GLM-5.3-Flash EXL3。GLM-5.3-Flash EXL3 on 2x DGX Spark with TensorFold

langfuse/langfuse-docs
MDX · 2026-09-11 LLM 基础设施 工具 实验 Stars 240 周增 +0

开源 agent 评估与可观测性:在一个开放平台上追踪、评估并改进 LLM 应用。🪢 Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.

agentllm-infra
pingmike2/freebuff2api-wokers
JavaScript · 2026-08-11 LLM 基础设施 工具 实验 Stars 236 周增 +0
carloslfu/slotstream
Swift · 2026-09-02 LLM 基础设施 工具 实验 Stars 232 周增 +560

在 Mac 上以远低于 104 GB 的内存运行 Qwen3.8-Flash-Next(125B MoE,4-bit 下约 104 GB),通过 SSD 流式加载专家。MLX + Swift,兼容 Ollama 的 API。Run Qwen3.8-Flash-Next (125B MoE, 104 GB at 4-bit) on Macs with a fraction of that RAM by streaming experts from SSD. MLX + Swift, Ollama-compatible API.

llm-infra
walkinglabs/modern-llm-notebook
Jupyter Notebook · 2026-09-30 LLM 基础设施 应用 实验 Stars 222 周增 +9

一门关于现代 LLM 架构、训练与推理的实战课程,配有逐步 PyTorch 实现和可运行的 notebook。A hands-on course on modern LLM architectures, training, and inference, with step-by-step PyTorch implementations and runnable notebooks.

llm-infra
MiaAI-Lab/DeepSeek-v4-Flash-One-DGX-Spark
Python · 2026-08-24 LLM 基础设施 模型 实验 Stars 215 周增 +0

DeepSeek v4 Flash EXL3 运行于单台 DGX SparkDeepSeek v4 Flash EXL3 on one DGX Spark