研究库 开源仓库
Repositories · organized/repo_cards

仓库/Skill 库

38 个 · LLM 基础设施 · 应用 · AI 核心

排序 Stars 周增
ggml-org/llama.cpp
C++ · 2026-08-03 LLM 基础设施 应用 生产可用 Stars 122572 周增 +0

基于 C/C++ 的 LLM 推理实现LLM inference in C/C++

llm-infra
onyx-dot-app/onyx
Python · 2026-09-24 LLM 基础设施 应用 生产可用 Stars 32229 周增 +116

开源 AI 平台——支持所有 LLM 的 AI Chat,并提供丰富的高级功能。Open Source AI Platform - AI Chat with advanced features that works with every LLM

ragdatabasellm-infra
karpathy/llm.c
Cuda · 2025-06-26 LLM 基础设施 应用 实验 Stars 31122 周增 +0

用简单、原生的 C/CUDA 进行 LLM 训练LLM training in simple, raw C/CUDA

llm-infra
dyad-sh/dyad
TypeScript · 2026-08-11 LLM 基础设施 应用 生产可用 Stars 21197 周增 +21

本地化、开源的 AI 应用构建器,面向高级用户 ✨ v0 / Lovable / Replit / Bolt 的替代方案 🌟 喜欢就点 Star!Local, open-source AI app builder for power users ✨ v0 / Lovable / Replit / Bolt alternative 🌟 Star if you like it!

llm-infra
GeeeekExplorer/nano-vllm
Python · 2026-04-26 LLM 基础设施 应用 生产可用 Stars 14623 周增 +28

Nano vLLM。Nano vLLM

llm-infra
codexu/note-gen
TypeScript · 2026-08-26 LLM 基础设施 应用 生产可用 Stars 12685 周增 +0

先记录,后整理。一款 local-first 的 Markdown 应用,借助 AI 将零散记录转化为清晰的笔记。Capture first. Organize later. A local-first Markdown app that turns scattered records into clear notes with AI.

agentragllm-infra
LMCache/LMCache
Python · 2026-08-11 LLM 基础设施 应用 生产可用 Stars 11101 周增 +49

LMCache:借助最快的 KV Cache 层为你的 LLM 加速。LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

llm-infra
triton-inference-server/server
Python · 2026-08-10 LLM 基础设施 应用 生产可用 Stars 10913 周增 +14

Triton Inference Server 提供针对云端和边缘推理优化的解决方案。The Triton Inference Server provides an optimized cloud and edge inferencing solution.

llm-infra
ultralytics/yolov3
Python · 2026-08-02 LLM 基础设施 应用 生产可用 Stars 10591 周增 +14

基于 PyTorch 的 YOLOv3、YOLOv3-SPP 和 YOLOv3-tiny 实时目标检测实现,支持训练、验证、推理与多格式导出。PyTorch implementation of YOLOv3, YOLOv3-SPP, and YOLOv3-tiny for real-time object detection with training, validation, inference, and multi-format export.

multimodalllm-infra
rockbenben/ChatGPT-Shortcut
TypeScript · 2026-08-03 LLM 基础设施 应用 生产可用 Stars 8679 周增 +7

🚀💪Maximize your efficiency and productivity. The ultimate hub to manage, customize, and share prompts. (English/中文/Español/العربية). 让生产力加倍的 AI 快捷指令。更高效地管理提示词,在分享社区中发现适用于不同场景的灵感。

llm-infra
OpenNMT/CTranslate2
C++ · 2026-08-05 LLM 基础设施 应用 研究原型 Stars 4617 周增 +7

Transformer 模型的高速推理引擎。Fast inference engine for Transformer models

llm-infra
algorithmicsuperintelligence/optillm
Python · 2026-07-18 LLM 基础设施 应用 研究原型 Stars 4236 周增 +7

为 LLM 优化推理代理Optimizing inference proxy for LLMs

agentllm-infra
pytorch/ao
Python · 2026-08-10 LLM 基础设施 应用 研究原型 Stars 2936 周增 +14

PyTorch 原生的量化和稀疏化方案,支持训练与推理PyTorch native quantization and sparsity for training and inference

llm-infra
vllm-project/vllm-ascend
C++ · 2026-08-11 LLM 基础设施 应用 研究原型 Stars 2600 周增 +49

为 vLLM 在 Ascend 上的社区维护硬件插件Community maintained hardware plugin for vLLM on Ascend

llm-infraengineering
pykeio/ort
Rust · 2026-08-08 LLM 基础设施 应用 研究原型 Stars 2443 周增 +14

基于 Rust 的 ONNX 模型快速 ML 推理与训练Fast ML inference & training for ONNX models in Rust

llm-infra
google/XNNPACK
C · 2026-08-10 LLM 基础设施 应用 研究原型 Stars 2421 周增 -7

高效的浮点神经网络推理算子,覆盖移动端、服务端和 WebHigh-efficiency floating-point neural network inference operators for mobile, server, and Web

llm-infra
roboflow/inference
Python · 2026-08-10 LLM 基础设施 应用 研究原型 Stars 2411 周增 +7

将任意电脑或边缘设备打造为计算机视觉项目的指挥中心Turn any computer or edge device into a command center for your computer vision projects.

agentmultimodalllm-infraengineering
bionic-gpt/bionic-gpt
Rust · 2026-08-03 LLM 基础设施 应用 研究原型 Stars 2349 周增 +7

Bionic 是 ChatGPT 的本地化部署替代方案,在保持严格数据机密性的同时提供生成式 AI 的能力Bionic is an on-premise replacement for ChatGPT, offering the advantages of Generative AI while maintaining strict data confidentiality

llm-infra
beam-cloud/beta9
Go · 2026-09-21 LLM 基础设施 应用 研究原型 Stars 1785 周增 +8

超快速 serverless GPU 推理、沙箱和后台任务Ultrafast serverless GPU inference, sandboxes, and background jobs

llm-infra
trymirai/uzu
Rust · 2026-08-10 LLM 基础设施 应用 研究原型 Stars 1670 周增 +0

高性能 AI 模型推理引擎A high-performance inference engine for AI models

llm-infra
alibaba/rtp-llm
Cuda · 2026-08-05 LLM 基础设施 应用 研究原型 Stars 1296 周增 +0

RTP-LLM:阿里巴巴面向多样化应用的高性能 LLM 推理引擎。RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.

llm-infra
StarlightSearch/EmbedAnything
Rust · 2026-08-11 LLM 基础设施 应用 研究原型 Stars 1294 周增 +5

基于 Rust 🦀 构建的高性能、模块化、内存安全、生产可用的推理、数据接入与索引系统Highly Performant, Modular, Memory Safe and Production-ready Inference, Ingestion and Indexing built in Rust 🦀

ragllm-infraengineeringdatabase
NVIDIA/kvpress
Python · 2026-10-01 LLM 基础设施 应用 研究原型 Stars 1222 周增 +0

让 LLM KV cache 压缩变得更简单LLM KV cache compression made easy

llm-infra
Yu-Yang-Li/StarWhisper
Python · 2026-08-17 LLM 基础设施 应用 研究原型 Stars 325 周增 +0

StarWhisper 天文 LLMs、StarWhisper Telescope、Virtual-GOTTA,以及面向 embodied observing workflow 的天文定制研究 skills。StarWhisper astronomy LLMs, StarWhisper Telescope, Virtual-GOTTA, and astronomy-adapted research skills for embodied observing workflows

agentllm-infra
amitshekhariitbhu/llm-inference-engineering
Markdown · 2026-10-05 LLM 基础设施 应用 实验 Stars 294 周增 +12

逐步学习 LLM 推理工程——从 KV cache、PagedAttention 和连续批处理,到 vLLM、SGLang 和 GPU。Learn LLM Inference Engineering step by step - from KV cache, PagedAttention, and continuous batching to vLLM, SGLang, and GPUs.

llm-infra
walkinglabs/modern-llm-notebook
Jupyter Notebook · 2026-09-30 LLM 基础设施 应用 实验 Stars 222 周增 +9

一门关于现代 LLM 架构、训练与推理的实战课程,配有逐步 PyTorch 实现和可运行的 notebook。A hands-on course on modern LLM architectures, training, and inference, with step-by-step PyTorch implementations and runnable notebooks.

llm-infra
UniversityOfHelsinkiCS/gptwrapper
TypeScript · 2026-08-26 LLM 基础设施 应用 生产可用 Stars 11 周增 +0

为赫尔辛基大学师生打造的 LLM 聊天工具,用于教育与研究。LLM chat built for University of Helsinki staff and students, for education and research.

ragllm-infra
ShaerWare/AI_Secretary_System
Python · 2026-08-11 LLM 基础设施 应用 实验 Stars 10 周增 +0

本地 AI 秘书、技术支持与销售一体化方案,基于 XTTS v2 语音克隆、Vosk/Whisper 实时语音识别与 vLLM + Qwen/Llama 等离线 LLM。配备 Vue 3 完整管理面板、Telegram Bot、网站挂件及 fine-tuning pipeline。支持自托管、数据隐私、短信与电话呼叫。📞 Локальный AI-секретарь, тех. поддержка и менеджер по продажам с клонированием голоса XTTS v2, real-time распознаванием речи (Vosk/Whisper) и offline LLM (vLLM + Qwen/Llama и тп). Полноценная админ-панель (Vue 3), Telegram-бот, виджет для сайта, fine-tuning pipeline. Self-hosted, приватность данных, СМС и телефонные звонки .

agentllm-infraengineering
Dreamer-Toby/STEPQuant
Python · 2026-09-30 LLM 基础设施 应用 实验 Stars 7 周增 +0

STEPQuant: Delta 规则循环状态量化中错误在何时何处重要STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization

llm-infra
konkomaji/bornomala
HTML · 2026-08-18 LLM 基础设施 应用 实验 Stars 5 周增 +0

以孟加拉语优先、面向方言的 LLM 研究。开源孟加拉语 tokenizer,性能优于 Sarvam、AI4Bharat 和 GPT-4o(fertility 1.52,几乎零破损 conjuncts)。非商用,保护孟加拉语及其方言。Bengali-first, dialect-aware LLM research. Open-source Bengali tokenizer that outperforms Sarvam, AI4Bharat, and GPT-4o (fertility 1.52, near-zero broken conjuncts). Non-commercial, preserving Bengali and its dialects.

llm-infra
bojobh609/TurboQuant
Python · 2026-09-21 LLM 基础设施 应用 实验 Stars 2 周增 +0

基于 TurboQuant 优化 FAISS 兼容的向量量化,实现快速、精准的向量检索。Optimize FAISS-compatible vector quantization for fast, accurate vector search with TurboQuant

ragllm-infradatabase
troycheng/learn-inference
Ruby · 2026-08-11 LLM 基础设施 应用 实验 Stars 1 周增 +0

《大模型推理原理与优化》:面向系统/架构/后端研发工程师的模型原理入门课,目标是通俗易懂的解释推理过程,理解原理有助于系统开发/维护工作

llm-infra
nick7nlp/OPDHub
HTML · 2026-10-05 LLM 基础设施 应用 研究原型 Stars 1 周增 +0

论文 A Survey of On-Policy Distillation for Large Language Models(arXiv:2604.00626)的配套网站。Companion website for A Survey of On-Policy Distillation for Large Language Models (arXiv:2604.00626).

llm-infra
Barist3142/deepseek-v4-flash-0731-in-c
C · 2026-10-02 LLM 基础设施 应用 实验 Stars 1 周增 +0

使用纯 C99 MoE 推理引擎在 CPU 上原生运行 DeepSeek-V4-Flash-0731,无需 GPU、CUDA 或 PyTorch。Run native DeepSeek-V4-Flash-0731 on CPU with a pure C99 MoE inference engine — no GPU, CUDA, or PyTorch needed.

llm-infra
Pogud/MegaQwen
Cuda · 2026-09-12 LLM 基础设施 应用 实验 Stars 0 周增 +0

通过 MegaQwen CUDA megakernel 加速 Qwen3-0.6B 推理,在 RTX 3090 上达到 531 tok/s decode,较 HuggingFace 提升 3.9×🚀 Achieve faster Qwen3-0.6B inference with the MegaQwen CUDA megakernel, delivering 531 tok/s decode on RTX 3090—3.9x faster than HuggingFace.

ragllm-infra
murilonerdx/anydoc-studio
JavaScript · 2026-08-14 LLM 基础设施 应用 实验 Stars 0 周增 +0

完全在浏览器中通过 WebAssembly 实现文档转换、检查、OCR 与翻译。保留版式的翻译,全程本地、离线优先。Convert, inspect, OCR, and translate any document entirely in the browser via WebAssembly. Layout-preserving translation, fully local, offline-first.

ragllm-infra