基于 C/C++ 的 LLM 推理实现LLM inference in C/C++
仓库/Skill 库
31 个 · LLM 基础设施 · 应用 · AI 核心
开源 AI 平台——支持所有 LLM 的 AI Chat,并提供丰富的高级功能。Open Source AI Platform - AI Chat with advanced features that works with every LLM
本地化、开源的 AI 应用构建器,面向高级用户 ✨ v0 / Lovable / Replit / Bolt 的替代方案 🌟 喜欢就点 Star!Local, open-source AI app builder for power users ✨ v0 / Lovable / Replit / Bolt alternative 🌟 Star if you like it!
LMCache:借助最快的 KV Cache 层为你的 LLM 加速。LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Triton Inference Server 提供针对云端和边缘推理优化的解决方案。The Triton Inference Server provides an optimized cloud and edge inferencing solution.
基于 PyTorch 的 YOLOv3、YOLOv3-SPP 和 YOLOv3-tiny 实时目标检测实现,支持训练、验证、推理与多格式导出。PyTorch implementation of YOLOv3, YOLOv3-SPP, and YOLOv3-tiny for real-time object detection with training, validation, inference, and multi-format export.
🚀💪Maximize your efficiency and productivity. The ultimate hub to manage, customize, and share prompts. (English/中文/Español/العربية). 让生产力加倍的 AI 快捷指令。更高效地管理提示词,在分享社区中发现适用于不同场景的灵感。
PyTorch 原生的量化和稀疏化方案,支持训练与推理PyTorch native quantization and sparsity for training and inference
为 vLLM 在 Ascend 上的社区维护硬件插件Community maintained hardware plugin for vLLM on Ascend
基于 Rust 的 ONNX 模型快速 ML 推理与训练Fast ML inference & training for ONNX models in Rust
高效的浮点神经网络推理算子,覆盖移动端、服务端和 WebHigh-efficiency floating-point neural network inference operators for mobile, server, and Web
将任意电脑或边缘设备打造为计算机视觉项目的指挥中心Turn any computer or edge device into a command center for your computer vision projects.
Bionic 是 ChatGPT 的本地化部署替代方案,在保持严格数据机密性的同时提供生成式 AI 的能力Bionic is an on-premise replacement for ChatGPT, offering the advantages of Generative AI while maintaining strict data confidentiality
超快速 serverless GPU 推理、沙箱和后台任务Ultrafast serverless GPU inference, sandboxes, and background jobs
RTP-LLM:阿里巴巴面向多样化应用的高性能 LLM 推理引擎。RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.
基于 Rust 🦀 构建的高性能、模块化、内存安全、生产可用的推理、数据接入与索引系统Highly Performant, Modular, Memory Safe and Production-ready Inference, Ingestion and Indexing built in Rust 🦀
StarWhisper 天文 LLMs、StarWhisper Telescope、Virtual-GOTTA,以及面向 embodied observing workflow 的天文定制研究 skills。StarWhisper astronomy LLMs, StarWhisper Telescope, Virtual-GOTTA, and astronomy-adapted research skills for embodied observing workflows
一门实战课程,用 PyTorch 从零构建现代 LLM,包含 26 个可运行的 Jupyter Notebook,涵盖 tokenizer、attention、MoE、RLHF、推理、评估和蒸馏。A hands-on course for building modern LLMs from scratch in PyTorch, with 26 runnable Jupyter Notebooks covering tokenizers, attention, MoE, RLHF, inference, evaluation, and distillation.
本地 AI 秘书、技术支持与销售一体化方案,基于 XTTS v2 语音克隆、Vosk/Whisper 实时语音识别与 vLLM + Qwen/Llama 等离线 LLM。配备 Vue 3 完整管理面板、Telegram Bot、网站挂件及 fine-tuning pipeline。支持自托管、数据隐私、短信与电话呼叫。📞 Локальный AI-секретарь, тех. поддержка и менеджер по продажам с клонированием голоса XTTS v2, real-time распознаванием речи (Vosk/Whisper) и offline LLM (vLLM + Qwen/Llama и тп). Полноценная админ-панель (Vue 3), Telegram-бот, виджет для сайта, fine-tuning pipeline. Self-hosted, приватность данных, СМС и телефонные звонки .
以孟加拉语优先、面向方言的 LLM 研究。开源孟加拉语 tokenizer,性能优于 Sarvam、AI4Bharat 和 GPT-4o(fertility 1.52,几乎零破损 conjuncts)。非商用,保护孟加拉语及其方言。Bengali-first, dialect-aware LLM research. Open-source Bengali tokenizer that outperforms Sarvam, AI4Bharat, and GPT-4o (fertility 1.52, near-zero broken conjuncts). Non-commercial, preserving Bengali and its dialects.
《大模型推理原理与优化》:面向系统/架构/后端研发工程师的模型原理入门课,目标是通俗易懂的解释推理过程,理解原理有助于系统开发/维护工作
通过 MegaQwen CUDA megakernel 加速 Qwen3-0.6B 推理,在 RTX 3090 上达到 531 tok/s decode,较 HuggingFace 提升 3.9×🚀 Achieve faster Qwen3-0.6B inference with the MegaQwen CUDA megakernel, delivering 531 tok/s decode on RTX 3090—3.9x faster than HuggingFace.
完全在浏览器中通过 WebAssembly 实现文档转换、检查、OCR 与翻译。保留版式的翻译,全程本地、离线优先。Convert, inspect, OCR, and translate any document entirely in the browser via WebAssembly. Layout-preserving translation, fully local, offline-first.
metaScreener——基于插件的桌面应用,用于 human-in-the-loop systematic literature screening。在顺序可审计的 pipeline 中结合确定性启发式过滤与 LLM 推理,通过 SHA-256 校验包实现完全可复现。MIT 协议。metaScreener — a plugin-based desktop application for human-in-the-loop systematic literature screening. Combines deterministic heuristic filters with LLM inference in a sequential, auditable pipeline. SHA-256 verified bundles for full reproducibility. MIT licensed.
🛠 通过此硬件插件提升 vLLM 在 Kunlun XPU 上的性能,无缝集成主流 AI 模型并优化执行效率🛠 Enhance vLLM performance on Kunlun XPU with this hardware plugin, offering seamless integration for popular AI models and optimized execution.
使用纯 C99 MoE 推理引擎在 CPU 上原生运行 DeepSeek-V4-Flash-0731,无需 GPU、CUDA 或 PyTorch。Run native DeepSeek-V4-Flash-0731 on CPU with a pure C99 MoE inference engine — no GPU, CUDA, or PyTorch needed.