⚙️🦀 使用 Rust 构建模块化、可扩展的 LLM 应用。⚙️🦀 Build modular and scalable LLM Applications in Rust
仓库/Skill 库
135 个 · LLM 基础设施 · AI 核心
基于 OpenTelemetry,为你的 GenAI 或 LLM 应用提供开源可观测性。Open-source observability for your GenAI or LLM application, based on OpenTelemetry
中文LLaMA-2 & Alpaca-2大模型二期项目 + 64K超长上下文模型 (Chinese LLaMA-2 & Alpaca-2 LLMs with 64K long context models)
面向并利用基础模型的数据处理!🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
Mooncake 是 Moonshot AI 旗下领先 LLM 服务 Kimi 的服务平台。Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
面向全模态模型的高效推理框架。A framework for efficient model inference with omni-modality models
为开发者精选的 LLMOps 工具 awesome 列表An awesome & curated list of best LLMOps tools for developers
🚀 新兴模型架构的高效实现🚀 Efficient implementations for emerging model architectures
用于高性能 AI 模型服务(vLLM、SGLang)和按需 SSH 访问 GPU 实例的 GPU 集群管理器。A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
cube studio开源云原生一站式机器学习/深度学习/大模型AI平台,mlops算法链路全流程,算力租赁平台,notebook在线开发,拖拉拽任务流pipeline编排,多机多卡分布式训练,超参搜索,推理服务VGPU虚拟化,边缘计算,标注平台自动化标注,deepseek等大模型sft微调/奖励模型/强化学习训练,vllm/ollama/mindie大模型多机推理,私有知识库,AI模型市场,支持国产cpu/gpu/npu 昇腾生态,支持RDMA,支持pytorch/tf/mxnet/deepspeed/paddle/colossalai/horovod/ray/volcano等分布式
高质量且极速的预训练深度学习模型与 demo。Pre-trained Deep Learning models and demos (high quality and extremely fast)
针对加速基础设施和微服务架构优化的生成式 AI 参考工作流。Generative AI reference workflows optimized for accelerated infrastructure and microservice architecture.
ChatGPT、Claude 等 LLM 的所有前端 GUI 客户端汇总。Every front-end GUI client for ChatGPT, Claude, and other LLMs
任意模型,任意硬件,零妥协。基于 @ziglang / @openxla / MLIR / @bazelbuild 构建。Any model. Any hardware. Zero compromise. Built with @ziglang / @openxla / MLIR / @bazelbuild
🚀 通过易用的硬件优化工具,加速 🤗 Transformers、Diffusers、TIMM 和 Sentence Transformers 的推理与训练。🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools
Apple Silicon 上最快的本地 AI 引擎。比 Ollama 快 4.2 倍,缓存 TTFT 仅 0.08s,工具调用支持率 100%。内置 17 种工具解析器、prompt cache、推理分离、云端路由。可作为 OpenAI 的即插即用替代,兼容 Claude Code、Cursor、Aider。The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
📚 OpenVINO™ 的 Jupyter notebook 教程。📚 Jupyter notebook tutorials for OpenVINO™
PyTorch 原生的量化和稀疏化方案,支持训练与推理PyTorch native quantization and sparsity for training and inference
SOTA 低比特 LLM 量化(INT8/FP8/MXFP8/INT4/MXFP4/NVFP4)与稀疏化方案;面向 PyTorch、TensorFlow 与 ONNX Runtime 的领先模型压缩技术SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity; leading model compression techniques on PyTorch, TensorFlow, and ONNX Runtime
为 vLLM 在 Ascend 上的社区维护硬件插件Community maintained hardware plugin for vLLM on Ascend
基于 Rust 的 ONNX 模型快速 ML 推理与训练Fast ML inference & training for ONNX models in Rust
高效的浮点神经网络推理算子,覆盖移动端、服务端和 WebHigh-efficiency floating-point neural network inference operators for mobile, server, and Web
将任意电脑或边缘设备打造为计算机视觉项目的指挥中心Turn any computer or edge device into a command center for your computer vision projects.
一个简洁、高性能且可扩展的 Jax LLM!A simple, performant, and scalable Jax LLM!
Bionic 是 ChatGPT 的本地化部署替代方案,在保持严格数据机密性的同时提供生成式 AI 的能力Bionic is an on-premise replacement for ChatGPT, offering the advantages of Generative AI while maintaining strict data confidentiality
通过统一接口对接 LLM 服务商Communicate with an LLM provider using a single interface
基于 Envoy Gateway 构建,提供生成式 AI 服务的统一接入管理Manages Unified Access to Generative AI Services built on Envoy Gateway
超快速 serverless GPU 推理、沙箱和后台任务Ultrafast serverless GPU inference, sandboxes, and background jobs
可扩展的 LLM 数据预处理与清洗工具集Scalable data pre processing and curation toolkit for LLMs
基于 nvidia-smi 二进制工具的 Nvidia GPU Prometheus exporterNvidia GPU exporter for prometheus using nvidia-smi binary
通过统一 API 接口路由、管理和分析跨多家服务商的 LLM 请求Route, manage, and analyze your LLM requests across multiple providers with a unified API interface.
RTP-LLM:阿里巴巴面向多样化应用的高性能 LLM 推理引擎。RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.