为 LLM 优化推理代理Optimizing inference proxy for LLMs
仓库/Skill 库
199 个 · LLM 基础设施
针对加速基础设施和微服务架构优化的生成式 AI 参考工作流。Generative AI reference workflows optimized for accelerated infrastructure and microservice architecture.
Fengshenbang-LM(封神榜大模型)是IDEA研究院认知计算与自然语言研究中心主导的大模型开源体系,成为中文AIGC和认知智能的基础设施。
ChatGPT、Claude 等 LLM 的所有前端 GUI 客户端汇总。Every front-end GUI client for ChatGPT, Claude, and other LLMs
任意模型,任意硬件,零妥协。基于 @ziglang / @openxla / MLIR / @bazelbuild 构建。Any model. Any hardware. Zero compromise. Built with @ziglang / @openxla / MLIR / @bazelbuild
🚀 通过易用的硬件优化工具,加速 🤗 Transformers、Diffusers、TIMM 和 Sentence Transformers 的推理与训练。🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools
Apple Silicon 上最快的本地 AI 引擎。比 Ollama 快 4.2 倍,缓存 TTFT 仅 0.08s,工具调用支持率 100%。内置 17 种工具解析器、prompt cache、推理分离、云端路由。可作为 OpenAI 的即插即用替代,兼容 Claude Code、Cursor、Aider。The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
📚 OpenVINO™ 的 Jupyter notebook 教程。📚 Jupyter notebook tutorials for OpenVINO™
AgentENV (AENV) 是一个用于大规模运行 Agent 环境的分布式平台。AgentENV (AENV) is a distributed platform for running agent environments at scale.
PyTorch 原生的量化和稀疏化方案,支持训练与推理PyTorch native quantization and sparsity for training and inference
SOTA 低比特 LLM 量化(INT8/FP8/MXFP8/INT4/MXFP4/NVFP4)与稀疏化方案;面向 PyTorch、TensorFlow 与 ONNX Runtime 的领先模型压缩技术SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity; leading model compression techniques on PyTorch, TensorFlow, and ONNX Runtime
为 vLLM 在 Ascend 上的社区维护硬件插件Community maintained hardware plugin for vLLM on Ascend
基于 Rust 的 ONNX 模型快速 ML 推理与训练Fast ML inference & training for ONNX models in Rust
高效的浮点神经网络推理算子,覆盖移动端、服务端和 WebHigh-efficiency floating-point neural network inference operators for mobile, server, and Web
将任意电脑或边缘设备打造为计算机视觉项目的指挥中心Turn any computer or edge device into a command center for your computer vision projects.
一个简洁、高性能且可扩展的 Jax LLM!A simple, performant, and scalable Jax LLM!
Bionic 是 ChatGPT 的本地化部署替代方案,在保持严格数据机密性的同时提供生成式 AI 的能力Bionic is an on-premise replacement for ChatGPT, offering the advantages of Generative AI while maintaining strict data confidentiality
通过统一接口对接 LLM 服务商Communicate with an LLM provider using a single interface
基于 Envoy Gateway 构建,提供生成式 AI 服务的统一接入管理Manages Unified Access to Generative AI Services built on Envoy Gateway
超快速 serverless GPU 推理、沙箱和后台任务Ultrafast serverless GPU inference, sandboxes, and background jobs
可扩展的 LLM 数据预处理与清洗工具集Scalable data pre processing and curation toolkit for LLMs
面向 AI 辅助草稿的可读性与自然节奏改进的开源 pipeline 与参考实现。Open-source pipeline and reference implementations for improving the readability and natural cadence of AI-assisted drafts.
基于 nvidia-smi 二进制工具的 Nvidia GPU Prometheus exporterNvidia GPU exporter for prometheus using nvidia-smi binary
通过统一 API 接口路由、管理和分析跨多家服务商的 LLM 请求Route, manage, and analyze your LLM requests across multiple providers with a unified API interface.
dsh-routing-suite——注入器 + 路由标准套件:先安装运行时注入器,再安装任务感知的推理模式路由预设(已实测 P1-P23)。dsh-routing-suite — injector + router-standard kit: install the runtime injector first, then the task-aware reasoning-mode router preset (measured P1-P23).
Prompt工程师指南,源自英文版,但增加了AIGC的prompt部分,为了降低同学们的学习门槛,翻译更新
RTP-LLM:阿里巴巴面向多样化应用的高性能 LLM 推理引擎。RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.
基于 Rust 🦀 构建的高性能、模块化、内存安全、生产可用的推理、数据接入与索引系统Highly Performant, Modular, Memory Safe and Production-ready Inference, Ingestion and Indexing built in Rust 🦀
轻量、可移植的 LLM 沙箱运行时(代码解释器)Python 库。Lightweight and portable LLM sandbox runtime (code interpreter) Python library.
面向 LLM 和 LVLM 的强大模型压缩框架,适配 NVIDIA GPU 和华为昇腾 NPUA powerful model compression framework for LLMs and LVLMs, adapted for NVIDIA GPUs and Huawei Ascend NPUs.
Krasis 是一个混合 LLM 运行时,专注于在消费级 VRAM 受限硬件上高效运行大模型Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware