基于 nvidia-smi 二进制工具的 Nvidia GPU Prometheus exporterNvidia GPU exporter for prometheus using nvidia-smi binary
仓库/Skill 库
149 个 · LLM 基础设施 · AI 核心
通过统一 API 接口路由、管理和分析跨多家服务商的 LLM 请求Route, manage, and analyze your LLM requests across multiple providers with a unified API interface.
RTP-LLM:阿里巴巴面向多样化应用的高性能 LLM 推理引擎。RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.
基于 Rust 🦀 构建的高性能、模块化、内存安全、生产可用的推理、数据接入与索引系统Highly Performant, Modular, Memory Safe and Production-ready Inference, Ingestion and Indexing built in Rust 🦀
轻量、可移植的 LLM 沙箱运行时(代码解释器)Python 库。Lightweight and portable LLM sandbox runtime (code interpreter) Python library.
面向 LLM 和 LVLM 的强大模型压缩框架,适配 NVIDIA GPU 和华为昇腾 NPUA powerful model compression framework for LLMs and LVLMs, adapted for NVIDIA GPUs and Huawei Ascend NPUs.
关于大语言模型 On-Policy Distillation 的精选论文与资源合集A curated collection of papers and resources on On-Policy Distillation for Large Language Models.
通用 LLM 网关:一个 API 对接所有 LLM。提供兼容 OpenAI/Anthropic 的端点,支持多 provider 转换与智能负载均衡。Universal LLM Gateway: One API, every LLM. OpenAI/Anthropic-compatible endpoints with multi-provider translation and intelligent load-balancing.
Krasis 是一个混合 LLM 运行时,专注于在消费级 VRAM 受限硬件上高效运行大模型Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware
StarWhisper 天文 LLMs、StarWhisper Telescope、Virtual-GOTTA,以及面向 embodied observing workflow 的天文定制研究 skills。StarWhisper astronomy LLMs, StarWhisper Telescope, Virtual-GOTTA, and astronomy-adapted research skills for embodied observing workflows
逐步学习 LLM 推理工程——从 KV cache、PagedAttention 和连续批处理,到 vLLM、SGLang 和 GPU。Learn LLM Inference Engineering step by step - from KV cache, PagedAttention, and continuous batching to vLLM, SGLang, and GPUs.
开源 agent 评估与可观测性:在一个开放平台上追踪、评估并改进 LLM 应用。🪢 Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.
一门关于现代 LLM 架构、训练与推理的实战课程,配有逐步 PyTorch 实现和可运行的 notebook。A hands-on course on modern LLM architectures, training, and inference, with step-by-step PyTorch implementations and runnable notebooks.
用于探索 Transformer 架构、张量及实时 LLM 推理的交互式 3D 可视化平台。Interactive 3D visualization platform for exploring transformer architectures, tensors, and real-time LLM inference.
LLM7.io 提供单一 API 网关,可连接来自多家供应商的众多领先 AI 模型LLM7.io offers a single API gateway that connects you to a wide array of leading AI models from various providers.
2026 年度顶级开源 AI 工程平台 Qyvaria KernelTop Open-Source AI Engineering Platform 2026 Qyvaria Kernel
一款便捷的 lib,用于流畅地与大语言模型(LLM)交互及构建 AI 应用。A handy lib for smooth interaction with large language models (LLMs) and crafting AI apps.
🪢 Langfuse API 的自动生成 Java 客户端。🪢 Auto-generated Java Client for Langfuse API
ELM 是将 LLM 应用于能源研究的一系列工具集合。ELM is a collection of utilities to apply Large Language Models (LLMs) to energy research.
类似 unix du 的命令行工具,用于统计每个文件和目录的 token 使用量a unix-like du command line tool to count token usage per files and directories
repo-map 生成由 LLM 增强的软件仓库摘要与分析,为开发者提供关于项目结构、文件用途以及跨编程语言潜在考量的洞察。repo-map generates LLM-enhanced summaries and analysis of software repositories, providing developers with valuable insights into project structures, file purposes, and potential considerations across various programming languages.
对人类而言,语言是表达的工具;对 AI 而言,语言是推理的基底。For humans, a language is a tool for expression. For AIs, it's a substrate for reasoning.
AI 原生的结构化数据 wire 格式。在每个前沿模型上实现 100% 理解,比 JSON 减少 50-92% token,跨 17 种格式完成 43B+ 无损往返。Spec v3.4 Stable。The AI-native wire format for structured data. 100% comprehension on every frontier model. 50-92% fewer tokens than JSON. 43B+ lossless round-trips across 17 formats. Spec v3.4 Stable.
精选 AI 模型及其 API 提供商列表,完全无需信用卡即可使用,欢迎贡献!Curated list of AI Models with their API Providers that you never ever require a Credit Card for. Feel free to Contribute!
一个以证据为引领的六语种 LLM 实战手册:包含可迁移的核心、Codex 旗舰路线,以及 ChatGPT、Claude Code、Gemini、DeepSeek 和 Grok 的适配器。An evidence-led, six-language LLM playbook: the transferable core, the Codex flagship track, and adapters for ChatGPT, Claude Code, Gemini, DeepSeek, and Grok.
Run the native 284B-A13B DeepSeek-V4-Flash-0731 LLM locally on one laptop CPU: pure C, 8 GB RAM minimum, no GPU, best TPOT 0.892 s/token, resident OpenAI-compatible API with function tools. | 在笔记本单颗 CPU 上本地运行原生 284B-A13B DeepSeek-V4-Flash-0731 大模型:纯 C,最低 8 GB 内存,无需 GPU,最优 TPOT 0.892 秒/token,支持模型常驻的 OpenAI 兼容接口与函数工具。
🦙 使用 Ollama CLI 配置 GitHub Actions。🦙 Set up GitHub Actions with Ollama CLI.
为赫尔辛基大学师生打造的 LLM 聊天工具,用于教育与研究。LLM chat built for University of Helsinki staff and students, for education and research.
本地 AI 秘书、技术支持与销售一体化方案,基于 XTTS v2 语音克隆、Vosk/Whisper 实时语音识别与 vLLM + Qwen/Llama 等离线 LLM。配备 Vue 3 完整管理面板、Telegram Bot、网站挂件及 fine-tuning pipeline。支持自托管、数据隐私、短信与电话呼叫。📞 Локальный AI-секретарь, тех. поддержка и менеджер по продажам с клонированием голоса XTTS v2, real-time распознаванием речи (Vosk/Whisper) и offline LLM (vLLM + Qwen/Llama и тп). Полноценная админ-панель (Vue 3), Telegram-бот, виджет для сайта, fine-tuning pipeline. Self-hosted, приватность данных, СМС и телефонные звонки .
运行 Ollama 本地 LLM 服务的 Docker 镜像。默认安全,所有 API 请求需 Bearer token(首次启动时自动生成)。OpenAI 兼容 API。支持首次启动模型预拉取、NVIDIA GPU (CUDA) 加速和持久化模型存储。多架构:amd64、arm64。Docker image to run an Ollama local LLM server. Secure by default, all API requests require a Bearer token (auto-generated on first start). OpenAI-compatible API. Supports first-start model pre-pull, NVIDIA GPU (CUDA) acceleration, and persistent model storage. Multi-arch: amd64, arm64.
并行运行 Claude Code、Codex 与 Gemini,并支持彼此交接任务。面向 AI CLI 的便携 Windows 终端。Run Claude Code, Codex and Gemini side by side — and let them hand work to each other. Portable Windows terminal for AI CLIs.
Gebo.ai —— 开源、企业级、与 AI 供应商无关的平台Gebo.ai The open source Enterprise AI vendor agnostic platform
STEPQuant: Delta 规则循环状态量化中错误在何时何处重要STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization
💻 在 macOS 上借助 Metal GPU 实现 Qwen3 Transformer 模型,获得加速且高效的性能,并支持关键架构特性💻 Implement Qwen3 transformer model on macOS using Metal GPU for accelerated, efficient performance with support for key architecture features.
本仓库包含用于执行不同任务的 Gen AI 项目temThis repository contains Gen AI projects that performs different tasks