统一 AI 模型聚合与分发中心,支持将各类 LLM 跨格式转换为 OpenAI 兼容、Claude 兼容或 Gemini 兼容格式,为个人和企业提供集中式模型管理网关。🍥A unified AI model hub for aggregation & distribution. It supports cross-converting various LLMs into OpenAI-compatible, Claude-compatible, or Gemini-compatible formats. A centralized gateway for personal and enterprise model management. 🍥
仓库/Skill 库
30 个 · LLM 基础设施 · 框架
Ray 是一个 AI 计算引擎,由核心分布式运行时和一套用于加速 ML 工作负载的 AI 库组成。Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
让大型 AI 模型更便宜、更快速、更易用。Making large AI models cheaper, faster and more accessible
LLM API 管理 & 分发系统,支持 OpenAI、Azure、Anthropic Claude、Google Gemini、DeepSeek、字节豆包、ChatGLM、文心一言、讯飞星火、通义千问、360 智脑、腾讯混元等主流模型,统一 API 适配,可用于 key 管理与二次分发。单可执行文件,提供 Docker 镜像,一键部署,开箱即用。LLM API management & key redistribution system, unifying multiple providers under a single API. Single binary, Docker-ready, with an English UI.
SGLang 是一个面向大语言模型和多模态模型的高性能 serving 框架。SGLang is a high-performance serving framework for large language models and multimodal models.
ncnn 是一个为移动平台优化的高性能神经网络推理框架。ncnn is a high-performance neural network inference framework optimized for the mobile platform
在云端以 OpenAI 兼容 API 端点形式运行任意开源 LLM,例如 DeepSeek 和 Llama。Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.
⚙️🦀 使用 Rust 构建模块化、可扩展的 LLM 应用。⚙️🦀 Build modular and scalable LLM Applications in Rust
Mooncake 是 Moonshot AI 旗下领先 LLM 服务 Kimi 的服务平台。Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
cube studio开源云原生一站式机器学习/深度学习/大模型AI平台,mlops算法链路全流程,算力租赁平台,notebook在线开发,拖拉拽任务流pipeline编排,多机多卡分布式训练,超参搜索,推理服务VGPU虚拟化,边缘计算,标注平台自动化标注,deepseek等大模型sft微调/奖励模型/强化学习训练,vllm/ollama/mindie大模型多机推理,私有知识库,AI模型市场,支持国产cpu/gpu/npu 昇腾生态,支持RDMA,支持pytorch/tf/mxnet/deepspeed/paddle/colossalai/horovod/ray/volcano等分布式
一个 2.78 万亿参数的 Kimi K3,在仅 8.24 GB 内存的单颗 CPU 上运行推理。可移植的 C99:无需 BLAS,无需框架,无需 GPU。A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
针对加速基础设施和微服务架构优化的生成式 AI 参考工作流。Generative AI reference workflows optimized for accelerated infrastructure and microservice architecture.
任意模型,任意硬件,零妥协。基于 @ziglang / @openxla / MLIR / @bazelbuild 构建。Any model. Any hardware. Zero compromise. Built with @ziglang / @openxla / MLIR / @bazelbuild
Apple Silicon 上最快的本地 AI 引擎。比 Ollama 快 4.2 倍,缓存 TTFT 仅 0.08s,工具调用支持率 100%。内置 17 种工具解析器、prompt cache、推理分离、云端路由。可作为 OpenAI 的即插即用替代,兼容 Claude Code、Cursor、Aider。The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
AgentENV (AENV) 是一个用于大规模运行 Agent 环境的分布式平台。AgentENV (AENV) is a distributed platform for running agent environments at scale.
一个简洁、高性能且可扩展的 Jax LLM!A simple, performant, and scalable Jax LLM!
基于 Envoy Gateway 构建,提供生成式 AI 服务的统一接入管理Manages Unified Access to Generative AI Services built on Envoy Gateway
通过统一 API 接口路由、管理和分析跨多家服务商的 LLM 请求Route, manage, and analyze your LLM requests across multiple providers with a unified API interface.
面向 LLM 和 LVLM 的强大模型压缩框架,适配 NVIDIA GPU 和华为昇腾 NPUA powerful model compression framework for LLMs and LVLMs, adapted for NVIDIA GPUs and Huawei Ascend NPUs.
OpenSkills:使用任意 LLM 在本地运行 Claude Skills。OpenSkills: Run Claude Skills Locally using any LLM
2026 年度顶级开源 AI 工程平台 Qyvaria KernelTop Open-Source AI Engineering Platform 2026 Qyvaria Kernel
纯 Go 编写的轻量神经网络框架,使用 AVX2 SIMD 内核(GOEXPERIMENT=simd)。A tiny neural-network framework in pure Go with AVX2 SIMD kernels (GOEXPERIMENT=simd)
对人类而言,语言是表达的工具;对 AI 而言,语言是推理的基底。For humans, a language is a tool for expression. For AIs, it's a substrate for reasoning.
Gebo.ai —— 开源、企业级、与 AI 供应商无关的平台Gebo.ai The open source Enterprise AI vendor agnostic platform
通过一次 wrap() 调用为 LLM 流水线构建运行时可靠性守卫,可配置地防御常见生产故障Build runtime reliability guards for LLM pipelines with one wrap() call and configurable protection against common production failures
WordPress 共享 AI 基础设施:provider 凭证、实时模型、提示词、标准化请求及 AI-Scribe 和兼容插件的使用记录。Shared AI infrastructure for WordPress: provider credentials, live models, prompts, normalised requests and usage records for AI-Scribe and compatible plugins.
构建安全、有治理、可观测、成本可控的云与 AI 平台能力的实用参考框架。A practical reference framework for building secure, governed, observable, and cost-aware Cloud & AI platform capabilities.
面向几何感知量子机器学习的可复现研究框架。A reproducible research framework for geometry-aware quantum machine learning.
使用纯 C 从零构建 LLM 推理引擎,不依赖任何框架。Build an LLM inference engine from scratch in pure C with no frameworks.
面向 LLM 的阶段感知上下文窗口治理框架。提供不变的上下文长度上限、基于熵的稳定性控制,以及针对降级(碎片化)状态的概率性保证,适用于生产级 LLM 系统。Phase-aware context window governance framework for Large Language Models (LLMs). Provides invariant context length caps, entropy-based stability control, and probabilistic guarantees against degraded (fragmentation) states for production LLM systems.