Ray 是一个 AI 计算引擎,由核心分布式运行时和一套用于加速 ML 工作负载的 AI 库组成。Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
仓库/Skill 库
22 个 · LLM 基础设施 · 框架 · AI 核心
让大型 AI 模型更便宜、更快速、更易用。Making large AI models cheaper, faster and more accessible
ncnn 是一个为移动平台优化的高性能神经网络推理框架。ncnn is a high-performance neural network inference framework optimized for the mobile platform
在云端以 OpenAI 兼容 API 端点形式运行任意开源 LLM,例如 DeepSeek 和 Llama。Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.
⚙️🦀 使用 Rust 构建模块化、可扩展的 LLM 应用。⚙️🦀 Build modular and scalable LLM Applications in Rust
Mooncake 是 Moonshot AI 旗下领先 LLM 服务 Kimi 的服务平台。Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
cube studio开源云原生一站式机器学习/深度学习/大模型AI平台,mlops算法链路全流程,算力租赁平台,notebook在线开发,拖拉拽任务流pipeline编排,多机多卡分布式训练,超参搜索,推理服务VGPU虚拟化,边缘计算,标注平台自动化标注,deepseek等大模型sft微调/奖励模型/强化学习训练,vllm/ollama/mindie大模型多机推理,私有知识库,AI模型市场,支持国产cpu/gpu/npu 昇腾生态,支持RDMA,支持pytorch/tf/mxnet/deepspeed/paddle/colossalai/horovod/ray/volcano等分布式
针对加速基础设施和微服务架构优化的生成式 AI 参考工作流。Generative AI reference workflows optimized for accelerated infrastructure and microservice architecture.
任意模型,任意硬件,零妥协。基于 @ziglang / @openxla / MLIR / @bazelbuild 构建。Any model. Any hardware. Zero compromise. Built with @ziglang / @openxla / MLIR / @bazelbuild
Apple Silicon 上最快的本地 AI 引擎。比 Ollama 快 4.2 倍,缓存 TTFT 仅 0.08s,工具调用支持率 100%。内置 17 种工具解析器、prompt cache、推理分离、云端路由。可作为 OpenAI 的即插即用替代,兼容 Claude Code、Cursor、Aider。The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
一个简洁、高性能且可扩展的 Jax LLM!A simple, performant, and scalable Jax LLM!
基于 Envoy Gateway 构建,提供生成式 AI 服务的统一接入管理Manages Unified Access to Generative AI Services built on Envoy Gateway
通过统一 API 接口路由、管理和分析跨多家服务商的 LLM 请求Route, manage, and analyze your LLM requests across multiple providers with a unified API interface.
面向 LLM 和 LVLM 的强大模型压缩框架,适配 NVIDIA GPU 和华为昇腾 NPUA powerful model compression framework for LLMs and LVLMs, adapted for NVIDIA GPUs and Huawei Ascend NPUs.
2026 年度顶级开源 AI 工程平台 Qyvaria KernelTop Open-Source AI Engineering Platform 2026 Qyvaria Kernel
对人类而言,语言是表达的工具;对 AI 而言,语言是推理的基底。For humans, a language is a tool for expression. For AIs, it's a substrate for reasoning.
Gebo.ai —— 开源、企业级、与 AI 供应商无关的平台Gebo.ai The open source Enterprise AI vendor agnostic platform
通过一次 wrap() 调用为 LLM 流水线构建运行时可靠性守卫,可配置地防御常见生产故障Build runtime reliability guards for LLM pipelines with one wrap() call and configurable protection against common production failures
WordPress 共享 AI 基础设施:provider 凭证、实时模型、提示词、标准化请求及 AI-Scribe 和兼容插件的使用记录。Shared AI infrastructure for WordPress: provider credentials, live models, prompts, normalised requests and usage records for AI-Scribe and compatible plugins.
构建安全、有治理、可观测、成本可控的云与 AI 平台能力的实用参考框架。A practical reference framework for building secure, governed, observable, and cost-aware Cloud & AI platform capabilities.
使用纯 C 从零构建 LLM 推理引擎,不依赖任何框架。Build an LLM inference engine from scratch in pure C with no frameworks.
面向 LLM 的阶段感知上下文窗口治理框架。提供不变的上下文长度上限、基于熵的稳定性控制,以及针对降级(碎片化)状态的概率性保证,适用于生产级 LLM 系统。Phase-aware context window governance framework for Large Language Models (LLMs). Provides invariant context length caps, entropy-based stability control, and probabilistic guarantees against degraded (fragmentation) states for production LLM systems.