结构化输出(Structured Outputs)Structured Outputs
仓库/Skill 库
199 个 · LLM 基础设施
🎨 TypeScript 的穷尽式 Pattern Matching 库,具备智能类型推断。🎨 The exhaustive Pattern Matching library for TypeScript, with smart type inference.
NVIDIA® TensorRT™ 是在 NVIDIA GPU 上进行高性能深度学习推理的 SDK。本仓库包含 TensorRT 的开源组件。NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT.
在云端以 OpenAI 兼容 API 端点形式运行任意开源 LLM,例如 DeepSeek 和 Llama。Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.
复旦大学开源的工具增强对话语言模型An open-source tool-augmented conversational language model from Fudan University
LMCache:借助最快的 KV Cache 层为你的 LLM 加速。LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
示例 📓 Jupyter notebooks,演示如何使用 🧠 Amazon SageMaker 构建、训练和部署机器学习模型。Example 📓 Jupyter notebooks that demonstrate how to build, train, and deploy machine learning models using 🧠 Amazon SageMaker.
Triton Inference Server 提供针对云端和边缘推理优化的解决方案。The Triton Inference Server provides an optimized cloud and edge inferencing solution.
OpenVINO™ 是用于优化和部署 AI 推理的开源工具包OpenVINO™ is an open source toolkit for optimizing and deploying AI inference
基于 PyTorch 的 YOLOv3、YOLOv3-SPP 和 YOLOv3-tiny 实时目标检测实现,支持训练、验证、推理与多格式导出。PyTorch implementation of YOLOv3, YOLOv3-SPP, and YOLOv3-tiny for real-time object detection with training, validation, inference, and multi-format export.
用于本地运行 AI 的生产就绪工具包。Production ready toolkit to run AI locally
LLM 实战指南资源的精选清单(LLMs Tree、示例、论文)A curated list of practical guide resources of LLMs (LLMs Tree, Examples, Papers)
面向本地部署的高速 LLM 服务High-speed Large Language Model Serving for Local Deployment
OpenAI Codex 与 Claude Code 的通用 provider 代理 —— 可在 Codex CLI、App、SDK 及 Claude Code 中使用任意 LLM(Claude、Gemini、Grok、DeepSeek、Ollama…)。Universal provider proxy for OpenAI Codex & Claude Code — use any LLM (Claude, Gemini, Grok, DeepSeek, Ollama…) with Codex CLI, App, SDK, and Claude Code
提供 AI 应用和模型服务的最简方式 —— 构建模型推理 API、任务队列、LLM 应用、多模型 pipeline 等。The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
🚀💪Maximize your efficiency and productivity. The ultimate hub to manage, customize, and share prompts. (English/中文/Español/العربية). 让生产力加倍的 AI 快捷指令。更高效地管理提示词,在分享社区中发现适用于不同场景的灵感。
Transformer 可视化解析:通过交互式可视化学习 LLM Transformer 模型的工作原理。Transformer Explained Visually: Learn How LLM Transformer Models Work with Interactive Visualization
⚙️🦀 使用 Rust 构建模块化、可扩展的 LLM 应用。⚙️🦀 Build modular and scalable LLM Applications in Rust
基于 OpenTelemetry,为你的 GenAI 或 LLM 应用提供开源可观测性。Open-source observability for your GenAI or LLM application, based on OpenTelemetry
中文LLaMA-2 & Alpaca-2大模型二期项目 + 64K超长上下文模型 (Chinese LLaMA-2 & Alpaca-2 LLMs with 64K long context models)
面向并利用基础模型的数据处理!🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
Mooncake 是 Moonshot AI 旗下领先 LLM 服务 Kimi 的服务平台。Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
面向全模态模型的高效推理框架。A framework for efficient model inference with omni-modality models
为开发者精选的 LLMOps 工具 awesome 列表An awesome & curated list of best LLMOps tools for developers
Gemma 4 26B-A4B 在任意 M 系列 MacBook 上以约 2 GB 内存进行推理。Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
🚀 新兴模型架构的高效实现🚀 Efficient implementations for emerging model architectures
用于高性能 AI 模型服务(vLLM、SGLang)和按需 SSH 访问 GPU 实例的 GPU 集群管理器。A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
cube studio开源云原生一站式机器学习/深度学习/大模型AI平台,mlops算法链路全流程,算力租赁平台,notebook在线开发,拖拉拽任务流pipeline编排,多机多卡分布式训练,超参搜索,推理服务VGPU虚拟化,边缘计算,标注平台自动化标注,deepseek等大模型sft微调/奖励模型/强化学习训练,vllm/ollama/mindie大模型多机推理,私有知识库,AI模型市场,支持国产cpu/gpu/npu 昇腾生态,支持RDMA,支持pytorch/tf/mxnet/deepspeed/paddle/colossalai/horovod/ray/volcano等分布式
一个 2.78 万亿参数的 Kimi K3,在仅 8.24 GB 内存的单颗 CPU 上运行推理。可移植的 C99:无需 BLAS,无需框架,无需 GPU。A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
精选的 LLM、AI 绘画等领域的教程和资源。Curated tutorials and resources for Large Language Models, AI Painting, and more.
高质量且极速的预训练深度学习模型与 demo。Pre-trained Deep Learning models and demos (high quality and extremely fast)