通过统一 API 接口路由、管理和分析跨多家服务商的 LLM 请求Route, manage, and analyze your LLM requests across multiple providers with a unified API interface.
仓库/Skill 库
263 个 · LLM 基础设施
Prompt工程师指南,源自英文版,但增加了AIGC的prompt部分,为了降低同学们的学习门槛,翻译更新
RTP-LLM:阿里巴巴面向多样化应用的高性能 LLM 推理引擎。RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.
基于 Rust 🦀 构建的高性能、模块化、内存安全、生产可用的推理、数据接入与索引系统Highly Performant, Modular, Memory Safe and Production-ready Inference, Ingestion and Indexing built in Rust 🦀
一个精选列表,收录基于 Jev(TypeSafe AI 的 System One 模型,用于类型化决策)构建的公开项目、集成与讨论。A curated list of public projects, integrations, and discussions built on Jev — TypeSafe AI's System One model for typed decisions.
AI 可观测性的 OpenTelemetry 插桩OpenTelemetry Instrumentation for AI Observability
轻量、可移植的 LLM 沙箱运行时(代码解释器)Python 库。Lightweight and portable LLM sandbox runtime (code interpreter) Python library.
面向 LLM 和 LVLM 的强大模型压缩框架,适配 NVIDIA GPU 和华为昇腾 NPUA powerful model compression framework for LLMs and LVLMs, adapted for NVIDIA GPUs and Huawei Ascend NPUs.
我们能在家里用 3090 跑类似 Jev 的东西吗?Can we run something like Jev on a 3090 at home?
关于大语言模型 On-Policy Distillation 的精选论文与资源合集A curated collection of papers and resources on On-Policy Distillation for Large Language Models.
通用 LLM 网关:一个 API 对接所有 LLM。提供兼容 OpenAI/Anthropic 的端点,支持多 provider 转换与智能负载均衡。Universal LLM Gateway: One API, every LLM. OpenAI/Anthropic-compatible endpoints with multi-provider translation and intelligent load-balancing.
在 dsh 里装上这个插件即可,无需登录、注册或填 API Key,就能使用包括 Muse Spark 1.3、MiMo V2.6 在内的前沿模型——完全免费,不限量。 All you do is install this plugin in dsh: no login, no sign-up, no API key — the frontier models are just there, Muse Spark 1.3 and MiMo V2.6 among them. Completely free, with no usage cap.
Krasis 是一个混合 LLM 运行时,专注于在消费级 VRAM 受限硬件上高效运行大模型Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware
本地运行开源决策模型:在兼容 TypeSafe 的 API 后面拉取并服务 Laya、decider、NLI 和 GLiClass。基于 Ollama 的决策模型。Run open decision models locally: pull and serve Laya, decider, NLI and GLiClass behind a TypeSafe-compatible API. Ollama for decision models.
OpenSkills:使用任意 LLM 在本地运行 Claude Skills。OpenSkills: Run Claude Skills Locally using any LLM
一个精选列表,收录 TypeSafe、System One 模型及 Jev 的官方资源与社区项目。A curated list of official resources and community projects for TypeSafe, System One models, and Jev.
Qwen3.8-27B 在单卡 RTX 3090 上使用 vLLM 部署:64 并发下约 1,000 tok/s(int8 张量核心 GEMM、fp16 DeltaNet 状态),默认采样下单用户约 114 tok/s/贪心约 124 tok/s(MTP 草稿、自输出草稿词表、校准 int4 lm_head、split-KV 校验注意力),150k–262k 上下文;附带补丁、重新量化脚本与基准测试Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks
一款围绕该模型构建的、面向 Apple silicon 的本地推理引擎。A local inference engine for Apple silicon, built around the model.
StarWhisper 天文 LLMs、StarWhisper Telescope、Virtual-GOTTA,以及面向 embodied observing workflow 的天文定制研究 skills。StarWhisper astronomy LLMs, StarWhisper Telescope, Virtual-GOTTA, and astronomy-adapted research skills for embodied observing workflows
把腾讯WorkBuddy账号变成 OpenAI 兼容 API 的多账号网关,同时自动完成任务中心全部任务,附 Web 管理面板(账号池可视化 / 积分任务 / 配置热更新)。基于 Sliverkiss/workbuddy2api 的增强分支
逐步学习 LLM 推理工程——从 KV cache、PagedAttention 和连续批处理,到 vLLM、SGLang 和 GPU。Learn LLM Inference Engineering step by step - from KV cache, PagedAttention, and continuous batching to vLLM, SGLang, and GPUs.
Qwen3.8-Flash-Next 部署于单台 DGX Spark(TP=1)Qwen3.8-Flash-Next on ONE DGX Spark (TP=1)
基于 2x DGX Spark 与 TensorFold 的 GLM-5.3-Flash EXL3。GLM-5.3-Flash EXL3 on 2x DGX Spark with TensorFold
开源 agent 评估与可观测性:在一个开放平台上追踪、评估并改进 LLM 应用。🪢 Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.
在 Mac 上以远低于 104 GB 的内存运行 Qwen3.8-Flash-Next(125B MoE,4-bit 下约 104 GB),通过 SSD 流式加载专家。MLX + Swift,兼容 Ollama 的 API。Run Qwen3.8-Flash-Next (125B MoE, 104 GB at 4-bit) on Macs with a fraction of that RAM by streaming experts from SSD. MLX + Swift, Ollama-compatible API.
一门关于现代 LLM 架构、训练与推理的实战课程,配有逐步 PyTorch 实现和可运行的 notebook。A hands-on course on modern LLM architectures, training, and inference, with step-by-step PyTorch implementations and runnable notebooks.
DeepSeek v4 Flash EXL3 运行于单台 DGX SparkDeepSeek v4 Flash EXL3 on one DGX Spark