友好的 AI 交互界面(支持 Ollama、OpenAI API 等)。User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
仓库/Skill 库
68 个 · LLM 基础设施 · 快速增长
21 节课程,开启生成式 AI 应用开发之旅21 Lessons, Get Started Building with Generative AI
面向 LLM 的高吞吐、内存高效的推理与 serving 引擎。A high-throughput and memory-efficient inference and serving engine for LLMs
Python SDK、代理服务器(AI Gateway),以 OpenAI(或原生)格式调用 100+ LLM API,支持成本追踪、guardrails、负载均衡和日志记录。[Bedrock, Azure, OpenAI, VertexAI, Cohere, Anthropic, Sagemaker, HuggingFace, VLLM, NVIDIA NIM]The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthropic, OpenAI, VertexAI, vLLM, Nvidia NIM]
Sub2API 是一款开源中继平台,将 Claude、OpenAI、Gemini 与 Antigravity 订阅统一为单一端点,支持账号共享与费用分摊,兼容原生工具调用。Sub2API 一站式开源中转服务,让 Claude、Openai 、Gemini、Grok订阅统一接入,支持拼车共享,更高效分摊成本,原生工具无缝使用。
SGLang 是一个面向大语言模型和多模态模型的高性能 serving 框架。SGLang is a high-performance serving framework for large language models and multimodal models.
通过 API 访问的免费 LLM 推理资源列表。A list of free LLM inference resources accessible via API.
在现有硬件上运行前沿 MoE 模型——纯 C 实现、零依赖、专家权重从磁盘流式加载。轻量引擎,海量模型。🐦Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦
非自回归的 System 1 决策引擎。单次前向传播即可对任意文本进行类型化选择、打分和 yes/no 决策,支持 100+ 语言,并通过路由按请求选择合适的 checkpoint。Non-autoregressive System 1 decision engine. Typed choice, score and yes/no decisions over any text in a single forward pass, in 100+ languages, with a router that picks the right checkpoint per request.
面向 Metal、CUDA 和 ROCm 的 DeepSeek 4 Flash 与 PRO 本地推理引擎。DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
OpenAI Codex 与 Claude Code 的通用 provider 代理 —— 可在 Codex CLI、App、SDK 及 Claude Code 中使用任意 LLM(Claude、Gemini、Grok、DeepSeek、Ollama…)。Universal provider proxy for OpenAI Codex & Claude Code — use any LLM (Claude, Gemini, Grok, DeepSeek, Ollama…) with Codex CLI, App, SDK, and Claude Code
dsh-routing-suite——注入器 + 路由标准套件:先安装运行时注入器,再安装任务感知的推理模式路由预设(已实测 P1-P23)。dsh-routing-suite — injector + router-standard kit: install the runtime injector first, then the task-aware reasoning-mode router preset (measured P1-P23).
基于 Qwen3.5 构建的小型 Jev 风格决策模型家族,可自行训练与部署。tiny Jev-like family of decision models built on top of Qwen3.5 you can train and run on your own
Gemma 4 26B-A4B 在任意 M 系列 MacBook 上以约 2 GB 内存进行推理。Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
一个 2.78 万亿参数的 Kimi K3,在仅 8.24 GB 内存的单颗 CPU 上运行推理。可移植的 C99:无需 BLAS,无需框架,无需 GPU。A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
在任意消费级硬件上运行 Qwen3.8-Flash-Next:Windows / Linux 一键安装。Strata 推理引擎,本地 localhost 提供 OpenAI / Anthropic API,支持可选图像输入。Qwen3.8-Flash-Next on any consumer hardware: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input.
AgentENV (AENV) 是一个用于大规模运行 Agent 环境的分布式平台。AgentENV (AENV) is a distributed platform for running agent environments at scale.
在家用 3090 上由开源模型生成语义化 if。独立项目,与 Jev 或 TypeSafe 无关。Semantic ifs from open models, on a 3090 at home. Independent; not affiliated with Jev or TypeSafe.
一个精选列表,收录基于 Jev(TypeSafe AI 的 System One 模型,用于类型化决策)构建的公开项目、集成与讨论。A curated list of public projects, integrations, and discussions built on Jev — TypeSafe AI's System One model for typed decisions.
我们能在家里用 3090 跑类似 Jev 的东西吗?Can we run something like Jev on a 3090 at home?
在 dsh 里装上这个插件即可,无需登录、注册或填 API Key,就能使用包括 Muse Spark 1.3、MiMo V2.6 在内的前沿模型——完全免费,不限量。 All you do is install this plugin in dsh: no login, no sign-up, no API key — the frontier models are just there, Muse Spark 1.3 and MiMo V2.6 among them. Completely free, with no usage cap.
本地运行开源决策模型:在兼容 TypeSafe 的 API 后面拉取并服务 Laya、decider、NLI 和 GLiClass。基于 Ollama 的决策模型。Run open decision models locally: pull and serve Laya, decider, NLI and GLiClass behind a TypeSafe-compatible API. Ollama for decision models.
一个精选列表,收录 TypeSafe、System One 模型及 Jev 的官方资源与社区项目。A curated list of official resources and community projects for TypeSafe, System One models, and Jev.
Qwen3.8-27B 在单卡 RTX 3090 上使用 vLLM 部署:64 并发下约 1,000 tok/s(int8 张量核心 GEMM、fp16 DeltaNet 状态),默认采样下单用户约 114 tok/s/贪心约 124 tok/s(MTP 草稿、自输出草稿词表、校准 int4 lm_head、split-KV 校验注意力),150k–262k 上下文;附带补丁、重新量化脚本与基准测试Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks
一款围绕该模型构建的、面向 Apple silicon 的本地推理引擎。A local inference engine for Apple silicon, built around the model.
把腾讯WorkBuddy账号变成 OpenAI 兼容 API 的多账号网关,同时自动完成任务中心全部任务,附 Web 管理面板(账号池可视化 / 积分任务 / 配置热更新)。基于 Sliverkiss/workbuddy2api 的增强分支