快速上手 Kimi-K2.6、GLM-5.2、MiniMax、DeepSeek、gpt-oss、Qwen、Gemma 等模型Get up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
仓库/Skill 库
69 个 · LLM 基础设施 · 工具
GPT4All:在任意设备上运行本地 LLM。开源且可用于商业用途。GPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use.
免费大模型API,支持免费调用GPT、DeepSeek等主流模型,免费额度10000点,每日刷新!另付费价格最低官方1-2折!
Sub2API 是一款开源中继平台,将 Claude、OpenAI、Gemini 与 Antigravity 订阅统一为单一端点,支持账号共享与费用分摊,兼容原生工具调用。Sub2API 一站式开源中转服务,让 Claude、Openai 、Gemini、Grok订阅统一接入,支持拼车共享,更高效分摊成本,原生工具无缝使用。
一个 AI 提示词优化器,用于编写更好的提示词并获得更好的 AI 结果。An AI prompt optimizer for writing better prompts and getting better AI results.
📦 Repomix 是一个强大的工具,能将整个仓库打包为单个对 AI 友好的文件。非常适合需要将代码库喂给 LLM 或其他 AI 工具(如 Claude、ChatGPT、DeepSeek、Perplexity、Gemini、Gemma、Llama、Grok 等)的场景。📦 Repomix is a powerful tool that packs your entire repository into a single, AI-friendly file. Perfect for when you need to feed your codebase to Large Language Models (LLMs) or other AI tools like Claude, ChatGPT, DeepSeek, Perplexity, Gemini, Gemma, Llama, Grok, and more.
无限免费 AI 编程。通过 40+ 提供商将 Claude Code、Codex、Cursor、Cline、Copilot、Antigravity 连接至免费的 Claude/GPT/Gemini。支持自动回退,RTK 减少 40% tokens,永不触及限额。Unlimited FREE AI coding. Connect Claude Code, Codex, Cursor, Cline, Copilot, Antigravity to FREE Claude/GPT/Gemini via 40+ providers. Auto-fallback, RTK -40% tokens, never hit limits.
在现有硬件上运行前沿 MoE 模型——纯 C 实现、零依赖、专家权重从磁盘流式加载。轻量引擎,海量模型。🐦Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦
OpenVINO™ 是用于优化和部署 AI 推理的开源工具包OpenVINO™ is an open source toolkit for optimizing and deploying AI inference
OpenAI Codex 与 Claude Code 的通用 provider 代理 —— 可在 Codex CLI、App、SDK 及 Claude Code 中使用任意 LLM(Claude、Gemini、Grok、DeepSeek、Ollama…)。Universal provider proxy for OpenAI Codex & Claude Code — use any LLM (Claude, Gemini, Grok, DeepSeek, Ollama…) with Codex CLI, App, SDK, and Claude Code
ODS V3 预发布:在 V3 正式发布前进行公开测试与打磨。将你的 PC、Mac 或 Linux 机器变为私有 AI 服务器。ODS V3 Pre-Release: Public testing and refinement ahead of the official V3 launch. Turn your PC, Mac, or Linux box into a private AI server.
dsh-routing-suite——注入器 + 路由标准套件:先安装运行时注入器,再安装任务感知的推理模式路由预设(已实测 P1-P23)。dsh-routing-suite — injector + router-standard kit: install the runtime injector first, then the task-aware reasoning-mode router preset (measured P1-P23).
🚀 通过易用的硬件优化工具,加速 🤗 Transformers、Diffusers、TIMM 和 Sentence Transformers 的推理与训练。🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools
🎨 NeMo Data Designer:从零或从种子数据生成高质量合成数据。🎨 NeMo Data Designer: Generate high-quality synthetic data from scratch or from seed data.
可扩展的 LLM 数据预处理与清洗工具集Scalable data pre processing and curation toolkit for LLMs
在家用 3090 上由开源模型生成语义化 if。独立项目,与 Jev 或 TypeSafe 无关。Semantic ifs from open models, on a 3090 at home. Independent; not affiliated with Jev or TypeSafe.
面向 AI 辅助草稿的可读性与自然节奏改进的开源 pipeline 与参考实现。Open-source pipeline and reference implementations for improving the readability and natural cadence of AI-assisted drafts.
基于 nvidia-smi 二进制工具的 Nvidia GPU Prometheus exporterNvidia GPU exporter for prometheus using nvidia-smi binary
我们能在家里用 3090 跑类似 Jev 的东西吗?Can we run something like Jev on a 3090 at home?
通用 LLM 网关:一个 API 对接所有 LLM。提供兼容 OpenAI/Anthropic 的端点,支持多 provider 转换与智能负载均衡。Universal LLM Gateway: One API, every LLM. OpenAI/Anthropic-compatible endpoints with multi-provider translation and intelligent load-balancing.
在 dsh 里装上这个插件即可,无需登录、注册或填 API Key,就能使用包括 Muse Spark 1.3、MiMo V2.6 在内的前沿模型——完全免费,不限量。 All you do is install this plugin in dsh: no login, no sign-up, no API key — the frontier models are just there, Muse Spark 1.3 and MiMo V2.6 among them. Completely free, with no usage cap.
Krasis 是一个混合 LLM 运行时,专注于在消费级 VRAM 受限硬件上高效运行大模型Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware
本地运行开源决策模型:在兼容 TypeSafe 的 API 后面拉取并服务 Laya、decider、NLI 和 GLiClass。基于 Ollama 的决策模型。Run open decision models locally: pull and serve Laya, decider, NLI and GLiClass behind a TypeSafe-compatible API. Ollama for decision models.
把腾讯WorkBuddy账号变成 OpenAI 兼容 API 的多账号网关,同时自动完成任务中心全部任务,附 Web 管理面板(账号池可视化 / 积分任务 / 配置热更新)。基于 Sliverkiss/workbuddy2api 的增强分支
Qwen3.8-Flash-Next 部署于单台 DGX Spark(TP=1)Qwen3.8-Flash-Next on ONE DGX Spark (TP=1)
开源 agent 评估与可观测性:在一个开放平台上追踪、评估并改进 LLM 应用。🪢 Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.
在 Mac 上以远低于 104 GB 的内存运行 Qwen3.8-Flash-Next(125B MoE,4-bit 下约 104 GB),通过 SSD 流式加载专家。MLX + Swift,兼容 Ollama 的 API。Run Qwen3.8-Flash-Next (125B MoE, 104 GB at 4-bit) on Macs with a fraction of that RAM by streaming experts from SSD. MLX + Swift, Ollama-compatible API.
LLM7.io 提供单一 API 网关,可连接来自多家供应商的众多领先 AI 模型LLM7.io offers a single API gateway that connects you to a wide array of leading AI models from various providers.
Codex 的每轮模型与推理路由,由 Jev(TypeSafe System One)驱动:为每一轮选择模型、思考深度和速度模式。Per-turn model & reasoning routing for Codex, driven by Jev (TypeSafe System One): picks the model, thinking depth and speed mode for every turn.
在单台 DGX Spark 上运行 Qwen3.8 Flash Next(TensorFold)Qwen3.8 Flash Next on one DGX Spark (TensorFold)
腾讯 CodeBuddy 账号池管理控制台 + OpenAI 兼容反代网关。扫码纳管账号、自动签到、多密钥分发、IP 管控、调用日志与用量统计。UI 对标 linux-do/cdk。
DeepSeek-V4.1-Flash 部署于 3–4 块 NVIDIA DGX Spark。DeepSeek-V4.1-Flash on 3-4x NVIDIA DGX Sparks
将你的 CodeBuddy 订阅作为本地 OpenAI API 使用。Use your CodeBuddy subscription as local OpenAI APIs.