快速上手 Kimi-K2.6、GLM-5.2、MiniMax、DeepSeek、gpt-oss、Qwen、Gemma 等模型Get up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
仓库/Skill 库
32 个 · LLM 基础设施 · 工具 · AI 核心
GPT4All:在任意设备上运行本地 LLM。开源且可用于商业用途。GPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use.
一个 AI 提示词优化器,用于编写更好的提示词并获得更好的 AI 结果。An AI prompt optimizer for writing better prompts and getting better AI results.
📦 Repomix 是一个强大的工具,能将整个仓库打包为单个对 AI 友好的文件。非常适合需要将代码库喂给 LLM 或其他 AI 工具(如 Claude、ChatGPT、DeepSeek、Perplexity、Gemini、Gemma、Llama、Grok 等)的场景。📦 Repomix is a powerful tool that packs your entire repository into a single, AI-friendly file. Perfect for when you need to feed your codebase to Large Language Models (LLMs) or other AI tools like Claude, ChatGPT, DeepSeek, Perplexity, Gemini, Gemma, Llama, Grok, and more.
OpenVINO™ 是用于优化和部署 AI 推理的开源工具包OpenVINO™ is an open source toolkit for optimizing and deploying AI inference
ODS V3 预发布:在 V3 正式发布前进行公开测试与打磨。将你的 PC、Mac 或 Linux 机器变为私有 AI 服务器。ODS V3 Pre-Release: Public testing and refinement ahead of the official V3 launch. Turn your PC, Mac, or Linux box into a private AI server.
🚀 通过易用的硬件优化工具,加速 🤗 Transformers、Diffusers、TIMM 和 Sentence Transformers 的推理与训练。🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools
可扩展的 LLM 数据预处理与清洗工具集Scalable data pre processing and curation toolkit for LLMs
基于 nvidia-smi 二进制工具的 Nvidia GPU Prometheus exporterNvidia GPU exporter for prometheus using nvidia-smi binary
通用 LLM 网关:一个 API 对接所有 LLM。提供兼容 OpenAI/Anthropic 的端点,支持多 provider 转换与智能负载均衡。Universal LLM Gateway: One API, every LLM. OpenAI/Anthropic-compatible endpoints with multi-provider translation and intelligent load-balancing.
Krasis 是一个混合 LLM 运行时,专注于在消费级 VRAM 受限硬件上高效运行大模型Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware
开源 agent 评估与可观测性:在一个开放平台上追踪、评估并改进 LLM 应用。🪢 Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.
LLM7.io 提供单一 API 网关,可连接来自多家供应商的众多领先 AI 模型LLM7.io offers a single API gateway that connects you to a wide array of leading AI models from various providers.
类似 unix du 的命令行工具,用于统计每个文件和目录的 token 使用量a unix-like du command line tool to count token usage per files and directories
repo-map 生成由 LLM 增强的软件仓库摘要与分析,为开发者提供关于项目结构、文件用途以及跨编程语言潜在考量的洞察。repo-map generates LLM-enhanced summaries and analysis of software repositories, providing developers with valuable insights into project structures, file purposes, and potential considerations across various programming languages.
Run the native 284B-A13B DeepSeek-V4-Flash-0731 LLM locally on one laptop CPU: pure C, 8 GB RAM minimum, no GPU, best TPOT 0.892 s/token, resident OpenAI-compatible API with function tools. | 在笔记本单颗 CPU 上本地运行原生 284B-A13B DeepSeek-V4-Flash-0731 大模型:纯 C,最低 8 GB 内存,无需 GPU,最优 TPOT 0.892 秒/token,支持模型常驻的 OpenAI 兼容接口与函数工具。
🦙 使用 Ollama CLI 配置 GitHub Actions。🦙 Set up GitHub Actions with Ollama CLI.
并行运行 Claude Code、Codex 与 Gemini,并支持彼此交接任务。面向 AI CLI 的便携 Windows 终端。Run Claude Code, Codex and Gemini side by side — and let them hand work to each other. Portable Windows terminal for AI CLIs.
💻 在 macOS 上借助 Metal GPU 实现 Qwen3 Transformer 模型,获得加速且高效的性能,并支持关键架构特性💻 Implement Qwen3 transformer model on macOS using Metal GPU for accelerated, efficient performance with support for key architecture features.
轻松优化面向 AI 系统的 prompt,借助直观工具与特性提升性能🚀 Optimize your prompts for AI systems easily and boost performance with intuitive tools and features designed for better results.
构建 LongCat-Flash-Prover,利用 LongCat 模型进行快速定理证明与形式化推理。Build LongCat-Flash-Prover for fast theorem proving and formal reasoning with LongCat models
The Convergence Gap 的代码与 artifact:指令微调模型何时收敛到下一 token 预测Code and artifacts for The Convergence Gap: when instruction-tuned models settle on next-token predictions.
生产级 LLM 网关:兼容 OpenAI 的 API,支持语义缓存,可在 Claude/OpenAI/Ollama 之间进行智能路由与成本分析。Production LLM gateway: OpenAI-compatible API, semantic caching, intelligent routing across Claude/OpenAI/Ollama, cost analytics.
使用 TurboQuant 和 MLX 在 Windows 上运行 Qwen3.5 语言模型,实现快速的本地推理。Run the Qwen3.5 language model on Windows using TurboQuant and MLX for fast local performance.
在双 RTX 3090 GPU 上以 262K 上下文长度并发流运行 Qwen3.6-27B,开发工作已迁移至 club-3090。Run Qwen3.6-27B at 262K context with concurrent streams on dual RTX 3090 GPUs. Development moved to club-3090.
使用支持 macOS 和 Windows 平台的开源工具高效构建与定制机器学习模型Build and customize machine learning models efficiently with an open-source tool that supports macOS and Windows platforms.
通过单一严格的 YAML 清单统一管理本地 LLM 运行时——支持状态查看、健康检查、启动、停止,基于原生 systemd 与 Docker。Manage local LLM runtimes from one strict YAML manifest — status, doctor, start, stop over native systemd and Docker.
🚀 通过 autopack 简化 Hugging Face 模型的运行、分享与发布,自动完成量化与多格式导出🚀 Simplify running, sharing, and shipping Hugging Face models with autopack; it quantizes and exports to multiple formats effortlessly.
AI-Core 2026:面向 OpenAI、Anthropic、Gemini 与 Grok API 管理的集中化 WordPress AI Provider 中枢。AI-Core 2026: Centralized WordPress AI Provider Hub for OpenAI, Anthropic, Gemini & Grok API Management
Arcee AI —— 由 API Evangelist 提供的独立第三方公开 API 画像。Arcee AI 是一家美国开放智能研究实验室,构建并发布小型、高效的开源权重语言模型(Trinity 系列、AFM-4.5B 以及 Virtuoso/Maestro 衍生模型),并提供运行这些模型的开发者平台。Arcee AI — independent third-party profile of a public API surface, by API Evangelist. Arcee AI is an American open-intelligence research lab that builds and releases small, efficient open-weight language models (the Trinity family, AFM-4.5B, and Virtuoso/Maestro derivatives) along with a developer platform for running them.