生产级 LLM 网关:兼容 OpenAI 的 API,支持语义缓存,可在 Claude/OpenAI/Ollama 之间进行智能路由与成本分析。Production LLM gateway: OpenAI-compatible API, semantic caching, intelligent routing across Claude/OpenAI/Ollama, cost analytics.
仓库/Skill 库
46 个 · LLM 基础设施 · 工具
使用 TurboQuant 和 MLX 在 Windows 上运行 Qwen3.5 语言模型,实现快速的本地推理。Run the Qwen3.5 language model on Windows using TurboQuant and MLX for fast local performance.
不要再重复解决同一段代码——一个本地优先的分布式推理缓存:来自真实环境的匿名兼容性证据,加上面向编码 LLM 的已验证最小示例。Stop solving the same code twice — a local-first distributed reasoning cache: anonymous compatibility evidence from real environments plus verified minimal samples for coding LLMs.
在双 RTX 3090 GPU 上以 262K 上下文长度并发流运行 Qwen3.6-27B,开发工作已迁移至 club-3090。Run Qwen3.6-27B at 262K context with concurrent streams on dual RTX 3090 GPUs. Development moved to club-3090.
本地优先、隐私至上的 AI 驱动学术写作人性化工具。AI-Powered Academic Writing Humanizer - Local-First, Privacy-Centric Solution
使用支持 macOS 和 Windows 平台的开源工具高效构建与定制机器学习模型Build and customize machine learning models efficiently with an open-source tool that supports macOS and Windows platforms.
通过单一严格的 YAML 清单统一管理本地 LLM 运行时——支持状态查看、健康检查、启动、停止,基于原生 systemd 与 Docker。Manage local LLM runtimes from one strict YAML manifest — status, doctor, start, stop over native systemd and Docker.
Codex skill,面向交通运输与低空出行领域的学术写作(知识蒸馏自 LT)。Codex skill for transportation and low-altitude mobility academic writing (Knowledge distillation from LT)
🚀 通过 autopack 简化 Hugging Face 模型的运行、分享与发布,自动完成量化与多格式导出🚀 Simplify running, sharing, and shipping Hugging Face models with autopack; it quantizes and exports to multiple formats effortlessly.
AI-Core 2026:面向 OpenAI、Anthropic、Gemini 与 Grok API 管理的集中化 WordPress AI Provider 中枢。AI-Core 2026: Centralized WordPress AI Provider Hub for OpenAI, Anthropic, Gemini & Grok API Management