友好的 AI 交互界面(支持 Ollama、OpenAI API 等)。User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
仓库/Skill 库
62 个 · LLM 基础设施 · 应用
面向 LLM 的高吞吐、内存高效的推理与 serving 引擎。A high-throughput and memory-efficient inference and serving engine for LLMs
为GPT/GLM等LLM大语言模型提供实用化交互接口,特别优化论文阅读/润色/写作体验,模块化设计,支持自定义快捷按钮&函数插件,支持Python和C++等项目剖析&自译解功能,PDF/LaTex论文翻译&总结功能,支持并行问询多种LLM模型,支持chatglm3等本地模型。接入通义千问, deepseekcoder, 讯飞星火, 文心一言, llama2, rwkv, claude2, moss等。
开源 AI 平台——支持所有 LLM 的 AI Chat,并提供丰富的高级功能。Open Source AI Platform - AI Chat with advanced features that works with every LLM
通过 API 访问的免费 LLM 推理资源列表。A list of free LLM inference resources accessible via API.
本地化、开源的 AI 应用构建器,面向高级用户 ✨ v0 / Lovable / Replit / Bolt 的替代方案 🌟 喜欢就点 Star!Local, open-source AI app builder for power users ✨ v0 / Lovable / Replit / Bolt alternative 🌟 Star if you like it!
面向 Metal、CUDA 和 ROCm 的 DeepSeek 4 Flash 与 PRO 本地推理引擎。DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
先记录,后整理。一款 local-first 的 Markdown 应用,借助 AI 将零散记录转化为清晰的笔记。Capture first. Organize later. A local-first Markdown app that turns scattered records into clear notes with AI.
LMCache:借助最快的 KV Cache 层为你的 LLM 加速。LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Triton Inference Server 提供针对云端和边缘推理优化的解决方案。The Triton Inference Server provides an optimized cloud and edge inferencing solution.
基于 PyTorch 的 YOLOv3、YOLOv3-SPP 和 YOLOv3-tiny 实时目标检测实现,支持训练、验证、推理与多格式导出。PyTorch implementation of YOLOv3, YOLOv3-SPP, and YOLOv3-tiny for real-time object detection with training, validation, inference, and multi-format export.
🚀💪Maximize your efficiency and productivity. The ultimate hub to manage, customize, and share prompts. (English/中文/Español/العربية). 让生产力加倍的 AI 快捷指令。更高效地管理提示词,在分享社区中发现适用于不同场景的灵感。
Gemma 4 26B-A4B 在任意 M 系列 MacBook 上以约 2 GB 内存进行推理。Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
在任意消费级硬件上运行 Qwen3.8-Flash-Next:Windows / Linux 一键安装。Strata 推理引擎,本地 localhost 提供 OpenAI / Anthropic API,支持可选图像输入。Qwen3.8-Flash-Next on any consumer hardware: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input.
PyTorch 原生的量化和稀疏化方案,支持训练与推理PyTorch native quantization and sparsity for training and inference
为 vLLM 在 Ascend 上的社区维护硬件插件Community maintained hardware plugin for vLLM on Ascend
基于 Rust 的 ONNX 模型快速 ML 推理与训练Fast ML inference & training for ONNX models in Rust
高效的浮点神经网络推理算子,覆盖移动端、服务端和 WebHigh-efficiency floating-point neural network inference operators for mobile, server, and Web
将任意电脑或边缘设备打造为计算机视觉项目的指挥中心Turn any computer or edge device into a command center for your computer vision projects.
Bionic 是 ChatGPT 的本地化部署替代方案,在保持严格数据机密性的同时提供生成式 AI 的能力Bionic is an on-premise replacement for ChatGPT, offering the advantages of Generative AI while maintaining strict data confidentiality
超快速 serverless GPU 推理、沙箱和后台任务Ultrafast serverless GPU inference, sandboxes, and background jobs
RTP-LLM:阿里巴巴面向多样化应用的高性能 LLM 推理引擎。RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.
基于 Rust 🦀 构建的高性能、模块化、内存安全、生产可用的推理、数据接入与索引系统Highly Performant, Modular, Memory Safe and Production-ready Inference, Ingestion and Indexing built in Rust 🦀
StarWhisper 天文 LLMs、StarWhisper Telescope、Virtual-GOTTA,以及面向 embodied observing workflow 的天文定制研究 skills。StarWhisper astronomy LLMs, StarWhisper Telescope, Virtual-GOTTA, and astronomy-adapted research skills for embodied observing workflows
逐步学习 LLM 推理工程——从 KV cache、PagedAttention 和连续批处理,到 vLLM、SGLang 和 GPU。Learn LLM Inference Engineering step by step - from KV cache, PagedAttention, and continuous batching to vLLM, SGLang, and GPUs.
一门关于现代 LLM 架构、训练与推理的实战课程,配有逐步 PyTorch 实现和可运行的 notebook。A hands-on course on modern LLM architectures, training, and inference, with step-by-step PyTorch implementations and runnable notebooks.