基于 Rust 🦀 构建的高性能、模块化、内存安全、生产可用的推理、数据接入与索引系统Highly Performant, Modular, Memory Safe and Production-ready Inference, Ingestion and Indexing built in Rust 🦀
仓库/Skill 库
135 个 · LLM 基础设施 · AI 核心
轻量、可移植的 LLM 沙箱运行时(代码解释器)Python 库。Lightweight and portable LLM sandbox runtime (code interpreter) Python library.
面向 LLM 和 LVLM 的强大模型压缩框架,适配 NVIDIA GPU 和华为昇腾 NPUA powerful model compression framework for LLMs and LVLMs, adapted for NVIDIA GPUs and Huawei Ascend NPUs.
Krasis 是一个混合 LLM 运行时,专注于在消费级 VRAM 受限硬件上高效运行大模型Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware
关于大语言模型 On-Policy Distillation 的精选论文与资源合集A curated collection of papers and resources on On-Policy Distillation for Large Language Models.
StarWhisper 天文 LLMs、StarWhisper Telescope、Virtual-GOTTA,以及面向 embodied observing workflow 的天文定制研究 skills。StarWhisper astronomy LLMs, StarWhisper Telescope, Virtual-GOTTA, and astronomy-adapted research skills for embodied observing workflows
🪢 Langfuse 文档——Langfuse 是开源 LLM 工程平台,提供可观测性、评估、Prompt 管理、Playground 与指标,用于调试与改进 LLM 应用🪢 Langfuse documentation -- Langfuse is the open source LLM Engineering Platform. Observability, evals, prompt management, playground and metrics to debug and improve LLM apps
LLM7.io 提供单一 API 网关,可连接来自多家供应商的众多领先 AI 模型LLM7.io offers a single API gateway that connects you to a wide array of leading AI models from various providers.
一门实战课程,用 PyTorch 从零构建现代 LLM,包含 26 个可运行的 Jupyter Notebook,涵盖 tokenizer、attention、MoE、RLHF、推理、评估和蒸馏。A hands-on course for building modern LLMs from scratch in PyTorch, with 26 runnable Jupyter Notebooks covering tokenizers, attention, MoE, RLHF, inference, evaluation, and distillation.
2026 年度顶级开源 AI 工程平台 Qyvaria KernelTop Open-Source AI Engineering Platform 2026 Qyvaria Kernel
一款便捷的 lib,用于流畅地与大语言模型(LLM)交互及构建 AI 应用。A handy lib for smooth interaction with large language models (LLMs) and crafting AI apps.
🪢 Langfuse API 的自动生成 Java 客户端。🪢 Auto-generated Java Client for Langfuse API
ELM 是将 LLM 应用于能源研究的一系列工具集合。ELM is a collection of utilities to apply Large Language Models (LLMs) to energy research.
repo-map 生成由 LLM 增强的软件仓库摘要与分析,为开发者提供关于项目结构、文件用途以及跨编程语言潜在考量的洞察。repo-map generates LLM-enhanced summaries and analysis of software repositories, providing developers with valuable insights into project structures, file purposes, and potential considerations across various programming languages.
对人类而言,语言是表达的工具;对 AI 而言,语言是推理的基底。For humans, a language is a tool for expression. For AIs, it's a substrate for reasoning.
AI 原生的结构化数据 wire 格式。在每个前沿模型上实现 100% 理解,比 JSON 减少 50-92% token,跨 17 种格式完成 43B+ 无损往返。Spec v3.4 Stable。The AI-native wire format for structured data. 100% comprehension on every frontier model. 50-92% fewer tokens than JSON. 43B+ lossless round-trips across 17 formats. Spec v3.4 Stable.
精选 AI 模型及其 API 提供商列表,完全无需信用卡即可使用,欢迎贡献!Curated list of AI Models with their API Providers that you never ever require a Credit Card for. Feel free to Contribute!
一个以证据为引领的六语种 LLM 实战手册:包含可迁移的核心、Codex 旗舰路线,以及 ChatGPT、Claude Code、Gemini、DeepSeek 和 Grok 的适配器。An evidence-led, six-language LLM playbook: the transferable core, the Codex flagship track, and adapters for ChatGPT, Claude Code, Gemini, DeepSeek, and Grok.
🦙 使用 Ollama CLI 配置 GitHub Actions。🦙 Set up GitHub Actions with Ollama CLI.
本地 AI 秘书、技术支持与销售一体化方案,基于 XTTS v2 语音克隆、Vosk/Whisper 实时语音识别与 vLLM + Qwen/Llama 等离线 LLM。配备 Vue 3 完整管理面板、Telegram Bot、网站挂件及 fine-tuning pipeline。支持自托管、数据隐私、短信与电话呼叫。📞 Локальный AI-секретарь, тех. поддержка и менеджер по продажам с клонированием голоса XTTS v2, real-time распознаванием речи (Vosk/Whisper) и offline LLM (vLLM + Qwen/Llama и тп). Полноценная админ-панель (Vue 3), Telegram-бот, виджет для сайта, fine-tuning pipeline. Self-hosted, приватность данных, СМС и телефонные звонки .
运行 Ollama 本地 LLM 服务的 Docker 镜像。默认安全,所有 API 请求需 Bearer token(首次启动时自动生成)。OpenAI 兼容 API。支持首次启动模型预拉取、NVIDIA GPU (CUDA) 加速和持久化模型存储。多架构:amd64、arm64。Docker image to run an Ollama local LLM server. Secure by default, all API requests require a Bearer token (auto-generated on first start). OpenAI-compatible API. Supports first-start model pre-pull, NVIDIA GPU (CUDA) acceleration, and persistent model storage. Multi-arch: amd64, arm64.
💻 在 macOS 上借助 Metal GPU 实现 Qwen3 Transformer 模型,获得加速且高效的性能,并支持关键架构特性💻 Implement Qwen3 transformer model on macOS using Metal GPU for accelerated, efficient performance with support for key architecture features.
Gebo.ai —— 开源、企业级、与 AI 供应商无关的平台Gebo.ai The open source Enterprise AI vendor agnostic platform
以孟加拉语优先、面向方言的 LLM 研究。开源孟加拉语 tokenizer,性能优于 Sarvam、AI4Bharat 和 GPT-4o(fertility 1.52,几乎零破损 conjuncts)。非商用,保护孟加拉语及其方言。Bengali-first, dialect-aware LLM research. Open-source Bengali tokenizer that outperforms Sarvam, AI4Bharat, and GPT-4o (fertility 1.52, near-zero broken conjuncts). Non-commercial, preserving Bengali and its dialects.
🚀 通过 70 个生成式 AI 实战项目转型技能,从入门到生产级架构师,涵盖真实场景应用。🚀 Transform your skills with 70 hands-on projects in Generative AI, guiding you from beginner to production-ready architect with real-world applications.
本仓库包含用于执行不同任务的 Gen AI 项目temThis repository contains Gen AI projects that performs different tasks
并行运行 Claude Code、Codex 与 Gemini,并支持彼此交接任务。面向 AI CLI 的便携 Windows 终端。Run Claude Code, Codex and Gemini side by side — and let them hand work to each other. Portable Windows terminal for AI CLIs.
大语言模型元数据的开放注册表——以单一机器可读的 models.json 提供身份、作者、模态、上下文/输出限制、能力与生命周期日期,并通过 JSON Schema 校验。CC BY 4.0Open registry of large-language-model metadata — identity, authorship, modalities, context/output limits, capabilities & lifecycle dates as one machine-readable models.json validated by JSON Schema. CC BY 4.0.
🛠️ 通过 Local-LLM(一款基于 BERT 模型的轻量级 Python 库)在安全环境中实现离线 NLP 工作流,确保可复现性与可靠性。🛠️ Enable offline NLP workflows with Local-LLM, a lightweight Python library for secure environments using the BERT model, ensuring reproducibility and reliability.
Run the native 284B-A13B DeepSeek-V4-Flash-0731 LLM locally on a single laptop CPU: pure C, 8 GB RAM minimum, no GPU, best TPOT 0.892 s/token. | 在笔记本单颗 CPU 上本地运行原生 284B-A13B DeepSeek-V4-Flash-0731 大模型:纯 C,最低 8 GB 内存,无需 GPU,最优 TPOT 0.892 秒/token。
轻松优化面向 AI 系统的 prompt,借助直观工具与特性提升性能🚀 Optimize your prompts for AI systems easily and boost performance with intuitive tools and features designed for better results.
《大模型推理原理与优化》:面向系统/架构/后端研发工程师的模型原理入门课,目标是通俗易懂的解释推理过程,理解原理有助于系统开发/维护工作
DP-ES 官方实现:面向 Prompt 优化的差分隐私进化策略(EMNLP 2026)。Official implementation of DP-ES: Differentially Private Evolution Strategies for Prompt Optimization (EMNLP 2026)
通过一次 wrap() 调用为 LLM 流水线构建运行时可靠性守卫,可配置地防御常见生产故障Build runtime reliability guards for LLM pipelines with one wrap() call and configurable protection against common production failures