Repositories · organized/repo_cards

仓库/Skill 库

1028 个 · AI 核心

排序 Stars 周增
decodingai-magazine/second-brain-ai-assistant-course
Jupyter Notebook · 2026-04-06 Agent 智能体 应用 研究原型 Stars 2877 周增 +21

学习使用 LLM、agent、RAG、微调、LLMOps 和 AI 系统技术构建你的 Second Brain AI 助手。Learn to build your Second Brain AI assistant with LLMs, agents, RAG, fine-tuning, LLMOps and AI systems techniques.

agentragengineeringllm-infra
ax-llm/ax
TypeScript · 2026-08-19 Agent 智能体 框架 生产可用 Stars 2876 周增 +0

几乎可视为 DSPy 在 TypeScript 上的"官方"框架。The pretty much "official" DSPy framework for Typescript

ragdatabasellm-infra
skyhook-io/radar
Go · 2026-08-11 Agent 智能体 应用 研究原型 Stars 2838 周增 +98

缺失的开源 Kubernetes UI,内置 MCP 服务器,面向 AI Agent。可查看故障、原因及变更。Issues、Topology、事件时间线、Helm、GitOps、实时服务流量和集群审计——全部打包在一个 Go 二进制文件中。The missing open-source Kubernetes UI with a built-in MCP server for AI agents. See what's broken, why, and what changed. Issues, Topology, event timeline, Helm, GitOps, live service traffic, and cluster audits - all in one Go binary.

agent
superlinked/sie
Python · 2026-08-10 Agent 智能体 应用 研究原型 Stars 2757 周增 +175

开源推理服务器与生产级集群,支持 Agent 所需的全部模型Open-source inference server and production cluster for all the models your agent needs.

agentragllm-infraengineering
intel/neural-compressor
Python · 2026-08-11 LLM 基础设施 模型 研究原型 Stars 2697 周增 +0

SOTA 低比特 LLM 量化(INT8/FP8/MXFP8/INT4/MXFP4/NVFP4)与稀疏化方案;面向 PyTorch、TensorFlow 与 ONNX Runtime 的领先模型压缩技术SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity; leading model compression techniques on PyTorch, TensorFlow, and ONNX Runtime

llm-infra
openlit/openlit
TypeScript · 2026-08-11 Agent 智能体 框架 研究原型 Stars 2680 周增 +7

开源 AI 工程平台:基于 OpenTelemetry 的 LLM 可观测性、GPU 监控、guardrails、评估、prompt 管理、Vault、Playground。🚀💻 集成 50+ LLM provider、VectorDB、agent 框架与 GPU。Open source platform for AI Engineering: OpenTelemetry-native LLM Observability, GPU Monitoring, Guardrails, Evaluations, Prompt Management, Vault, Playground. 🚀💻 Integrates with 50+ LLM Providers, VectorDBs, Agent Frameworks and GPUs.

agentevaluationdatabasellm-infra
vllm-project/vllm-ascend
C++ · 2026-08-11 LLM 基础设施 应用 研究原型 Stars 2600 周增 +49

为 vLLM 在 Ascend 上的社区维护硬件插件Community maintained hardware plugin for vLLM on Ascend

llm-infraengineering
apache/hamilton
Jupyter Notebook · 2026-08-10 工程化 框架 研究原型 Stars 2561 周增 +7

Apache Hamilton 帮助数据科学家和工程师定义可测试、模块化、自文档化的数据流,内置 lineage/tracing 与 metadata,可在任何支持 Python 的环境中运行和扩展。Apache Hamilton helps data scientists and engineers define testable, modular, self-documenting dataflows, that encode lineage/tracing and metadata. Runs and scales everywhere python does.

ragengineeringllm-infra
apache/burr
Python · 2026-08-09 Agent 智能体 应用 研究原型 Stars 2506 周增 +0

构建可自主决策的应用(chatbot、agent、simulation 等)。支持在你的基础设施上进行监控、追踪、持久化与执行。Build applications that make decisions (chatbots, agents, simulations, etc...). Monitor, trace, persist, and execute on your own infrastructure.

agentengineeringllm-infra
huggingface/huggingface.js
TypeScript · 2026-08-10 LLM 基础设施 研究原型 Stars 2487 周增 -7

在 JavaScript 中使用 Hugging FaceUse Hugging Face with JavaScript

llm-infra
pykeio/ort
Rust · 2026-08-08 LLM 基础设施 应用 研究原型 Stars 2443 周增 +14

基于 Rust 的 ONNX 模型快速 ML 推理与训练Fast ML inference & training for ONNX models in Rust

llm-infra
google/XNNPACK
C · 2026-08-10 LLM 基础设施 应用 研究原型 Stars 2421 周增 -7

高效的浮点神经网络推理算子,覆盖移动端、服务端和 WebHigh-efficiency floating-point neural network inference operators for mobile, server, and Web

llm-infra
roboflow/inference
Python · 2026-08-10 LLM 基础设施 应用 研究原型 Stars 2411 周增 +7

将任意电脑或边缘设备打造为计算机视觉项目的指挥中心Turn any computer or edge device into a command center for your computer vision projects.

agentmultimodalllm-infraengineering
data-infra/cube-studio
Python · 2026-08-04 工程化 框架 研究原型 Stars 2407 周增 +7

cubestudio开源云原生一站式机器学习/深度学习/大模型AI平台/MaaS/mlops/人工智能平台/训推平台,算法全链路流程,多租户,算力租赁平台,token中转,拖拉拽任务流pipeline编排,多机多卡分布式训练,超参搜索,推理服务,VGPU虚拟化,云边端协同,边缘计算,自动化标注平台,deepseek等大模型sft微调/奖励模型/强化学习训练,vllm/ollama/mindie大模型多机推理,私有知识库llmops智能体,AI模型市场,支持国产异构算力调度,昇腾/寒武纪/海光/摩尔/沐曦等,支持ib/roce/RDMA,信创支持

llm-infraengineering
AI-Hypercomputer/maxtext
Python · 2026-08-21 LLM 基础设施 框架 研究原型 Stars 2401 周增 +28

一个简洁、高性能且可扩展的 Jax LLM!A simple, performant, and scalable Jax LLM!

llm-infra
bionic-gpt/bionic-gpt
Rust · 2026-08-03 LLM 基础设施 应用 研究原型 Stars 2349 周增 +7

Bionic 是 ChatGPT 的本地化部署替代方案,在保持严格数据机密性的同时提供生成式 AI 的能力Bionic is an on-premise replacement for ChatGPT, offering the advantages of Generative AI while maintaining strict data confidentiality

llm-infra
tensorchord/envd
Go · 2026-07-25 Agent 智能体 应用 研究原型 Stars 2219 周增 +0

🏕️ 为开发者和 Agent 提供可复现的开发环境🏕️ Reproducible development environment for humans and agents

agentllm-infraengineering
dstackai/dstack
Python · 2026-08-11 工程化 应用 研究原型 Stars 2210 周增 +7

厂商无关的训练、推理与 Agent 负载编排,覆盖 NVIDIA、AMD、TPU 和 Tenstorrent,部署在云端、Kubernetes 和裸金属环境上Vendor-agnostic orchestration for training, inference and agentic workloads across NVIDIA, AMD, TPU, and Tenstorrent on clouds, Kubernetes, and bare metal.

agentllm-infra
mozilla-ai/any-llm
Python · 2026-08-10 LLM 基础设施 研究原型 Stars 2150 周增 +0

通过统一接口对接 LLM 服务商Communicate with an LLM provider using a single interface

llm-infra
mcp-router/mcp-router
TypeScript · 2026-08-03 Agent 智能体 应用 研究原型 Stars 2120 周增 +7

统一管理 MCP Server 的应用(MCP Manager)。A Unified MCP Server Management App (MCP Manager).

llm-infra
envoyproxy/ai-gateway
Go · 2026-08-10 LLM 基础设施 框架 研究原型 Stars 1919 周增 +14

基于 Envoy Gateway 构建,提供生成式 AI 服务的统一接入管理Manages Unified Access to Generative AI Services built on Envoy Gateway

llm-infra
LetsFG/LetsFG
Python · 2026-08-25 Agent 智能体 工具 研究原型 Stars 1863 周增 +0

原生支持 Agent 的机票与酒店搜索与预订 —— 提供 MCP server、CLI 以及 Python/JS SDK。覆盖数百家航空公司及主要预订平台,并附带每趟航班的可靠性历史。免费取消的酒店房价:先以小额预付款锁定房间,后续可通过链接在酒店截止日期前支付余款。Agent-native flight & hotel search and booking — MCP server, CLI, and Python/JS SDKs. Hundreds of airlines plus the major booking sites, with per-flight reliability history. Free-cancellation hotel rates: hold the room with a small upfront charge, then pay the balance later by link, up to the hotel's own deadline.

agent
AntigmaLabs/ante
Rust · 2026-08-20 Agent 智能体 模型 研究原型 Stars 1825 周增 +0

潜入 Shell 的幽灵。Ante 是一个自包含的 Agent 框架,核心高度优化。体验类似 Claude Code 或 Codex,但无其依赖与模型限制。Ghost in your shell. Ante is a self-contained agent harness with a highly optimized core. It works like Claude Code or Codex, with none of their dependencies or model constraints.

agent
onestardao/WFGY
Jupyter Notebook · 2026-08-15 Agent 智能体 应用 研究原型 Stars 1778 周增 +0

WFGY 正迈向 WFGY 5.0 Polaris Protocol,面向 AI reasoning、RAG、agents 与真实工作流的重要开源版本,包含 Problem Map、Global Debug Card、WFGY 4.0 与 CFV Easter EggWFGY is heading toward WFGY 5.0 Polaris Protocol, a major open-source release for AI reasoning, RAG, agents, and real-world workflows. Includes Problem Map, Global Debug Card, WFGY 4.0, and the CFV Easter Egg.

agentragevaluationrisk
supervc-stack/VectorChord
Rust · 2026-08-06 数据与向量库 研究原型 Stars 1767 周增 +0

在 Postgres 中可扩展、快速且节省磁盘的向量搜索,pgvecto.rs 的继任者。Scalable, fast, and disk-friendly vector search in Postgres, the successor of pgvecto.rs.

databasellm-infra
beam-cloud/beta9
Go · 2026-08-18 LLM 基础设施 应用 研究原型 Stars 1746 周增 +0

超快速 serverless GPU 推理、沙箱和后台任务Ultrafast serverless GPU inference, sandboxes, and background jobs

llm-infra
NVIDIA-NeMo/Curator
Python · 2026-08-19 LLM 基础设施 工具 研究原型 Stars 1723 周增 +7

可扩展的 LLM 数据预处理与清洗工具集Scalable data pre processing and curation toolkit for LLMs

engineeringllm-infra
trymirai/uzu
Rust · 2026-08-10 LLM 基础设施 应用 研究原型 Stars 1670 周增 +0

高性能 AI 模型推理引擎A high-performance inference engine for AI models

llm-infra
intentee/paddler
Rust · 2026-07-19 多模态 模型 研究原型 Stars 1651 周增 +7

开源 LLM/VLM 负载均衡器与服务平台,用于规模化自托管 LLM(和 VLM)🏓🦙 作为 llm-d、Docker Model Runner 等项目的替代方案,组件更少、部署更简单,基于 ggml 生态构建。支持 CPU 和 GPU。Open-source LLM/VLM load balancer and serving platform for self-hosting LLMs (and VLMs) at scale 🏓🦙 Alternative to projects like llm-d, Docker Model Runner, etc but with less moving parts and simple deployments built around ggml ecosystem. Runs on CPU and GPU.

llm-infraengineering
future-agi/future-agi
Python · 2026-08-11 Agent 智能体 框架 研究原型 Stars 1643 周增 +35

开源、端到端的 LLM 与 AI agent 应用评估、观测与改进平台。Tracing · Evals · Simulations · Datasets · Gateway · Guardrails。可自托管,Apache 2.0 许可。Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.

agentragllm-infraevaluation
AgentEra/Agently
Python · 2026-07-31 Agent 智能体 模型 研究原型 Stars 1636 周增 +0

[GenAI 应用开发框架] 🚀 快速轻松构建 GenAI 应用 💬 在代码中使用结构化数据和链式调用语法与 GenAI Agent 交互 🧩 使用事件驱动的 *TriggerFlow* 管理复杂的 GenAI 业务逻辑 🔀 无需重写代码即可切换任意模型[GenAI Application Development Framework] 🚀 Build GenAI application quick and easy 💬 Easy to interact with GenAI agent in code using structure data and chained-calls syntax 🧩 Use Event-Driven Flow *TriggerFlow* to manage complex GenAI working logic 🔀 Switch to any model without rewrite application code

agentllm-infra
langwatch/better-agents
TypeScript · 2026-06-03 Agent 智能体 应用 研究原型 Stars 1553 周增 +0

构建 Agent 的更高标准Standards for building agents, better

agentllm-infra
utkuozdemir/nvidia_gpu_exporter
Go · 2026-08-07 LLM 基础设施 工具 生产可用 Stars 1527 周增 +14

基于 nvidia-smi 二进制工具的 Nvidia GPU Prometheus exporterNvidia GPU exporter for prometheus using nvidia-smi binary

llm-infra
theopenco/llmgateway
TypeScript · 2026-08-11 LLM 基础设施 框架 研究原型 Stars 1520 周增 +7

通过统一 API 接口路由、管理和分析跨多家服务商的 LLM 请求Route, manage, and analyze your LLM requests across multiple providers with a unified API interface.

llm-infra
xLLM-AI/xllm
C++ · 2026-08-12 多模态 应用 研究原型 Stars 1516 周增 +0

面向 LLM、VLM、DiT 和 REC 模型的高性能推理引擎,针对多种 AI 加速器进行了优化。该项目托管于 OpenAtom 基金会。A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.

llm-infra
waybarrios/vllm-mlx
Python · 2026-06-28 多模态 工具 研究原型 Stars 1447 周增 +7

兼容 OpenAI 和 Anthropic 协议的 Apple Silicon 服务端。可运行 LLM 和视觉语言模型(Llama、Qwen-VL、LLaVA),支持 continuous batching、MCP 工具调用与多模态。原生 MLX 后端,速度达 400+ tok/s,兼容 Claude Code。OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous batching, MCP tool calling, and multimodal support. Native MLX backend, 400+ tok/s. Works with Claude Code.

multimodalllm-infra