学习使用 LLM、agent、RAG、微调、LLMOps 和 AI 系统技术构建你的 Second Brain AI 助手。Learn to build your Second Brain AI assistant with LLMs, agents, RAG, fine-tuning, LLMOps and AI systems techniques.
仓库/Skill 库
1028 个 · AI 核心
几乎可视为 DSPy 在 TypeScript 上的"官方"框架。The pretty much "official" DSPy framework for Typescript
缺失的开源 Kubernetes UI,内置 MCP 服务器,面向 AI Agent。可查看故障、原因及变更。Issues、Topology、事件时间线、Helm、GitOps、实时服务流量和集群审计——全部打包在一个 Go 二进制文件中。The missing open-source Kubernetes UI with a built-in MCP server for AI agents. See what's broken, why, and what changed. Issues, Topology, event timeline, Helm, GitOps, live service traffic, and cluster audits - all in one Go binary.
开源推理服务器与生产级集群,支持 Agent 所需的全部模型Open-source inference server and production cluster for all the models your agent needs.
SOTA 低比特 LLM 量化(INT8/FP8/MXFP8/INT4/MXFP4/NVFP4)与稀疏化方案;面向 PyTorch、TensorFlow 与 ONNX Runtime 的领先模型压缩技术SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity; leading model compression techniques on PyTorch, TensorFlow, and ONNX Runtime
开源 AI 工程平台:基于 OpenTelemetry 的 LLM 可观测性、GPU 监控、guardrails、评估、prompt 管理、Vault、Playground。🚀💻 集成 50+ LLM provider、VectorDB、agent 框架与 GPU。Open source platform for AI Engineering: OpenTelemetry-native LLM Observability, GPU Monitoring, Guardrails, Evaluations, Prompt Management, Vault, Playground. 🚀💻 Integrates with 50+ LLM Providers, VectorDBs, Agent Frameworks and GPUs.
为 vLLM 在 Ascend 上的社区维护硬件插件Community maintained hardware plugin for vLLM on Ascend
Apache Hamilton 帮助数据科学家和工程师定义可测试、模块化、自文档化的数据流,内置 lineage/tracing 与 metadata,可在任何支持 Python 的环境中运行和扩展。Apache Hamilton helps data scientists and engineers define testable, modular, self-documenting dataflows, that encode lineage/tracing and metadata. Runs and scales everywhere python does.
构建可自主决策的应用(chatbot、agent、simulation 等)。支持在你的基础设施上进行监控、追踪、持久化与执行。Build applications that make decisions (chatbots, agents, simulations, etc...). Monitor, trace, persist, and execute on your own infrastructure.
基于 Rust 的 ONNX 模型快速 ML 推理与训练Fast ML inference & training for ONNX models in Rust
高效的浮点神经网络推理算子,覆盖移动端、服务端和 WebHigh-efficiency floating-point neural network inference operators for mobile, server, and Web
将任意电脑或边缘设备打造为计算机视觉项目的指挥中心Turn any computer or edge device into a command center for your computer vision projects.
cubestudio开源云原生一站式机器学习/深度学习/大模型AI平台/MaaS/mlops/人工智能平台/训推平台,算法全链路流程,多租户,算力租赁平台,token中转,拖拉拽任务流pipeline编排,多机多卡分布式训练,超参搜索,推理服务,VGPU虚拟化,云边端协同,边缘计算,自动化标注平台,deepseek等大模型sft微调/奖励模型/强化学习训练,vllm/ollama/mindie大模型多机推理,私有知识库llmops智能体,AI模型市场,支持国产异构算力调度,昇腾/寒武纪/海光/摩尔/沐曦等,支持ib/roce/RDMA,信创支持
一个简洁、高性能且可扩展的 Jax LLM!A simple, performant, and scalable Jax LLM!
Bionic 是 ChatGPT 的本地化部署替代方案,在保持严格数据机密性的同时提供生成式 AI 的能力Bionic is an on-premise replacement for ChatGPT, offering the advantages of Generative AI while maintaining strict data confidentiality
🏕️ 为开发者和 Agent 提供可复现的开发环境🏕️ Reproducible development environment for humans and agents
厂商无关的训练、推理与 Agent 负载编排,覆盖 NVIDIA、AMD、TPU 和 Tenstorrent,部署在云端、Kubernetes 和裸金属环境上Vendor-agnostic orchestration for training, inference and agentic workloads across NVIDIA, AMD, TPU, and Tenstorrent on clouds, Kubernetes, and bare metal.
通过统一接口对接 LLM 服务商Communicate with an LLM provider using a single interface
统一管理 MCP Server 的应用(MCP Manager)。A Unified MCP Server Management App (MCP Manager).
基于 Envoy Gateway 构建,提供生成式 AI 服务的统一接入管理Manages Unified Access to Generative AI Services built on Envoy Gateway
原生支持 Agent 的机票与酒店搜索与预订 —— 提供 MCP server、CLI 以及 Python/JS SDK。覆盖数百家航空公司及主要预订平台,并附带每趟航班的可靠性历史。免费取消的酒店房价:先以小额预付款锁定房间,后续可通过链接在酒店截止日期前支付余款。Agent-native flight & hotel search and booking — MCP server, CLI, and Python/JS SDKs. Hundreds of airlines plus the major booking sites, with per-flight reliability history. Free-cancellation hotel rates: hold the room with a small upfront charge, then pay the balance later by link, up to the hotel's own deadline.
潜入 Shell 的幽灵。Ante 是一个自包含的 Agent 框架,核心高度优化。体验类似 Claude Code 或 Codex,但无其依赖与模型限制。Ghost in your shell. Ante is a self-contained agent harness with a highly optimized core. It works like Claude Code or Codex, with none of their dependencies or model constraints.
WFGY 正迈向 WFGY 5.0 Polaris Protocol,面向 AI reasoning、RAG、agents 与真实工作流的重要开源版本,包含 Problem Map、Global Debug Card、WFGY 4.0 与 CFV Easter EggWFGY is heading toward WFGY 5.0 Polaris Protocol, a major open-source release for AI reasoning, RAG, agents, and real-world workflows. Includes Problem Map, Global Debug Card, WFGY 4.0, and the CFV Easter Egg.
在 Postgres 中可扩展、快速且节省磁盘的向量搜索,pgvecto.rs 的继任者。Scalable, fast, and disk-friendly vector search in Postgres, the successor of pgvecto.rs.
超快速 serverless GPU 推理、沙箱和后台任务Ultrafast serverless GPU inference, sandboxes, and background jobs
可扩展的 LLM 数据预处理与清洗工具集Scalable data pre processing and curation toolkit for LLMs
开源 LLM/VLM 负载均衡器与服务平台,用于规模化自托管 LLM(和 VLM)🏓🦙 作为 llm-d、Docker Model Runner 等项目的替代方案,组件更少、部署更简单,基于 ggml 生态构建。支持 CPU 和 GPU。Open-source LLM/VLM load balancer and serving platform for self-hosting LLMs (and VLMs) at scale 🏓🦙 Alternative to projects like llm-d, Docker Model Runner, etc but with less moving parts and simple deployments built around ggml ecosystem. Runs on CPU and GPU.
开源、端到端的 LLM 与 AI agent 应用评估、观测与改进平台。Tracing · Evals · Simulations · Datasets · Gateway · Guardrails。可自托管,Apache 2.0 许可。Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.
[GenAI 应用开发框架] 🚀 快速轻松构建 GenAI 应用 💬 在代码中使用结构化数据和链式调用语法与 GenAI Agent 交互 🧩 使用事件驱动的 *TriggerFlow* 管理复杂的 GenAI 业务逻辑 🔀 无需重写代码即可切换任意模型[GenAI Application Development Framework] 🚀 Build GenAI application quick and easy 💬 Easy to interact with GenAI agent in code using structure data and chained-calls syntax 🧩 Use Event-Driven Flow *TriggerFlow* to manage complex GenAI working logic 🔀 Switch to any model without rewrite application code
基于 nvidia-smi 二进制工具的 Nvidia GPU Prometheus exporterNvidia GPU exporter for prometheus using nvidia-smi binary
通过统一 API 接口路由、管理和分析跨多家服务商的 LLM 请求Route, manage, and analyze your LLM requests across multiple providers with a unified API interface.
面向 LLM、VLM、DiT 和 REC 模型的高性能推理引擎,针对多种 AI 加速器进行了优化。该项目托管于 OpenAtom 基金会。A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
兼容 OpenAI 和 Anthropic 协议的 Apple Silicon 服务端。可运行 LLM 和视觉语言模型(Llama、Qwen-VL、LLaVA),支持 continuous batching、MCP 工具调用与多模态。原生 MLX 后端,速度达 400+ tok/s,兼容 Claude Code。OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous batching, MCP tool calling, and multimodal support. Native MLX backend, 400+ tok/s. Works with Claude Code.