Crawl4AI:开源、对 LLM 友好的网络爬虫与抓取工具。🚀🤖 Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper. Don't be shy, join here: https://discord.gg/jP8KfhDhyN
仓库/Skill 库
119 个 · RAG 检索增强 · AI 核心
开箱即用的云模板,支持 RAG、AI 流水线与企业级搜索,接入实时数据。🐳Docker 友好。⚡始终与 Sharepoint、Google Drive、S3、Kafka、PostgreSQL 以及各类实时数据 API 保持同步。Ready-to-run cloud templates for RAG, AI pipelines, and enterprise search with live data. 🐳Docker-friendly.⚡Always in sync with Sharepoint, Google Drive, S3, Kafka, PostgreSQL, real-time data APIs, and more.
[EMNLP2025] "LightRAG:简单且快速的检索增强生成"[EMNLP2025] LightRAG: Simple and Fast Retrieval-Augmented Generation
模块化的、基于图结构的 Retrieval-Augmented Generation (RAG) 系统。A modular graph-based Retrieval-Augmented Generation (RAG) system
📑 PageIndex:面向无向量、基于推理的 RAG 的文档索引📑 PageIndex: Document Index for Vectorless, Reasoning-based RAG
FastGPT 是基于 LLM 构建的知识平台,提供开箱即用的全套能力,包括数据处理、RAG 检索与可视化 AI 工作流编排,让你无需复杂配置即可轻松开发并部署复杂的问答系统。FastGPT is a knowledge-based platform built on the LLMs, offers a comprehensive suite of out-of-the-box capabilities such as data processing, RAG retrieval, and visual AI workflow orchestration, letting you easily develop and deploy complex question-answering systems without the need for extensive setup or configuration.
该仓库展示了多种面向 Retrieval-Augmented Generation (RAG) 系统的高级技术,每个技术均配有详细的 notebook 教程。This repository showcases various advanced techniques for Retrieval-Augmented Generation (RAG) systems. Each technique has a detailed notebook tutorial.
面向 AI 就绪数据的 PDF 解析器,自动化 PDF 无障碍处理,开源。PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
基于 RAG 的开源文档对话工具。An open-source RAG-based tool for chatting with your documents.
💡 面向语义搜索、LLM 编排与语言模型工作流的一体化 AI 框架💡 All-in-one AI framework for semantic search, LLM orchestration and language model workflows
LLM 实战指南:从基础到使用 LLMOps 最佳实践将 LLM 和 RAG 应用部署到 AWSThe LLM's practical guide: From the fundamentals to deploying advanced LLM and RAG apps to AWS using LLMOps best practices
面向 monorepo 的终极 RAG。借助 AI 与知识图谱的能力,对多语言代码库进行查询、理解与编辑。The ultimate RAG for your monorepo. Query, understand, and edit multi-language codebases with the power of AI and knowledge graphs
免费学习如何使用 LLMOps 最佳实践构建端到端生产级 LLM & RAG 系统:源码 + 12 个实操课程。🤖 𝗟𝗲𝗮𝗿𝗻 for 𝗳𝗿𝗲𝗲 how to 𝗯𝘂𝗶𝗹𝗱 an end-to-end 𝗽𝗿𝗼𝗱𝘂𝗰𝘁𝗶𝗼𝗻-𝗿𝗲𝗮𝗱𝘆 𝗟𝗟𝗠 & 𝗥𝗔𝗚 𝘀𝘆𝘀𝘁𝗲𝗺 using 𝗟𝗟𝗠𝗢𝗽𝘀 best practices: ~ 𝘴𝘰𝘶𝘳𝘤𝘦 𝘤𝘰𝘥𝘦 + 12 𝘩𝘢𝘯𝘥𝘴-𝘰𝘯 𝘭𝘦𝘴𝘴𝘰𝘯𝘴
⚡FlashRAG:面向高效 RAG 研究的 Python 工具包(WWW2025 Resource)⚡FlashRAG: A Python Toolkit for Efficient RAG Research (WWW2025 Resource)
RAG 领域新 SOTA —— 一种全新的原创检索架构,以及面向人类与 Agent 的开源知识库。A new SOTA for RAG — an original retrieval architecture and an open-source knowledge base for humans and agents.
基于 OpenSearch 构建的开源、自托管企业及站内搜索服务器,爬取网页、文件、数据库与云端数据源,支持 20+ 语言、REST API,以及 AI/RAG 与语义搜索。Apache-2.0。Open-source, self-hosted enterprise & site search server built on OpenSearch. Crawls web / file / DB / cloud sources, 20+ languages, REST API, and AI/RAG & semantic search. Apache-2.0.
Rust 库,用于本地生成向量 embedding 和重排序!Rust library for generating vector embeddings and reranking locally!
使用开源 LLM 摄取文件用于检索增强生成(RAG),无需第三方,数据不出你的网络。Ingest files for retrieval augmented generation (RAG) with open-source Large Language Models (LLMs), all without 3rd parties or sensitive data leaving your network.
极简网络搜索平台,配有可在浏览器直接运行的 AI 助手。Demo:https://felladrin-minisearch.hf.spaceMinimalist web-searching platform with an AI assistant that runs directly from your browser. Demo: https://felladrin-minisearch.hf.space
デジタル庁のガバメントAI「源内(GENAI)」を完全ローカル(ローカルLLM/OpenAI互換)で動かす非公式プロジェクト。SAML認証(Keycloak)・RAG(Qdrant)・文字起こし(Whisper)・画像生成(SD)・チーム単位ナレッジをローカル完結。
本地代码搜索,结合 BM25、向量相似度与 cross-encoder 重排序。使用 tree-sitter 解析 60+ 种语言,完全离线运行,返回包含文件路径、行号范围与符号元数据的结构化结果。使用 Rust 构建。Local code search combining BM25, vector similarity, and cross-encoder reranking. Parses 60+ languages with tree-sitter, runs entirely offline, and returns structured results with file paths, line ranges, and symbol metadata. Built in Rust.
自托管 AI 聊天平台,支持多步工具调用、持久化 Python 沙箱、RAG 和 Deep Research。支持 Claude、GPT、Gemini 及任何 OpenAI 兼容服务商。Self-hosted AI chat platform with multi-step tool calls, persistent Python sandbox, RAG, and Deep Research. Supports Claude, GPT, Gemini and any OpenAI-compatible provider.
面向实验的 GenAI/RAG 优化器与工具包,基于 Oracle Database AI Vector Search 与 NL2SQLGenAI/RAG Optimizer and Toolkit for experimentation using Oracle Database AI Vector Search and NL2SQL
为大型语言模型(LLM)和 RAG 系统转换并优化你的 markdown 文档,自动生成 llms.txt。Transform and optimize your markdown documentation for Large Language Models (LLMs) and RAG systems. Generate llms.txt automatically.
ICLR 2026 Oral 论文 "Q-RAG: Long Context Multi-Step Retrieval via Value-Based Embedder Training" 的官方仓库。Official repository for the ICLR 2026 Oral Paper🔥 “Q-RAG: Long Context Multi-Step Retrieval via Value-Based Embedder Training”
Cognee Rust——快速且高性能的 AI 记忆引擎Cognee Rust - fast and performant AI memory engine
SPY put-credit 验证与真金白银的 control plane,具备券商账本支持、确定性 risk gates、hybrid RAG、对账及每月 $1k 税后收益证据追踪。SPY put-credit validation and real-money control plane with broker-backed ledgers, deterministic risk gates, hybrid RAG, reconciliation, and $1k/mo after-tax evidence tracking.
🔐 Replayable RAG —— 本地优先、可审计的知识推理平台。Graph-RAG + 仅追加审计日志 + 确定性矛盾仲裁 + 人类把关的自演化。可逐字节回放任意历史知识状态。100% 本地(Ollama)。MIT 许可。🔐 Replayable RAG — a local-first, auditable knowledge-reasoning platform. Graph-RAG + append-only audit log + deterministic contradiction arbitration + human-gated self-evolution. Replay any past knowledge state byte-for-byte. 100% local (Ollama). MIT.
邮件的大脑:MailFathom 将 IMAP 邮箱转变为自托管、AI-native 的服务。邮件同步至你自己的 PostgreSQL,建立索引以支持搜索与检索,并通过 Model Context Protocol 提供给 AI agent。具备语义检索、问答与受控的写入工具。基于 .NET 10,AGPL-3.0-only 许可。A brain for your mail: MailFathom turns IMAP mailboxes into a self-hosted, AI-native service. Mail synchronizes into your own PostgreSQL, is indexed for search and retrieval, and is served to AI agents over the Model Context Protocol. Semantic retrieval, answering, and gated write tools. .NET 10, AGPL-3.0-only.
工厂本体驱动的数据问答框架(开源,Apache-2.0):把中小企业台账/MES/ERP 变成可自然语言问答的语义知识图谱。CSV进答案出,确定性规则引擎+本体GraphRAG+向量混合检索,本地化数据不出厂,换行业只换词典。
RAG 管道的回归测试与配置扫描,附带统计学指标以判断变更是否真正带来改进。Regression testing and configuration sweeps for RAG pipelines, with the statistics to know whether a change actually helped.