Crawl4AI:开源、对 LLM 友好的网络爬虫与抓取工具。🚀🤖 Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper. Don't be shy, join here: https://discord.gg/jP8KfhDhyN
仓库/Skill 库
96 个 · RAG 检索增强 · AI 核心
开箱即用的云模板,支持 RAG、AI 流水线与企业级搜索,接入实时数据。🐳Docker 友好。⚡始终与 Sharepoint、Google Drive、S3、Kafka、PostgreSQL 以及各类实时数据 API 保持同步。Ready-to-run cloud templates for RAG, AI pipelines, and enterprise search with live data. 🐳Docker-friendly.⚡Always in sync with Sharepoint, Google Drive, S3, Kafka, PostgreSQL, real-time data APIs, and more.
模块化的、基于图结构的 Retrieval-Augmented Generation (RAG) 系统。A modular graph-based Retrieval-Augmented Generation (RAG) system
📑 PageIndex:面向无向量、基于推理的 RAG 的文档索引📑 PageIndex: Document Index for Vectorless, Reasoning-based RAG
FastGPT 是基于 LLM 构建的知识平台,提供开箱即用的全套能力,包括数据处理、RAG 检索与可视化 AI 工作流编排,让你无需复杂配置即可轻松开发并部署复杂的问答系统。FastGPT is a knowledge-based platform built on the LLMs, offers a comprehensive suite of out-of-the-box capabilities such as data processing, RAG retrieval, and visual AI workflow orchestration, letting you easily develop and deploy complex question-answering systems without the need for extensive setup or configuration.
该仓库展示了多种面向 Retrieval-Augmented Generation (RAG) 系统的高级技术,每个技术均配有详细的 notebook 教程。This repository showcases various advanced techniques for Retrieval-Augmented Generation (RAG) systems. Each technique has a detailed notebook tutorial.
面向 AI 就绪数据的 PDF 解析器,自动化 PDF 无障碍处理,开源。PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
基于 RAG 的开源文档对话工具。An open-source RAG-based tool for chatting with your documents.
💡 面向语义搜索、LLM 编排与语言模型工作流的一体化 AI 框架💡 All-in-one AI framework for semantic search, LLM orchestration and language model workflows
LLM 实战指南:从基础到使用 LLMOps 最佳实践将 LLM 和 RAG 应用部署到 AWSThe LLM's practical guide: From the fundamentals to deploying advanced LLM and RAG apps to AWS using LLMOps best practices
免费学习如何使用 LLMOps 最佳实践构建端到端生产级 LLM & RAG 系统:源码 + 12 个实操课程。🤖 𝗟𝗲𝗮𝗿𝗻 for 𝗳𝗿𝗲𝗲 how to 𝗯𝘂𝗶𝗹𝗱 an end-to-end 𝗽𝗿𝗼𝗱𝘂𝗰𝘁𝗶𝗼𝗻-𝗿𝗲𝗮𝗱𝘆 𝗟𝗟𝗠 & 𝗥𝗔𝗚 𝘀𝘆𝘀𝘁𝗲𝗺 using 𝗟𝗟𝗠𝗢𝗽𝘀 best practices: ~ 𝘴𝘰𝘶𝘳𝘤𝘦 𝘤𝘰𝘥𝘦 + 12 𝘩𝘢𝘯𝘥𝘴-𝘰𝘯 𝘭𝘦𝘴𝘴𝘰𝘯𝘴
⚡FlashRAG:面向高效 RAG 研究的 Python 工具包(WWW2025 Resource)⚡FlashRAG: A Python Toolkit for Efficient RAG Research (WWW2025 Resource)
面向实验的 GenAI/RAG 优化器与工具包,基于 Oracle Database AI Vector Search 与 NL2SQLGenAI/RAG Optimizer and Toolkit for experimentation using Oracle Database AI Vector Search and NL2SQL
为大型语言模型(LLM)和 RAG 系统转换并优化你的 markdown 文档,自动生成 llms.txt。Transform and optimize your markdown documentation for Large Language Models (LLMs) and RAG systems. Generate llms.txt automatically.
ICLR 2026 Oral 论文 "Q-RAG: Long Context Multi-Step Retrieval via Value-Based Embedder Training" 的官方仓库。Official repository for the ICLR 2026 Oral Paper🔥 “Q-RAG: Long Context Multi-Step Retrieval via Value-Based Embedder Training”
SPY put-credit 验证与真金白银的 control plane,具备券商账本支持、确定性 risk gates、hybrid RAG、对账及每月 $1k 税后收益证据追踪。SPY put-credit validation and real-money control plane with broker-backed ledgers, deterministic risk gates, hybrid RAG, reconciliation, and $1k/mo after-tax evidence tracking.
RAG 管道的回归测试与配置扫描,附带统计学指标以判断变更是否真正带来改进。Regression testing and configuration sweeps for RAG pipelines, with the statistics to know whether a change actually helped.
SharpAI 是基于 llama.cpp(通过 LlamaSharp)构建的可嵌入 Embedding、补全与模型管理平台,内置 Ollama 兼容的 Web 服务。SharpAI is an embeddable embeddings, completions, and model management platform using llama.cpp via LlamaSharp, with a built-in Ollama-compatible webserver.
面向 Node.js 与 TypeScript 的 AI firewall。阻止 prompt injection、音频幻觉与 RAG 数据爬取。零依赖。MIT 许可。AI firewall for Node.js & TypeScript. Stop prompt injection, audio hallucinations, and RAG data scraping. Zero dependencies. MIT
工厂本体驱动的数据问答框架(开源,Apache-2.0):把中小企业台账/MES/ERP 变成可自然语言问答的语义知识图谱。CSV进答案出,确定性规则引擎+本体GraphRAG+向量混合检索,本地化数据不出厂,换行业只换词典。
动态 README,包含 AI 生成的 SCP 基金会与 Wikipedia 随机文章摘要。A dynamic README with AI-generated summaries of random articles from the SCP Foundation and Wikipedia.
邮件的"大脑":MailFathom 将 IMAP 邮箱转化为自托管、AI 原生服务。邮件同步至自有 PostgreSQL,建立索引以支持搜索与检索,并通过 Model Context Protocol 服务于 AI Agent。当前为只读;后续将支持语义检索、问答与受控写入工具。.NET 10,Apache-2.0。A brain for your mail: MailFathom turns IMAP mailboxes into a self-hosted, AI-native service. Mail synchronizes into your own PostgreSQL, is indexed for search and retrieval, and is served to AI agents over the Model Context Protocol. Read-only today; semantic retrieval, answering, and gated write tools next. .NET 10, Apache-2.0.
自托管 AI 聊天平台,支持多步工具调用、持久化 Python 沙箱、RAG 和 Deep Research。支持 Claude、GPT、Gemini 及任何 OpenAI 兼容服务商。Self-hosted AI chat platform with multi-step tool calls, persistent Python sandbox, RAG, and Deep Research. Supports Claude, GPT, Gemini and any OpenAI-compatible provider.
本地优先、可审计的技术文档与源代码知识编译器。Local-first, auditable knowledge compiler for technical docs and source code
🔬 基于 AI 与研究论文对话,通过高级语义搜索与 RAG(检索增强生成)技术提取洞见与摘要。🔬 Chat with research papers using AI, extracting insights and summaries through advanced semantic search and Retrieval-Augmented Generation techniques.
用于评估检索增强视觉语言模型在循证医学视觉问答中表现的研究框架Research framework evaluating retrieval-augmented vision-language models for evidence-grounded medical visual question answering.
支持从多源智能检索、筛选与总结科学论文,提升研究与报告生成效率Enable intelligent retrieval, filtering, and summarization of scientific papers from multiple sources for efficient research and report generation.
金融与生活交易学习 RAG 与 QuantConnect LEAN 回测工作流(FastAPI、Qdrant、pgvector、Docker)。금융·생활거래 학습 RAG와 QuantConnect LEAN 백테스트 워크플로우 (FastAPI, Qdrant, pgvector, Docker)
一条 RAG pipeline,从 Stack Overflow 抓取问答内容并转化为 RAG 格式,存储到 huggingface 数据集中。This is RAG pipeline that take question answer from Stack Overflow and convert it to rag and get stored in the dataset in huggingface
AI Research Wiki 2026:通过深度引用综合自动构建知识库AI Research Wiki 2026: Auto-Building Knowledge Base with Deep Citation Syntheses
基于 LlamaIndex、Redis 与 PII 脱敏的 RAG 搜索引擎。RAG search engine using LlamaIndex, Redis, and PII masking.