PaddlePaddle/PaddleOCR

  • 类型:github
  • 标识:PaddlePaddle/PaddleOCR
  • 链接:https://github.com/PaddlePaddle/PaddleOCR
  • 主题:rag, multimodal, llm-infra
  • 主分类:multimodal
  • 形态:tool
  • 分类:trending
  • Stars:87395
  • 周增:+238
  • 语言:Python
  • 许可:Apache-2.0
  • 最近提交:2026-07-22
  • 简介:Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
  • 上次采集:2026-08-11
  • 简介中文:将任意 PDF 或图片文档转换为结构化数据供 AI 使用。强大而轻量的 OCR 工具集,弥合图像/PDF 与 LLM 之间的鸿沟,支持 100+ 种语言。
  • 待LLM分类:否
  • 成熟度:production
  • 场景:ocr、documents
  • 首次采集:2026-06-30
  • 来源文件
  • [GitHub Search]