Repositories · organized/repo_cards

仓库/Skill 库

39 个 · 多模态 · 工具

排序 Stars 周增
PaddlePaddle/PaddleOCR
Python · 2026-07-22 多模态 工具 生产可用 Stars 87395 周增 +238

将任意 PDF 或图片文档转换为结构化数据供 AI 使用。强大而轻量的 OCR 工具集,弥合图像/PDF 与 LLM 之间的鸿沟,支持 100+ 种语言。Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

ragmultimodalllm-infra
oobabooga/textgen
Python · 2026-06-02 多模态 工具 生产可用 Stars 47519 周增 +0

开源桌面应用,面向本地 LLM。支持文本、视觉、tool-calling,提供 OpenAI/Anthropic 兼容 API。100% 隐私保护。Open-source desktop app for local LLMs. Text, vision, tool-calling, OpenAI/Anthropic-compatible API. 100% private.

multimodalllm-infra
ocrmypdf/OCRmyPDF
Python · 2026-08-03 多模态 工具 生产可用 Stars 34348 周增 +0

OCRmyPDF 为扫描的 PDF 文件添加 OCR 文本层,使其可被搜索OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched

multimodal
HumanSignal/labelImg
Python · 2024-06-07 多模态 工具 实验 Stars 25069 周增 +0

LabelImg 现已成为 Label Studio 社区的一部分。由 Tzutalin 创建的热门图像标注工具已不再积极开发,但你可以查看 Label Studio——这款开源数据标注工具支持图像、文本、超文本、音频、视频和时序数据。LabelImg is now part of the Label Studio community. The popular image annotation tool created by Tzutalin is no longer actively being developed, but you can check out Label Studio, the open source data labeling tool for images, text, hypertext, audio, video and time-series data.

multimodal
lukas-blecher/LaTeX-OCR
Python · 2025-01-18 多模态 工具 实验 Stars 16530 周增 +7

pix2tex:使用 ViT 将公式图像转换为 LaTeX 代码pix2tex: Using a ViT to convert images of equations into LaTeX code.

multimodalengineering
img2threejs/img2threejs
Python · 2026-08-10 多模态 工具 实验 Stars 10708 周增 +1330

将参考图像中的物体重建为纯代码、程序化、质量可控、可直接用于动画的 Three.js 模型。Token 高效的图像转 3D。Rebuild the object in a reference image as a code-only, procedural, quality-gated, animation-ready Three.js model. Token-efficient image-to-3D.

agentmultimodal
helloianneo/ian-xiaohei-illustrations
未知语言 · 2026-06-03 多模态 工具 生产可用 Stars 8494 周增 +119

中文小黑怪诞正文配图生成 Skill | 16:9 白底手绘 | 少量红橙蓝批注 | Codex Skill

agentmultimodal
Moonvy/OpenPromptStudio
Vue · 2024-04-28 多模态 工具 生产可用 Stars 6648 周增 +21

🥣 AIGC 提示词可视化编辑器 | OPS | Open Prompt Studio

teamchong/pxpipe
TypeScript · 2026-07-18 多模态 工具 研究原型 Stars 6425 周增 +140

通过将文本上下文渲染为图像,将 Fable 5 的 token 使用量降低cut Fable 5 token usage by rendering text context as images

multimodal
op7418/guizang-social-card-skill
HTML · 2026-07-01 多模态 工具 研究原型 Stars 5512 周增 +70

🪧 Claude Code / Codex Skill——生成小红书图文卡片与公众号 21:9+1:1 封面配对。编辑 × Swiss 视觉系统,28 套版式,10 种主题,单文件 HTML → PNG。小红书图文 + 公众号封面对🪧 Claude Code / Codex skill — generate Xiaohongshu carousels & WeChat 21:9+1:1 cover pairs. Editorial × Swiss visual systems, 28 layouts, 10 themes, single-file HTML → PNG. 小红书图文 + 公众号封面对

agentmultimodal
LiamGvchi/gc-minimal-zine-poster
未知语言 · 2026-08-09 多模态 工具 实验 Stars 5385 周增 +693

Codex skill,用于生成安静极简的 zine 风格编辑海报提示词与图像。Codex skill for generating quiet minimal zine-style editorial poster prompts and images.

multimodal
aigc-apps/sd-webui-EasyPhoto
Python · 2024-07-10 多模态 工具 生产可用 Stars 5153 周增 +0

📷 EasyPhoto | 你的智能 AI 照片生成器。📷 EasyPhoto | Your Smart AI Photo Generator.

EvolvingLMMs-Lab/lmms-eval
Python · 2026-08-06 多模态 工具 研究原型 Stars 4355 周增 +7

一个覆盖文本、图像、视频与音频任务的统一多模态评估工具包。One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks

multimodalevaluationllm-infra
hoainho/img2threejs
Python · 2026-07-25 多模态 工具 实验 Stars 4176 周增 +2051

将参考图像中的对象重建为纯代码、程序化、带质量门控、可动画化的 Three.js 模型。token 高效的图像到三维转换。Rebuild the object in a reference image as a code-only, procedural, quality-gated, animation-ready Three.js model. Token-efficient image-to-3D.

agentmultimodal
PaddlePaddle/FastDeploy
Python · 2026-08-10 多模态 工具 研究原型 Stars 3703 周增 +7

基于 PaddlePaddle 的高性能 LLM 与 VLM 推理部署工具包。High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle

llm-infraengineering
breezedeus/Pix2Text
Jupyter Notebook · 2026-02-07 多模态 工具 研究原型 Stars 3213 周增 +7

一款开源 Python3 工具,使用小模型识别图像中的版面、表格、数学公式(LaTeX)和文本,并转换为 Markdown 格式。Mathpix 的免费替代方案,可将视觉内容无缝转换为文本表示,支持 80+ 种语言。An Open-Source Python3 tool with SMALL models for recognizing layouts, tables, math formulas (LaTeX), and text in images, converting them into Markdown format. A free alternative to Mathpix, empowering seamless conversion of visual content into text-based representations. 80+ languages are supported.

multimodalengineering
kingyiusuen/image-to-latex
Python · 2022-10-04 多模态 工具 实验 Stars 2160 周增 +0

将 LaTeX 数学公式的图像转换为 LaTeX 代码。Convert images of LaTex math equations into LaTex code.

multimodalengineering
huangserva/skill-prompt-generator
Python · 2026-05-10 多模态 工具 实验 Stars 1451 周增 +14

这是一个基于Claude Skill的**AI人像Prompt生成系统**,能够从特征库中智能组合生成高质量的人像描述Prompt,并具备自动学习和库扩展能力。 核心能力: Prompt生成、特征提取、自动学习、智能审核、版本控制

waybarrios/vllm-mlx
Python · 2026-06-28 多模态 工具 研究原型 Stars 1447 周增 +7

兼容 OpenAI 和 Anthropic 协议的 Apple Silicon 服务端。可运行 LLM 和视觉语言模型(Llama、Qwen-VL、LLaVA),支持 continuous batching、MCP 工具调用与多模态。原生 MLX 后端,速度达 400+ tok/s,兼容 Claude Code。OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous batching, MCP tool calling, and multimodal support. Native MLX backend, 400+ tok/s. Works with Claude Code.

multimodalllm-infra
liyue-aigc/female-portrait-director
未知语言 · 2026-07-15 多模态 工具 实验 Stars 1336 周增 +21

用于引导和扩展详细 AI 女性肖像 prompt 的模块化 Codex Skill。A modular Codex Skill for directing and expanding detailed AI female portrait prompts.

multimodal
FutureUniant/Tailor
Python · 2025-06-03 多模态 工具 生产可用 Stars 1107 周增 +0

Tailor是一款视频智能裁剪、视频生成和视频优化的视频剪辑工具。目前的目标是通过人工智能技术减少视频剪辑的繁琐操作,让普通人也能简单实现专业剪辑人的水准!长远目标是让视频剪辑实现真正的AIGC!

multimodal
poleHansen/baibaiAIGC
Python · 2026-05-15 多模态 工具 研究原型 Stars 945 周增 +7

baibaiAIGCbaibaiAIGC

ysr666/dsh-vision-router
JavaScript · 2026-08-19 多模态 工具 实验 Stars 841 周增 +763

为纯文本 DeepSeek Harness Agent 装上眼睛:内置免费视觉链路(无需 key)+ 像素级视觉工具(问答、定位、裁剪、像素 diff、取色、OCR、SVG 矢量化、抠图、截图)。一条命令安装,无需 Python,图像回合与普通工具调用回合一致。Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.

agentmultimodal
all-in-aigc/aiwallpaper
TypeScript · 2024-08-15 多模态 工具 生产可用 Stars 716 周增 +0

AI 壁纸生成器。AI Wallpaper Generator

aldegad/sprite-gen
Python · 2026-08-08 多模态 工具 生产可用 Stars 667 周增 +14

生成干净的 2D 游戏精灵图与动画图集——组件行流水线:状态行、alpha 清理、帧提取、运行时图集。Codex/Claude skill。Generate clean 2D game sprites & animation atlases — component-row pipeline: state rows, alpha cleanup, frame extraction, runtime atlases. Codex/Claude skill.

engineering
ItsssssJack/power-design
未知语言 · 2026-08-10 多模态 工具 生产可用 Stars 603 周增 +14

让幻灯片摆脱 AI 味的 Claude skill。品牌 DNA × 20 条编码化设计原则。A Claude skill for slides that don't look like AI made them. Brand DNA × 20 codified design principles.

rookiestar28/ComfyUI-OpenClaw
Python · 2026-08-08 多模态 工具 生产可用 Stars 557 周增 +0

你的个人 AIGC 工厂。任意图像,任意视频,以 Comfy 方式实现。©️Your own personal AIGC Factory. Any picture. Any reel. The Comfy way. ©️

agent
alonw0/web-asset-generator
Python · 2026-01-28 多模态 工具 实验 Stars 482 周增 +0

Claude skill,可基于 logo、文字或 emoji 生成 favicon、应用图标与社交媒体图片。支持 emoji 推荐、校验与框架自动集成。Claude skill to generate favicons, app icons, and social media images from logos, text, or emojis. Supports emoji suggestions, validation, and framework auto-integration.

multimodal
NitroxNova/humanizer
GDScript · 2025-11-07 多模态 工具 实验 Stars 283 周增 +0

将 MakeHuman 模型转换为 Godot4 格式convert MakeHuman to Godot4

JayJokerr/arknights-pixel-autofill
Python · 2026-08-12 多模态 工具 实验 Stars 279 周增 +0

明日方舟 24×24 像素画转换、手动编辑与自动填色工具

thiagotigaz/ocr-it
JavaScript · 2026-08-25 多模态 工具 生产可用 Stars 234 周增 +175

Chrome 扩展:固定一个屏幕区域后,可用快捷键翻阅分页文档。OCR 通过内置 Tesseract 100% 离线运行。Chrome extension: pin a screen region once, then hotkey your way through a paginated document. OCR runs 100% offline via bundled Tesseract.

huangserva/servasyy_skills
Python · 2026-02-05 多模态 工具 实验 Stars 155 周增 +0

AI驱动的多媒体内容生产skills集合:document-writer(写作)、illustration-generator(配图)、ppt-generator(PPT风格)、podcast-generator(TTS)、remoti on-dev(视频制作)、twitter-crawler(推文爬取)、markdown-illustrator(Markdown配图)、comic-generator(漫画生成)、media-downloader(媒体下载)、tts-script-generator(TTS脚本)、md-t o-pdf(文档转换)、wechat-formatter(微信格式化)、humanizer-zh(中文人性化)、shared-lib(核心API库)

Ashlixy17/PCB_lightgraph_Portable
HTML · 2026-08-23 多模态 工具 实验 Stars 81 周增 +0

PCB_lightgraph_portable 是一款点击即可运行的 PCB 图片智能分层与图纸导出的html工具。

ttthanh2044/voxdub
Python · 2026-08-12 多模态 工具 实验 Stars 77 周增 +0
jlucasmcrell/ComfyUI-H3-Multishot
Python · 2026-08-11 多模态 工具 实验 Stars 62 周增 +0
mananp-2730/BridgeBuild-AI-PM-Tool
Python · 2026-08-11 多模态 工具 实验 Stars 1 周增 +0

AI 驱动的企业级 Agile OS,使用 Gemini 1.5 将原始客户音频与笔记即时转化为客户 pitch deck、PM epic、UI 规格与后端工程架构。An AI-powered Enterprise Agile OS that instantly translates raw client audio and notes into client pitch decks, PM epics, UI specs, and backend engineering architectures using Gemini 1.5.

rag