研究库 开源仓库
Repositories · organized/repo_cards

仓库/Skill 库

33 个 · 多模态 · 工具 · 快速增长

排序 Stars 周增
PaddlePaddle/PaddleOCR
Python · 2026-07-22 多模态 工具 生产可用 Stars 87395 周增 +238

将任意 PDF 或图片文档转换为结构化数据供 AI 使用。强大而轻量的 OCR 工具集,弥合图像/PDF 与 LLM 之间的鸿沟,支持 100+ 种语言。Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

ragmultimodalllm-infra
img2threejs/img2threejs
Python · 2026-08-10 多模态 工具 实验 Stars 10708 周增 +1330

将参考图像中的物体重建为纯代码、程序化、质量可控、可直接用于动画的 Three.js 模型。Token 高效的图像转 3D。Rebuild the object in a reference image as a code-only, procedural, quality-gated, animation-ready Three.js model. Token-efficient image-to-3D.

agentmultimodal
helloianneo/ian-xiaohei-illustrations
未知语言 · 2026-06-03 多模态 工具 生产可用 Stars 8494 周增 +119

中文小黑怪诞正文配图生成 Skill | 16:9 白底手绘 | 少量红橙蓝批注 | Codex Skill

agentmultimodal
teamchong/pxpipe
TypeScript · 2026-07-18 多模态 工具 研究原型 Stars 6425 周增 +140

通过将文本上下文渲染为图像,将 Fable 5 的 token 使用量降低cut Fable 5 token usage by rendering text context as images

multimodal
op7418/guizang-social-card-skill
HTML · 2026-07-01 多模态 工具 研究原型 Stars 5512 周增 +70

🪧 Claude Code / Codex Skill——生成小红书图文卡片与公众号 21:9+1:1 封面配对。编辑 × Swiss 视觉系统,28 套版式,10 种主题,单文件 HTML → PNG。小红书图文 + 公众号封面对🪧 Claude Code / Codex skill — generate Xiaohongshu carousels & WeChat 21:9+1:1 cover pairs. Editorial × Swiss visual systems, 28 layouts, 10 themes, single-file HTML → PNG. 小红书图文 + 公众号封面对

agentmultimodal
LiamGvchi/gc-minimal-zine-poster
未知语言 · 2026-08-09 多模态 工具 实验 Stars 5385 周增 +693

Codex skill,用于生成安静极简的 zine 风格编辑海报提示词与图像。Codex skill for generating quiet minimal zine-style editorial poster prompts and images.

multimodal
hoainho/img2threejs
Python · 2026-07-25 多模态 工具 实验 Stars 4176 周增 +2051

将参考图像中的对象重建为纯代码、程序化、带质量门控、可动画化的 Three.js 模型。token 高效的图像到三维转换。Rebuild the object in a reference image as a code-only, procedural, quality-gated, animation-ready Three.js model. Token-efficient image-to-3D.

agentmultimodal
ysr666/dsh-vision-router
JavaScript · 2026-08-19 多模态 工具 实验 Stars 841 周增 +763

为纯文本 DeepSeek Harness Agent 装上眼睛:内置免费视觉链路(无需 key)+ 像素级视觉工具(问答、定位、裁剪、像素 diff、取色、OCR、SVG 矢量化、抠图、截图)。一条命令安装,无需 Python,图像回合与普通工具调用回合一致。Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.

agentmultimodal
oil-oil/oil-ui
Python · 2026-10-05 多模态 工具 实验 Stars 581 周增 +0

把 AI 的 UI 设计能力推到极限。

perseval-BLR/DLSS5-NeuralScreen
Python · 2026-09-12 多模态 工具 研究原型 Stars 516 周增 +814

DLSS 5 神经渲染,实时运行在你的整个 Windows 桌面上。支持 RTX 30/40/50,12 种语言,用户预设,带音频录制,单窗口模式。DLSS 5 Neural Rendering on your whole Windows desktop, in real time. RTX 30/40/50, 12 languages, user presets, recording with audio, one-window mode.

eternityspring/reelbench-skills
HTML · 2026-09-13 多模态 工具 研究原型 Stars 373 周增 +0

Learning notes and tooling skills for AI video - AI 视频相关的学习与工具 skill

multimodal
JayJokerr/arknights-pixel-autofill
Python · 2026-08-12 多模态 工具 实验 Stars 279 周增 +0

明日方舟 24×24 像素画转换、手动编辑与自动填色工具

thiagotigaz/ocr-it
JavaScript · 2026-08-25 多模态 工具 生产可用 Stars 234 周增 +175

Chrome 扩展:固定一个屏幕区域后,可用快捷键翻阅分页文档。OCR 通过内置 Tesseract 100% 离线运行。Chrome extension: pin a screen region once, then hotkey your way through a paginated document. OCR runs 100% offline via bundled Tesseract.

EverettFish/holo-card-studio
Python · 2026-09-07 多模态 工具 实验 Stars 231 周增 +0

将用户的描述或上传的参考素材转化为一张成品的、可编辑的 Blender 卡片以及一个交互式 Three.js 页面。保留所要求的主题、风格、字体与目标场景。本 skill 仅包含代码与文本;生成的艺术作品属于用户的输出项目。Turn the user's description or uploaded reference into a finished, editable Blender card and an interactive Three.js page. Preserve the requested subject, style, typography and destination. This skill contains code and text only; generated artwork belongs in the user's output project.

mizorewww/course2md
Rust · 2026-09-03 多模态 工具 实验 Stars 227 周增 +224

将 YouTube、Bilibili 或本地课程/会议录像转换为带幻灯片配图的 Markdown 与 HTML 讲义。Turn YouTube, Bilibili, or local course/meeting recordings into slide-illustrated Markdown and HTML lecture notes.

programasweights/claudish
Python · 2026-08-27 多模态 工具 实验 Stars 224 周增 +0

通过小型 ProgramAsWeights 函数实现英语与 Claudish 之间的相互翻译Translate between English and Claudish with tiny ProgramAsWeights functions.

lihaoyun6/ComfyUI-H3VAE_TRT
Python · 2026-09-06 多模态 工具 实验 Stars 180 周增 +0

在 ComfyUI 中运行 MiniMax-H3 VAE 的 ONNX/TRT 版本,速度最高提升至 1.7 倍。Running ONNX/TRT version of the MiniMax-H3 VAE in ComfyUI, which increase speed by up to 1.7x

AlLHHH/ALH-Pro
C# · 2026-09-05 多模态 工具 生产可用 Stars 155 周增 +0

图片视频等增强工具

GENEXIS-AI/gpt-image-skill
JavaScript · 2026-08-28 多模态 工具 生产可用 Stars 126 周增 +0

使用 ChatGPT 订阅在 Codex 或 Claude Code 中生成 GPT 图像,无需 Images API。Generate GPT images from Codex or Claude Code using a ChatGPT subscription, without the Images API.

agentmultimodal
facok/comfyui-SelfLift
Python · 2026-09-12 多模态 工具 实验 Stars 106 周增 +124
filliptm/ComfyUI-FL-YuE2
Python · 2026-09-15 多模态 工具 生产可用 Stars 86 周增 +0

YuE2 音乐生成,以及面向 ComfyUI 的可编辑钢琴卷帘界面。YuE2 music generation and an editable piano roll for ComfyUI

Ashlixy17/PCB_lightgraph_Portable
HTML · 2026-08-23 多模态 工具 实验 Stars 81 周增 +0

PCB_lightgraph_portable 是一款点击即可运行的 PCB 图片智能分层与图纸导出的html工具。

ttthanh2044/voxdub
Python · 2026-08-12 多模态 工具 实验 Stars 77 周增 +0
tig3rmast3r/OFXR-Bridge
C++ · 2026-09-07 多模态 工具 实验 Stars 73 周增 +0

基于 Optical Flow 的 VR 帧生成。Frame generation for VR using Optical Flow

ayandexyz/glance-linux
Python · 2026-09-22 多模态 工具 实验 Stars 67 周增 +0

面向 Linux 的真活体检测人脸解锁:一个 PAM 模块、一个守护进程,以及一个 Omarchy 锁屏指示器。Face unlock for Linux with real liveness detection: a PAM module, a daemon, and an Omarchy lock screen indicator

quanru/doubao-say
Python · 2026-09-20 多模态 工具 实验 Stars 64 周增 +4

由豆包驱动的语音输入工具,适用于 Linux、Wayland 和 OmarchyDoubao-powered voice input for Linux, Wayland and Omarchy

awdqwdasdg/Comfyui-Spectrum-Qwen2.1
Python · 2026-09-22 多模态 工具 实验 Stars 63 周增 +0

字面意义上的 Slop。Literal Slop.

jlucasmcrell/ComfyUI-H3-Multishot
Python · 2026-08-11 多模态 工具 实验 Stars 62 周增 +0
JGRFW/comfyui-AICG3D
JavaScript · 2026-09-17 多模态 工具 实验 Stars 60 周增 +0

把 ComfyUI-MiniMaxH3-Easy 与 Goohai-MiniMax-H3_Integration 合并成一个统一的 MiniMax H3 创作工作台(非官方)

multimodal
bhouston/three-dlss-nr
TypeScript · 2026-10-05 多模态 工具 实验 Stars 58 周增 +0

NVIDIA DLSS 5 神经渲染网络开源复现 OpenDLSS-NR 的 Three.js(TSL / WebGPU)移植版。A Three.js (TSL / WebGPU) port of OpenDLSS-NR, the open-source reimplementation of NVIDIA's DLSS 5 neural rendering network

siddzzzz/Music-Decoder
TypeScript · 2026-09-04 多模态 工具 实验 Stars 57 周增 +0
Cuongyd196/cit-voice-studio
未知语言 · 2026-09-28 多模态 工具 生产可用 Stars 53 周增 +0

CIT Voice Studio —— 将越南语文本转为语音,完全在本地运行。CIT Voice Stuido - Chuyển văn bản tiếng Việt thành giọng nói, chạy hoàn toàn trên máy bạn.

binghe1980/PromptLens
未知语言 · 2026-09-08 多模态 工具 生产可用 Stars 51 周增 +0

一款把任意图片转换为专业生图提示词的 Chrome / Edge 浏览器插件。支持网页右键识图、本地上传与剪贴板粘贴,生成中文、英文和 JSON 提示词,并提供历史记录与收藏。

multimodal