Repositories · organized/repo_cards

仓库/Skill 库

35 个 · 多模态 · 快速增长

排序 Stars 周增
ruvnet/RuView
Rust · 2026-08-11 多模态 应用 生产可用 Stars 89443 周增 +1218

π RuView 将现成 WiFi 信号转化为实时空间智能、生命体征监测和存在检测,全程无需任何视频画面。π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video.

multimodalrisk
PaddlePaddle/PaddleOCR
Python · 2026-07-22 多模态 工具 生产可用 Stars 87395 周增 +238

将任意 PDF 或图片文档转换为结构化数据供 AI 使用。强大而轻量的 OCR 工具集,弥合图像/PDF 与 LLM 之间的鸿沟,支持 100+ 种语言。Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

ragmultimodalllm-infra
Anil-matcha/Open-Generative-AI
JavaScript · 2026-08-10 多模态 框架 生产可用 Stars 26047 周增 +280

开源无限制的 AI 视频平台替代方案 —— 免费 AI 图像与视频生成工作室,内置 200+ 模型(Flux、Midjourney、Kling、Sora、Veo)。无内容过滤,自托管,MIT 许可。Unrestricted Open-source alternative to AI video platforms — Free AI image & video generation studio with 500+ models (Flux, Midjourney, Kling, Sora, Veo). No content filters. Self-hosted, MIT licensed.

multimodal
baidu/Unlimited-OCR
Python · 2026-07-29 多模态 模型 生产可用 Stars 23399 周增 +581

Unlimited OCR Works:迈入一键长文档解析的时代。Unlimited OCR Works: Welcome the Era of One-shot Long-horizon Parsing.

img2threejs/img2threejs
Python · 2026-08-10 多模态 工具 实验 Stars 10708 周增 +1330

将参考图像中的物体重建为纯代码、程序化、质量可控、可直接用于动画的 Three.js 模型。Token 高效的图像转 3D。Rebuild the object in a reference image as a code-only, procedural, quality-gated, animation-ready Three.js model. Token-efficient image-to-3D.

agentmultimodal
XxHuberrr/Mineradio
JavaScript · 2026-07-28 多模态 应用 生产可用 Stars 9412 周增 +196

一款以电影镜头、粒子视觉和歌词舞台为核心的沉浸式音乐播放器。

helloianneo/ian-xiaohei-illustrations
未知语言 · 2026-06-03 多模态 工具 生产可用 Stars 8494 周增 +119

中文小黑怪诞正文配图生成 Skill | 16:9 白底手绘 | 少量红橙蓝批注 | Codex Skill

agentmultimodal
oso95/scroll-world
JavaScript · 2026-07-29 多模态 应用 生产可用 Stars 7924 周增 +189

将任意品牌转化为可滚动 3D 世界落地页的 skillA skill that turn any brand into a scrollable 3D world landing page

teamchong/pxpipe
TypeScript · 2026-07-18 多模态 工具 研究原型 Stars 6425 周增 +140

通过将文本上下文渲染为图像,将 Fable 5 的 token 使用量降低cut Fable 5 token usage by rendering text context as images

multimodal
op7418/guizang-social-card-skill
HTML · 2026-07-01 多模态 工具 研究原型 Stars 5512 周增 +70

🪧 Claude Code / Codex Skill——生成小红书图文卡片与公众号 21:9+1:1 封面配对。编辑 × Swiss 视觉系统,28 套版式,10 种主题,单文件 HTML → PNG。小红书图文 + 公众号封面对🪧 Claude Code / Codex skill — generate Xiaohongshu carousels & WeChat 21:9+1:1 cover pairs. Editorial × Swiss visual systems, 28 layouts, 10 themes, single-file HTML → PNG. 小红书图文 + 公众号封面对

agentmultimodal
LiamGvchi/gc-minimal-zine-poster
未知语言 · 2026-08-09 多模态 工具 实验 Stars 5385 周增 +693

Codex skill,用于生成安静极简的 zine 风格编辑海报提示词与图像。Codex skill for generating quiet minimal zine-style editorial poster prompts and images.

multimodal
diffusionstudio/lottie
TypeScript · 2026-07-25 多模态 应用 研究原型 Stars 5020 周增 +28

使用 Claude Code 或 Codex 生成可投产的 Lottie 动画Generate production-ready Lottie animations with Claude Code or Codex

engineering
Vincentwei1021/video-shotcraft
TypeScript · 2026-08-09 多模态 教程 研究原型 Stars 4535 周增 +350

面向 Claude Code 和 Codex 的 AI 视频 skill —— 基于 Remotion 制作电影级产品视频:含 152 张分镜配方卡、209 个动效预览,以及一套开箱即用的模板。AI video skill for Claude Code & Codex — cinematic product videos with Remotion: 152 shot recipe cards, 209 motion previews, a production-ready template

agentmultimodalengineering
hoainho/img2threejs
Python · 2026-07-25 多模态 工具 实验 Stars 4176 周增 +2051

将参考图像中的对象重建为纯代码、程序化、带质量门控、可动画化的 Three.js 模型。token 高效的图像到三维转换。Rebuild the object in a reference image as a code-only, procedural, quality-gated, animation-ready Three.js model. Token-efficient image-to-3D.

agentmultimodal
nexu-io/html-video
HTML · 2026-06-21 多模态 教程 研究原型 Stars 4160 周增 +21

面向编码 Agent 的程序化视频方案——在本地将 HTML 转视频。把 HTML、CSS 与数据渲染为真实 MP4,支持可插拔渲染引擎、21 套模板与 AI 配乐。Apache-2.0,无按次计费。Open Design 团队的官方项目。Programmatic video for coding agents — HTML to video on your laptop. Turn HTML, CSS & data into real MP4s with pluggable render engines, 21 templates, AI soundtrack. Apache-2.0, no per-render fees. An official project by the Open Design team.

agentmultimodal
WUBING2023/PaperSpine
Python · 2026-07-01 多模态 应用 研究原型 Stars 4041 周增 +105

PaperSpine 是以动机驱动的 Skill,用于研读高质量学术论文、构建论文核心论点,并通过证据感知蓝图、修订矩阵与 LaTeX 安全审计来重写稿件。PaperSpine is a motivation-driven skill for learning from strong academic papers, building a paper’s central argument, and rewriting manuscripts through evidence-aware blueprints, revision matrices, and LaTeX-safe audits.

multimodalengineering
digimata/quill
Swift · 2026-07-30 多模态 应用 生产可用 Stars 3763 周增 +56

极简的 macOS 录音 + 转写工具。Ultra-minimalist macOS recording + transcription.

facebookresearch/vggt-omega
Python · 2026-07-02 多模态 模型 研究原型 Stars 3471 周增 +63

[CVPR 2026 Oral] VGGT Omega[CVPR 2026 Oral] VGGT Omega

MisoLabsAI/MisoTTS
Python · 2026-06-09 多模态 模型 研究原型 Stars 3131 周增 +70

Miso TTS:拥有 80 亿参数、表现力强的文本转语音模型。Miso TTS is an 8 billion, highly emotive text-to-speech model

ysr666/dsh-vision-router
JavaScript · 2026-08-19 多模态 工具 实验 Stars 841 周增 +763

为纯文本 DeepSeek Harness Agent 装上眼睛:内置免费视觉链路(无需 key)+ 像素级视觉工具(问答、定位、裁剪、像素 diff、取色、OCR、SVG 矢量化、抠图、截图)。一条命令安装,无需 Python,图像回合与普通工具调用回合一致。Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.

agentmultimodal
Small-tailqwq/dsh-deep-whale
TypeScript · 2026-08-15 多模态 应用 研究原型 Stars 669 周增 +0

DSH Web 鲸鱼娘皮肤系列(深海女仆工坊 maid-atelier)——CC BY-NC-SA 4.0

Alain00/blobatar
TypeScript · 2026-08-20 多模态 应用 研究原型 Stars 599 周增 +0
JayJokerr/arknights-pixel-autofill
Python · 2026-08-12 多模态 工具 实验 Stars 279 周增 +0

明日方舟 24×24 像素画转换、手动编辑与自动填色工具

thiagotigaz/ocr-it
JavaScript · 2026-08-25 多模态 工具 生产可用 Stars 234 周增 +175

Chrome 扩展:固定一个屏幕区域后,可用快捷键翻阅分页文档。OCR 通过内置 Tesseract 100% 离线运行。Chrome extension: pin a screen region once, then hotkey your way through a paginated document. OCR runs 100% offline via bundled Tesseract.

tritant/ComfyUI_MiniMax_H3_Extender
Python · 2026-08-21 多模态 应用 实验 Stars 129 周增 +0

ComfyUI 节点,为 MiniMax H3 而设计,可串联多个视频片段,具备运动上下文、磁盘缓存、动态图像参考、音频参考支持以及最终视频/音频的无缝解码ComfyUI node for MiniMax H3 that chains multiple video clips with motion context, disk caching, dynamic image references, audio reference support, and seamless final video/audio decoding.

multimodal
kasturikhanke/generative-loaders
TypeScript · 2026-08-12 多模态 应用 实验 Stars 111 周增 +0

面向生成式接口的无障碍 React 加载状态:流式文本、内联活动指示与图像生成。Accessible React loading states for generative interfaces: streamed text, inline activity, and image generation.

multimodal
TarunTomar122/better-voice
Swift · 2026-08-24 多模态 应用 实验 Stars 103 周增 +0

语音听写,同时结合你所指屏幕位置的上下文。Voice dictation with the screen context you point at.

jaredrhod/barehands
HTML · 2026-08-17 多模态 应用 实验 Stars 103 周增 +0

用双手操控屏幕——基于摄像头的免穿戴、免手柄手势追踪界面,让你的 AI 直接响应动作Move things on your screen with your bare hands. A webcam-powered, hand-tracked interface for your AI. No headset. No controllers.

ModelTC/Minimax-H3-Turbo
Python · 2026-08-11 多模态 模型 实验 Stars 91 周增 +0

将 Minimax-H3 蒸馏为 4 步。Distill Minimax-H3 into 4 steps

Ashlixy17/PCB_lightgraph_Portable
HTML · 2026-08-23 多模态 工具 实验 Stars 81 周增 +0

PCB_lightgraph_portable 是一款点击即可运行的 PCB 图片智能分层与图纸导出的html工具。

ttthanh2044/voxdub
Python · 2026-08-12 多模态 工具 实验 Stars 77 周增 +0
A-Box-of-Tools/website
HTML · 2026-08-24 多模态 应用 实验 Stars 74 周增 +0

https://abox.tools/ 的源码——一个面向图像、视频、音频、PDF 与文本的小型 Web 工具集。文件永不离开本机,因为不存在任何能将它们外发的代码路径。Source for https://abox.tools/ — a box of small web tools for images, video, audio, PDFs and text. Your files never leave your machine, because there is no code path that could send them anywhere.

multimodal
arrival-space/splat.js
JavaScript · 2026-08-25 多模态 实验 Stars 67 周增 +112
jlucasmcrell/ComfyUI-H3-Multishot
Python · 2026-08-11 多模态 工具 实验 Stars 62 周增 +0
matlowai/ComfyUI-MAINodes
Python · 2026-08-14 多模态 应用 实验 Stars 54 周增 +0

MatlowAI 的 MiniMax-H3 ComfyUI 节点:Contact-Sheet diffusion + Motion Lab(针对快速运动的 test-time 去绳状畸变)MatlowAI's MiniMax-H3 ComfyUI nodes: Contact-Sheet diffusion + Motion Lab (test-time de-roping of fast motion)