研究库 开源仓库
Repositories · organized/repo_cards

仓库/Skill 库

90 个 · 多模态 · 快速增长

排序 Stars 周增
ruvnet/RuView
Rust · 2026-08-11 多模态 应用 生产可用 Stars 89443 周增 +1218

π RuView 将现成 WiFi 信号转化为实时空间智能、生命体征监测和存在检测,全程无需任何视频画面。π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video.

multimodalrisk
PaddlePaddle/PaddleOCR
Python · 2026-07-22 多模态 工具 生产可用 Stars 87395 周增 +238

将任意 PDF 或图片文档转换为结构化数据供 AI 使用。强大而轻量的 OCR 工具集,弥合图像/PDF 与 LLM 之间的鸿沟,支持 100+ 种语言。Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

ragmultimodalllm-infra
Anil-matcha/Open-Generative-AI
JavaScript · 2026-08-10 多模态 框架 生产可用 Stars 26047 周增 +280

开源无限制的 AI 视频平台替代方案 —— 免费 AI 图像与视频生成工作室,内置 200+ 模型(Flux、Midjourney、Kling、Sora、Veo)。无内容过滤,自托管,MIT 许可。Unrestricted Open-source alternative to AI video platforms — Free AI image & video generation studio with 500+ models (Flux, Midjourney, Kling, Sora, Veo). No content filters. Self-hosted, MIT licensed.

multimodal
baidu/Unlimited-OCR
Python · 2026-07-29 多模态 模型 生产可用 Stars 23399 周增 +581

Unlimited OCR Works:迈入一键长文档解析的时代。Unlimited OCR Works: Welcome the Era of One-shot Long-horizon Parsing.

img2threejs/img2threejs
Python · 2026-08-10 多模态 工具 实验 Stars 10708 周增 +1330

将参考图像中的物体重建为纯代码、程序化、质量可控、可直接用于动画的 Three.js 模型。Token 高效的图像转 3D。Rebuild the object in a reference image as a code-only, procedural, quality-gated, animation-ready Three.js model. Token-efficient image-to-3D.

agentmultimodal
XxHuberrr/Mineradio
JavaScript · 2026-07-28 多模态 应用 生产可用 Stars 9412 周增 +196

一款以电影镜头、粒子视觉和歌词舞台为核心的沉浸式音乐播放器。

storytold/photocraft
Rust · 2026-10-07 多模态 库 生产可用 Stars 8993 周增 +27836

纯 Rust 编写的 Adobe Photoshop 开源 clean-room 重新实现An open-source, clean-room reimplementation of Adobe Photoshop in pure Rust

multimodal
helloianneo/ian-xiaohei-illustrations
未知语言 · 2026-06-03 多模态 工具 生产可用 Stars 8494 周增 +119

中文小黑怪诞正文配图生成 Skill | 16:9 白底手绘 | 少量红橙蓝批注 | Codex Skill

agentmultimodal
oso95/scroll-world
JavaScript · 2026-07-29 多模态 应用 生产可用 Stars 7924 周增 +189

将任意品牌转化为可滚动 3D 世界落地页的 skillA skill that turn any brand into a scrollable 3D world landing page

teamchong/pxpipe
TypeScript · 2026-07-18 多模态 工具 研究原型 Stars 6425 周增 +140

通过将文本上下文渲染为图像,将 Fable 5 的 token 使用量降低cut Fable 5 token usage by rendering text context as images

multimodal
op7418/guizang-social-card-skill
HTML · 2026-07-01 多模态 工具 研究原型 Stars 5512 周增 +70

🪧 Claude Code / Codex Skill——生成小红书图文卡片与公众号 21:9+1:1 封面配对。编辑 × Swiss 视觉系统,28 套版式,10 种主题,单文件 HTML → PNG。小红书图文 + 公众号封面对🪧 Claude Code / Codex skill — generate Xiaohongshu carousels & WeChat 21:9+1:1 cover pairs. Editorial × Swiss visual systems, 28 layouts, 10 themes, single-file HTML → PNG. 小红书图文 + 公众号封面对

agentmultimodal
LiamGvchi/gc-minimal-zine-poster
未知语言 · 2026-08-09 多模态 工具 实验 Stars 5385 周增 +693

Codex skill,用于生成安静极简的 zine 风格编辑海报提示词与图像。Codex skill for generating quiet minimal zine-style editorial poster prompts and images.

multimodal
diffusionstudio/lottie
TypeScript · 2026-07-25 多模态 应用 研究原型 Stars 5020 周增 +28

使用 Claude Code 或 Codex 生成可投产的 Lottie 动画Generate production-ready Lottie animations with Claude Code or Codex

engineering
Vincentwei1021/video-shotcraft
TypeScript · 2026-08-09 多模态 教程 研究原型 Stars 4535 周增 +350

面向 Claude Code 和 Codex 的 AI 视频 skill —— 基于 Remotion 制作电影级产品视频:含 152 张分镜配方卡、209 个动效预览,以及一套开箱即用的模板。AI video skill for Claude Code & Codex — cinematic product videos with Remotion: 152 shot recipe cards, 209 motion previews, a production-ready template

agentmultimodalengineering
hoainho/img2threejs
Python · 2026-07-25 多模态 工具 实验 Stars 4176 周增 +2051

将参考图像中的对象重建为纯代码、程序化、带质量门控、可动画化的 Three.js 模型。token 高效的图像到三维转换。Rebuild the object in a reference image as a code-only, procedural, quality-gated, animation-ready Three.js model. Token-efficient image-to-3D.

agentmultimodal
nexu-io/html-video
HTML · 2026-06-21 多模态 教程 研究原型 Stars 4160 周增 +21

面向编码 Agent 的程序化视频方案——在本地将 HTML 转视频。把 HTML、CSS 与数据渲染为真实 MP4,支持可插拔渲染引擎、21 套模板与 AI 配乐。Apache-2.0,无按次计费。Open Design 团队的官方项目。Programmatic video for coding agents — HTML to video on your laptop. Turn HTML, CSS & data into real MP4s with pluggable render engines, 21 templates, AI soundtrack. Apache-2.0, no per-render fees. An official project by the Open Design team.

agentmultimodal
WUBING2023/PaperSpine
Python · 2026-07-01 多模态 应用 研究原型 Stars 4041 周增 +105

PaperSpine 是以动机驱动的 Skill,用于研读高质量学术论文、构建论文核心论点,并通过证据感知蓝图、修订矩阵与 LaTeX 安全审计来重写稿件。PaperSpine is a motivation-driven skill for learning from strong academic papers, building a paper’s central argument, and rewriting manuscripts through evidence-aware blueprints, revision matrices, and LaTeX-safe audits.

multimodalengineering
digimata/quill
Swift · 2026-07-30 多模态 应用 生产可用 Stars 3763 周增 +56

极简的 macOS 录音 + 转写工具。Ultra-minimalist macOS recording + transcription.

facebookresearch/vggt-omega
Python · 2026-07-02 多模态 模型 研究原型 Stars 3471 周增 +63

[CVPR 2026 Oral] VGGT Omega[CVPR 2026 Oral] VGGT Omega

MisoLabsAI/MisoTTS
Python · 2026-06-09 多模态 模型 研究原型 Stars 3131 周增 +70

Miso TTS:拥有 80 亿参数、表现力强的文本转语音模型。Miso TTS is an 8 billion, highly emotive text-to-speech model

pierrenade/short-video-generator-AI
Python · 2026-09-07 多模态 应用 研究原型 Stars 1148 周增 +0

免费开源项目,将 youtube 视频转化为爆款短视频。集成高光检测、字幕、翻译与配音,一站式内容制作。Free open-source project designed for turning youtube-viedos into viral short videos. Highlight detection, subtitles, translation, voiceover, all in one for your content.

multimodal
Vincentwei1021/anything2explainer
TypeScript · 2026-09-12 多模态 应用 研究原型 Stars 1040 周增 +0

输入主题,输出带解说的解释视频。一款 Claude Code / Codex Skill,可将任意主题转化为黑底动态图形解释视频,配备 TTS 配音、字幕和章节进度条。支持中文或英文;每一帧均通过 Remotion 以代码方式绘制。Topic in, narrated explainer video out. A Claude Code / Codex skill that turns any topic into a black-canvas motion-graphics explainer video with TTS voiceover, subtitles and a chapter progress bar. Chinese or English; every frame drawn in code with Remotion.

agentmultimodal
ysr666/dsh-vision-router
JavaScript · 2026-08-19 多模态 工具 实验 Stars 841 周增 +763

为纯文本 DeepSeek Harness Agent 装上眼睛:内置免费视觉链路(无需 key)+ 像素级视觉工具(问答、定位、裁剪、像素 diff、取色、OCR、SVG 矢量化、抠图、截图)。一条命令安装,无需 Python,图像回合与普通工具调用回合一致。Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.

agentmultimodal
Small-tailqwq/dsh-deep-whale
TypeScript · 2026-08-15 多模态 应用 研究原型 Stars 669 周增 +0

DSH Web 鲸鱼娘皮肤系列(深海女仆工坊 maid-atelier)——CC BY-NC-SA 4.0

Alain00/blobatar
TypeScript · 2026-08-20 多模态 应用 研究原型 Stars 599 周增 +0
oil-oil/oil-ui
Python · 2026-10-05 多模态 工具 实验 Stars 581 周增 +0

把 AI 的 UI 设计能力推到极限。

glanderness/BeefTV
TypeScript · 2026-09-30 多模态 应用 实验 Stars 566 周增 +1232

Local-first、轻量级、AI-native 的视频工作空间。Local-first, lightweight, AI-native video workspace.

multimodal
perseval-BLR/DLSS5-NeuralScreen
Python · 2026-09-12 多模态 工具 研究原型 Stars 516 周增 +814

DLSS 5 神经渲染,实时运行在你的整个 Windows 桌面上。支持 RTX 30/40/50,12 种语言,用户预设,带音频录制,单窗口模式。DLSS 5 Neural Rendering on your whole Windows desktop, in real time. RTX 30/40/50, 12 languages, user presets, recording with audio, one-window mode.

eternityspring/reelbench-skills
HTML · 2026-09-13 多模态 工具 研究原型 Stars 373 周增 +0

Learning notes and tooling skills for AI video - AI 视频相关的学习与工具 skill

multimodal
LBEILC/RhineLabUI
TypeScript · 2026-09-11 多模态 应用 实验 Stars 370 周增 +0

使用 TypeScript 与 Three.js 构建的 Rhine Lab 档案界面。Rhine Lab archive interface built with TypeScript and Three.js

oboroge0/hayamimi
Python · 2026-08-31 多模态 应用 实验 Stars 313 周增 +91

早耳——仅依赖 CPU 的实时多语种语音转文字。支持实时字幕、浏览器仪表盘、说话人标签与翻译,无需 GPU,无需云端。早耳 - Real-time multilingual speech-to-text on CPU only. Live subtitles, browser dashboard, speaker labels, translation. No GPU, no cloud.

JayJokerr/arknights-pixel-autofill
Python · 2026-08-12 多模态 工具 实验 Stars 279 周增 +0

明日方舟 24×24 像素画转换、手动编辑与自动填色工具

timoncool/YuE2-Studio
TypeScript · 2026-09-28 多模态 应用 实验 Stars 275 周增 +371

本地 AI 歌曲生成器,支持可编辑乐谱 —— 在你的 GPU 上运行 YuE2:带人声的全曲、五线谱、翻唱与精确回放。原生 Windows 应用,无需 Python,安装包自带自动更新。Local AI song generator with an editable score — YuE2 on your GPU: full songs with vocals, sheet music, covers, exact replay. Native Windows app, no Python, installer with auto-update.

soirihiroka/shrimply
Rust · 2026-08-31 多模态 应用 实验 Stars 264 周增 +0

你是说这条视频是一只虾做的?you're telling me a shrimp made this video?

multimodal
gatewai-dev/framefields
TypeScript · 2026-10-06 多模态 应用 实验 Stars 251 周增 +0

代码优先的视频,原生渲染于 WebGPUCode-first video, rendered natively on WebGPU.

multimodal
Merserk/dlss5-visual-enhancer
Python · 2026-09-02 多模态 应用 实验 Stars 237 周增 +0

DLSS 5 神经视频与图像增强器DLSS 5 Neural Video & Image Enhancer

multimodal