Repositories · organized/repo_cards

仓库/Skill 库

167 个 · 多模态

排序 Stars 周增
ysr666/dsh-vision-router
JavaScript · 2026-08-19 多模态 工具 实验 Stars 841 周增 +763

为纯文本 DeepSeek Harness Agent 装上眼睛:内置免费视觉链路(无需 key)+ 像素级视觉工具(问答、定位、裁剪、像素 diff、取色、OCR、SVG 矢量化、抠图、截图)。一条命令安装,无需 Python,图像回合与普通工具调用回合一致。Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.

agentmultimodal
coderonion/awesome-llm-and-aigc
未知语言 · 2025-08-01 多模态 收藏榜 实验 Stars 811 周增 +0

🚀🚀🚀 收录关于 LLM、VLM、VLA、AIGC 及相关数据集与应用的一些优秀开源项目合集。🚀🚀🚀A collection of some awesome public projects about Large Language Model(LLM), Vision Language Model(VLM), Vision Language Action(VLA), AI Generated Content(AIGC), the related Datasets and Applications.

multimodalllm-infra
hitcslj/Awesome-AIGC-3D
Python · 2026-05-04 多模态 收藏榜 研究原型 Stars 787 周增 +0

精选的 AIGC 3D 优秀论文列表。A curated list of awesome AIGC 3D papers

TypeTale/TypeTale
未知语言 · 2026-01-06 多模态 应用 生产可用 Stars 776 周增 +14

字字动画 - 完全免费的AIGC视频生成软件,主要用于AI短剧,AI电影,小说推文

all-in-aigc/aiwallpaper
TypeScript · 2024-08-15 多模态 工具 生产可用 Stars 716 周增 +0

AI 壁纸生成器。AI Wallpaper Generator

Kobaayyy/Awesome-CVPR2026-CVPR2025-ICCV2025-CVPR2024-ECCV2026-ECCV2024-AIGC
未知语言 · 2026-08-05 多模态 收藏榜 研究原型 Stars 674 周增 +0

CVPR2026/CVPR2025/ICCV2025/CVPR2024/ECCV2026/ECCV2024 AIGC 方向的论文与代码合集。A Collection of Papers and Codes for CVPR2026/CVPR2025/ICCV2025/CVPR2024/ECCV2026/ECCV2024 AIGC

multimodalllm-infra
Small-tailqwq/dsh-deep-whale
TypeScript · 2026-08-15 多模态 应用 研究原型 Stars 669 周增 +0

DSH Web 鲸鱼娘皮肤系列(深海女仆工坊 maid-atelier)——CC BY-NC-SA 4.0

aldegad/sprite-gen
Python · 2026-08-08 多模态 工具 生产可用 Stars 667 周增 +14

生成干净的 2D 游戏精灵图与动画图集——组件行流水线:状态行、alpha 清理、帧提取、运行时图集。Codex/Claude skill。Generate clean 2D game sprites & animation atlases — component-row pipeline: state rows, alpha cleanup, frame extraction, runtime atlases. Codex/Claude skill.

engineering
ItsssssJack/power-design
未知语言 · 2026-08-10 多模态 工具 生产可用 Stars 603 周增 +14

让幻灯片摆脱 AI 味的 Claude skill。品牌 DNA × 20 条编码化设计原则。A Claude skill for slides that don't look like AI made them. Brand DNA × 20 codified design principles.

Alain00/blobatar
TypeScript · 2026-08-20 多模态 应用 研究原型 Stars 599 周增 +0
gongminmin/awesome-aigc
未知语言 · 2023-10-21 多模态 收藏榜 实验 Stars 569 周增 +0

AIGC 优秀作品列表。A list of awesome AIGC works

rookiestar28/ComfyUI-OpenClaw
Python · 2026-08-08 多模态 工具 生产可用 Stars 557 周增 +0

你的个人 AIGC 工厂。任意图像,任意视频,以 Comfy 方式实现。©️Your own personal AIGC Factory. Any picture. Any reel. The Comfy way. ©️

agent
all-in-aigc/melodisco
TypeScript · 2024-09-16 多模态 应用 生产可用 Stars 514 周增 +0

AI Music PlayerAI Music Player

Moonlit-Pages/AIGC-Detector-Rewriter-Skill
未知语言 · 2026-06-19 多模态 应用 研究原型 Stars 488 周增 +0

一个面向英文学术写作的保守型 AIGC 检测器指导的论文改写 Skill。支持 Turnitin AI、CNKI AIGC、最小化编辑修订、保留学术要素、定性/定量路由,以及逐章降低 AI 写作风险,且不宣称绕过检测器。A conservative AIGC detector-informed thesis rewriting skill for English and Chinese academic writing. Supports Turnitin AI, CNKI AIGC, minimal-edit revision, protected academic elements, qualitative/quantitative routing, and chapter-by-chapter AI-writing risk reduction without detector-bypass claims.

multimodalriskengineering
alonw0/web-asset-generator
Python · 2026-01-28 多模态 工具 实验 Stars 482 周增 +0

Claude skill,可基于 logo、文字或 emoji 生成 favicon、应用图标与社交媒体图片。支持 emoji 推荐、校验与框架自动集成。Claude skill to generate favicons, app icons, and social media images from logos, text, or emojis. Supports emoji suggestions, validation, and framework auto-integration.

multimodal
modelscope/richdreamer
Python · 2024-09-27 多模态 模型 实验 Stars 478 周增 +0

[CVPR2024 (Highlight)] RichDreamer:一种可泛化的法线-深度扩散模型,用于生成细节丰富的文本到 3D 内容。Live Demo:https://modelscope.cn/studios/Damo_XR_Lab/3D_AIGC[CVPR2024 (Highlight)] RichDreamer: A Generalizable Normal-Depth Diffusion Model for Detail Richness in Text-to-3D. Live Demo:https://modelscope.cn/studios/Damo_XR_Lab/3D_AIGC

lycohana/BiliSum
Python · 2026-08-15 多模态 研究原型 Stars 471 周增 +0

为 Bilibili、YouTube 及本地视频提供 AI 视频摘要和知识库.AI video summarizer and knowledge base for Bilibili, YouTube and local videos.

ragmultimodal
fajarhide/omni
Rust · 2026-08-11 多模态 数据集 研究原型 Stars 320 周增 +0

决定你的 AI 编码 Agent 阅读哪些内容,并保留凭证。每次裁剪都可逐字节恢复,从不杜撰结果,所有数据从你自己的语料库回放。Decides what your AI coding agent reads, and keeps receipts. Every cut is recoverable byte for byte, it never invents a result, and the numbers are replayed from your own corpus.

agentllm-infra
NitroxNova/humanizer
GDScript · 2025-11-07 多模态 工具 实验 Stars 283 周增 +0

将 MakeHuman 模型转换为 Godot4 格式convert MakeHuman to Godot4

JayJokerr/arknights-pixel-autofill
Python · 2026-08-12 多模态 工具 实验 Stars 279 周增 +0

明日方舟 24×24 像素画转换、手动编辑与自动填色工具

modu-ai/cowork-plugins
HTML · 2026-06-19 多模态 应用 实验 Stars 267 周增 +7

所有人可用的 AI(MoAI)——面向韩语实务领域的 Claude Cowork 与 Claude Code AI harness 与插件市场。覆盖商业计划书、税务、法律、HR、营销、电商、BI、内容等领域,提供 Skill、Agent 与工作流。支持韩语 B2B 场景与办公文档(HWPX/DOCX/XLSX/PPTX/PDF)及 AI 多模态生成(图像/视频/语音)。内置 AI 痕迹审核与韩语 humanize-korean모두의 AI (MoAI) — Claude Cowork & Claude Code 한국 실무 도메인 AI 하네스·플러그인 마켓플레이스. 사업계획서·세무·법률·HR·마케팅·커머스·BI·콘텐츠 도메인 스킬·에이전트·워크플로우. Korean B2B + office docs (HWPX/DOCX/XLSX/PPTX/PDF) + AI media (image/video/voice). AI-slop 검수 + humanize-korean 내장.

agentmultimodal
thiagotigaz/ocr-it
JavaScript · 2026-08-25 多模态 工具 生产可用 Stars 234 周增 +175

Chrome 扩展:固定一个屏幕区域后,可用快捷键翻阅分页文档。OCR 通过内置 Tesseract 100% 离线运行。Chrome extension: pin a screen region once, then hotkey your way through a paginated document. OCR runs 100% offline via bundled Tesseract.

VisionVerse/RemoteSensing-Restoration-Survey
Python · 2026-08-09 多模态 应用 实验 Stars 231 周增 +0

[ISPRS 2026] 遥感图像去雾:进展、挑战与前景的系统综述[ISPRS 2026] Remote Sensing Image Dehazing: A Systematic Review of Progress, Challenges, and Prospects

multimodal
hubertjb/dl-eeg-review
Python · 2020-02-12 多模态 收藏榜 研究原型 Stars 159 周增 +0

深度学习与 EEG 系统文献综述的补充材料Supplementary material for systematic literature review on deep learning and EEG.

huangserva/servasyy_skills
Python · 2026-02-05 多模态 工具 实验 Stars 155 周增 +0

AI驱动的多媒体内容生产skills集合:document-writer(写作)、illustration-generator(配图)、ppt-generator(PPT风格)、podcast-generator(TTS)、remoti on-dev(视频制作)、twitter-crawler(推文爬取)、markdown-illustrator(Markdown配图)、comic-generator(漫画生成)、media-downloader(媒体下载)、tts-script-generator(TTS脚本)、md-t o-pdf(文档转换)、wechat-formatter(微信格式化)、humanizer-zh(中文人性化)、shared-lib(核心API库)

tritant/ComfyUI_MiniMax_H3_Extender
Python · 2026-08-21 多模态 应用 实验 Stars 129 周增 +0

ComfyUI 节点,为 MiniMax H3 而设计,可串联多个视频片段,具备运动上下文、磁盘缓存、动态图像参考、音频参考支持以及最终视频/音频的无缝解码ComfyUI node for MiniMax H3 that chains multiple video clips with motion context, disk caching, dynamic image references, audio reference support, and seamless final video/audio decoding.

multimodal
kasturikhanke/generative-loaders
TypeScript · 2026-08-12 多模态 应用 实验 Stars 111 周增 +0

面向生成式接口的无障碍 React 加载状态:流式文本、内联活动指示与图像生成。Accessible React loading states for generative interfaces: streamed text, inline activity, and image generation.

multimodal
TarunTomar122/better-voice
Swift · 2026-08-24 多模态 应用 实验 Stars 103 周增 +0

语音听写,同时结合你所指屏幕位置的上下文。Voice dictation with the screen context you point at.

jaredrhod/barehands
HTML · 2026-08-17 多模态 应用 实验 Stars 103 周增 +0

用双手操控屏幕——基于摄像头的免穿戴、免手柄手势追踪界面,让你的 AI 直接响应动作Move things on your screen with your bare hands. A webcam-powered, hand-tracked interface for your AI. No headset. No controllers.

synlp/RRG-Review
TeX · 2025-05-17 多模态 收藏榜 研究原型 Stars 100 周增 +0

综述论文"A Systematic Review of Deep Learning-based Research on Radiology Report Generation"的官方 GitHub 仓库The official GitHub repository of the survey paper "A Systematic Review of Deep Learning-based Research on Radiology Report Generation".

ModelTC/Minimax-H3-Turbo
Python · 2026-08-11 多模态 模型 实验 Stars 91 周增 +0

将 Minimax-H3 蒸馏为 4 步。Distill Minimax-H3 into 4 steps

Ashlixy17/PCB_lightgraph_Portable
HTML · 2026-08-23 多模态 工具 实验 Stars 81 周增 +0

PCB_lightgraph_portable 是一款点击即可运行的 PCB 图片智能分层与图纸导出的html工具。

stabgan/openrouter-mcp-multimodal
TypeScript · 2026-08-22 多模态 应用 实验 Stars 78 周增 +0

OpenRouter 的 MCP server——与 300+ LLM(Claude、Gemini、GPT)对话,分析图像/音频/视频,生成图像/语音/音乐/视频(Veo 3.1、Sora、Seedance、Wan),支持 Claude Desktop、Cursor、Kiro、VS CodeMCP server for OpenRouter — chat with 300+ LLMs (Claude, Gemini, GPT), analyze images / audio / video, generate images / speech / music / video (Veo 3.1, Sora, Seedance, Wan) from Claude Desktop, Cursor, Kiro, VS Code.

agentmultimodalllm-infra
ttthanh2044/voxdub
Python · 2026-08-12 多模态 工具 实验 Stars 77 周增 +0
A-Box-of-Tools/website
HTML · 2026-08-24 多模态 应用 实验 Stars 74 周增 +0

https://abox.tools/ 的源码——一个面向图像、视频、音频、PDF 与文本的小型 Web 工具集。文件永不离开本机,因为不存在任何能将它们外发的代码路径。Source for https://abox.tools/ — a box of small web tools for images, video, audio, PDFs and text. Your files never leave your machine, because there is no code path that could send them anywhere.

multimodal
arrival-space/splat.js
JavaScript · 2026-08-25 多模态 实验 Stars 67 周增 +112