Repositories · organized/repo_cards

仓库/Skill 库

167 个 · 多模态

排序 Stars 周增
argmaxinc/argmax-oss-swift
Swift · 2026-08-06 多模态 应用 生产可用 Stars 6317 周增 +7

面向 Apple Silicon 的端侧语音 AI。On-device Speech AI for Apple Silicon

llm-infra
op7418/guizang-social-card-skill
HTML · 2026-07-01 多模态 工具 研究原型 Stars 5512 周增 +70

🪧 Claude Code / Codex Skill——生成小红书图文卡片与公众号 21:9+1:1 封面配对。编辑 × Swiss 视觉系统,28 套版式,10 种主题,单文件 HTML → PNG。小红书图文 + 公众号封面对🪧 Claude Code / Codex skill — generate Xiaohongshu carousels & WeChat 21:9+1:1 cover pairs. Editorial × Swiss visual systems, 28 layouts, 10 themes, single-file HTML → PNG. 小红书图文 + 公众号封面对

agentmultimodal
modstart-lib/aigcpanel
TypeScript · 2026-07-16 多模态 应用 生产可用 Stars 5434 周增 +14

AIGCPanel 是一个简单易用的一站式AI数字人系统,支持视频合成、声音合成、声音克隆,简化本地模型管理、一键导入和使用AI模型。

LiamGvchi/gc-minimal-zine-poster
未知语言 · 2026-08-09 多模态 工具 实验 Stars 5385 周增 +693

Codex skill,用于生成安静极简的 zine 风格编辑海报提示词与图像。Codex skill for generating quiet minimal zine-style editorial poster prompts and images.

multimodal
aigc-apps/sd-webui-EasyPhoto
Python · 2024-07-10 多模态 工具 生产可用 Stars 5153 周增 +0

📷 EasyPhoto | 你的智能 AI 照片生成器。📷 EasyPhoto | Your Smart AI Photo Generator.

diffusionstudio/lottie
TypeScript · 2026-07-25 多模态 应用 研究原型 Stars 5020 周增 +28

使用 Claude Code 或 Codex 生成可投产的 Lottie 动画Generate production-ready Lottie animations with Claude Code or Codex

engineering
Vincentwei1021/video-shotcraft
TypeScript · 2026-08-09 多模态 教程 研究原型 Stars 4535 周增 +350

面向 Claude Code 和 Codex 的 AI 视频 skill —— 基于 Remotion 制作电影级产品视频:含 152 张分镜配方卡、209 个动效预览,以及一套开箱即用的模板。AI video skill for Claude Code & Codex — cinematic product videos with Remotion: 152 shot recipe cards, 209 motion previews, a production-ready template

agentmultimodalengineering
EvolvingLMMs-Lab/lmms-eval
Python · 2026-08-06 多模态 工具 研究原型 Stars 4355 周增 +7

一个覆盖文本、图像、视频与音频任务的统一多模态评估工具包。One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks

multimodalevaluationllm-infra
hoainho/img2threejs
Python · 2026-07-25 多模态 工具 实验 Stars 4176 周增 +2051

将参考图像中的对象重建为纯代码、程序化、带质量门控、可动画化的 Three.js 模型。token 高效的图像到三维转换。Rebuild the object in a reference image as a code-only, procedural, quality-gated, animation-ready Three.js model. Token-efficient image-to-3D.

agentmultimodal
nexu-io/html-video
HTML · 2026-06-21 多模态 教程 研究原型 Stars 4160 周增 +21

面向编码 Agent 的程序化视频方案——在本地将 HTML 转视频。把 HTML、CSS 与数据渲染为真实 MP4,支持可插拔渲染引擎、21 套模板与 AI 配乐。Apache-2.0,无按次计费。Open Design 团队的官方项目。Programmatic video for coding agents — HTML to video on your laptop. Turn HTML, CSS & data into real MP4s with pluggable render engines, 21 templates, AI soundtrack. Apache-2.0, no per-render fees. An official project by the Open Design team.

agentmultimodal
WUBING2023/PaperSpine
Python · 2026-07-01 多模态 应用 研究原型 Stars 4041 周增 +105

PaperSpine 是以动机驱动的 Skill,用于研读高质量学术论文、构建论文核心论点,并通过证据感知蓝图、修订矩阵与 LaTeX 安全审计来重写稿件。PaperSpine is a motivation-driven skill for learning from strong academic papers, building a paper’s central argument, and rewriting manuscripts through evidence-aware blueprints, revision matrices, and LaTeX-safe audits.

multimodalengineering
digimata/quill
Swift · 2026-07-30 多模态 应用 生产可用 Stars 3763 周增 +56

极简的 macOS 录音 + 转写工具。Ultra-minimalist macOS recording + transcription.

PaddlePaddle/FastDeploy
Python · 2026-08-10 多模态 工具 研究原型 Stars 3703 周增 +7

基于 PaddlePaddle 的高性能 LLM 与 VLM 推理部署工具包。High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle

llm-infraengineering
dramaclaw/dramaclaw
TypeScript · 2026-08-11 多模态 应用 研究原型 Stars 3493 周增 -126

A general-purpose AIGC video engine: script to finished film in one pipeline — dramas, ads, product videos, otome games, and more. | 通用 AIGC 视频引擎 —— 从剧本到成片一条流水线,漫剧、广告、电商、乙游皆可

agentmultimodalengineering
facebookresearch/vggt-omega
Python · 2026-07-02 多模态 模型 研究原型 Stars 3471 周增 +63

[CVPR 2026 Oral] VGGT Omega[CVPR 2026 Oral] VGGT Omega

breezedeus/Pix2Text
Jupyter Notebook · 2026-02-07 多模态 工具 研究原型 Stars 3213 周增 +7

一款开源 Python3 工具,使用小模型识别图像中的版面、表格、数学公式(LaTeX)和文本,并转换为 Markdown 格式。Mathpix 的免费替代方案,可将视觉内容无缝转换为文本表示,支持 80+ 种语言。An Open-Source Python3 tool with SMALL models for recognizing layouts, tables, math formulas (LaTeX), and text in images, converting them into Markdown format. A free alternative to Mathpix, empowering seamless conversion of visual content into text-based representations. 80+ languages are supported.

multimodalengineering
MisoLabsAI/MisoTTS
Python · 2026-06-09 多模态 模型 研究原型 Stars 3131 周增 +70

Miso TTS:拥有 80 亿参数、表现力强的文本转语音模型。Miso TTS is an 8 billion, highly emotive text-to-speech model

HiLab-git/SSL4MIS
Python · 2025-06-07 多模态 收藏榜 研究原型 Stars 2671 周增 +0

半监督学习在医学图像分割中的应用,含文献综述与代码实现合集Semi Supervised Learning for Medical Image Segmentation, a collection of literature reviews and code implementations.

multimodal
LearnPrompt/LearnPrompt
MDX · 2026-08-10 多模态 应用 研究原型 Stars 2580 周增 +0

永久免费开源的 AIGC 课程, 目前已支持Claude Code,Codex,Hermes,OpenClaw,Obsidian,Prompt Engineering, ChatGPT, Midjourney, Runway, Stable Diffusion, AI数字人,AI声音&音乐,开源大模型

agentllm-infra
aigc-apps/EasyAnimate
Python · 2025-03-06 多模态 应用 实验 Stars 2266 周增 +0

📺 基于 Transformer Diffusion 的高分辨率长视频生成端到端解决方案。📺 An End-to-End Solution for High-Resolution and Long Video Generation Based on Transformer Diffusion

multimodal
aigc-apps/VideoX-Fun
Python · 2026-08-10 多模态 框架 生产可用 Stars 2192 周增 +0

📹 一个更灵活的框架,支持任意分辨率视频生成以及从图像生成视频。📹 A more flexible framework that can generate videos at any resolution and creates videos from images.

multimodal
kingyiusuen/image-to-latex
Python · 2022-10-04 多模态 工具 实验 Stars 2160 周增 +0

将 LaTeX 数学公式的图像转换为 LaTeX 代码。Convert images of LaTex math equations into LaTex code.

multimodalengineering
all-in-aigc/aicover
TypeScript · 2025-01-24 多模态 应用 生产可用 Stars 1807 周增 +0

AI 封面生成器。ai cover generator

NVIDIA-AI-Blueprints/video-search-and-summarization
C++ · 2026-08-11 多模态 应用 研究原型 Stars 1790 周增 +7

NVIDIA AI Blueprint for video search and summarization(VSS)是一个 GPU 加速参考架构,用于构建具备实时验证告警、视觉问答与自动报告能力的视频分析 Agent。VSS Blueprint 采用 NVIDIA Cosmos 等视觉语言模型(VLM)、NVIDIA Nemotron 等 LLM,并结合 RAG 与 NVIDIA NIM。NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts, visual Q&A, and automated reporting. The VSS Blueprint uses vision language models (VLMs) such as NVIDIA Cosmos, LLMs such as NVIDIA Nemotron, RAG, and NVIDIA NIMs.

agentragmultimodalllm-infra
HITsz-TMG/VideoClaw
Python · 2026-07-17 多模态 框架 研究原型 Stars 1685 周增 +21

🚀 AI 全自动化视频生成员工 | Your First AIGC Coworker. Chat an Idea. Get a Film. 🦞

agentmultimodal
intentee/paddler
Rust · 2026-07-19 多模态 模型 研究原型 Stars 1651 周增 +7

开源 LLM/VLM 负载均衡器与服务平台,用于规模化自托管 LLM(和 VLM)🏓🦙 作为 llm-d、Docker Model Runner 等项目的替代方案,组件更少、部署更简单,基于 ggml 生态构建。支持 CPU 和 GPU。Open-source LLM/VLM load balancer and serving platform for self-hosting LLMs (and VLMs) at scale 🏓🦙 Alternative to projects like llm-d, Docker Model Runner, etc but with less moving parts and simple deployments built around ggml ecosystem. Runs on CPU and GPU.

llm-infraengineering
xLLM-AI/xllm
C++ · 2026-08-12 多模态 应用 研究原型 Stars 1516 周增 +0

面向 LLM、VLM、DiT 和 REC 模型的高性能推理引擎,针对多种 AI 加速器进行了优化。该项目托管于 OpenAtom 基金会。A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.

llm-infra
huangserva/skill-prompt-generator
Python · 2026-05-10 多模态 工具 实验 Stars 1451 周增 +14

这是一个基于Claude Skill的**AI人像Prompt生成系统**,能够从特征库中智能组合生成高质量的人像描述Prompt,并具备自动学习和库扩展能力。 核心能力: Prompt生成、特征提取、自动学习、智能审核、版本控制

waybarrios/vllm-mlx
Python · 2026-06-28 多模态 工具 研究原型 Stars 1447 周增 +7

兼容 OpenAI 和 Anthropic 协议的 Apple Silicon 服务端。可运行 LLM 和视觉语言模型(Llama、Qwen-VL、LLaVA),支持 continuous batching、MCP 工具调用与多模态。原生 MLX 后端,速度达 400+ tok/s,兼容 Claude Code。OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous batching, MCP tool calling, and multimodal support. Native MLX backend, 400+ tok/s. Works with Claude Code.

multimodalllm-infra
jd-opensource/xllm
C++ · 2026-07-06 多模态 应用 研究原型 Stars 1390 周增 +28

高性能推理引擎,支持 LLM、VLM、DiT 和 REC 模型,针对多种 AI 加速器优化A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators.

llm-infra
liyue-aigc/female-portrait-director
未知语言 · 2026-07-15 多模态 工具 实验 Stars 1336 周增 +21

用于引导和扩展详细 AI 女性肖像 prompt 的模块化 Codex Skill。A modular Codex Skill for directing and expanding detailed AI female portrait prompts.

multimodal
all-in-aigc/sorafm
TypeScript · 2024-08-15 多模态 应用 实验 Stars 1152 周增 +14

Sora.FM 推出的 Sora AI 视频生成器。Sora AI Video Generator by Sora.FM

multimodal
FutureUniant/Tailor
Python · 2025-06-03 多模态 工具 生产可用 Stars 1107 周增 +0

Tailor是一款视频智能裁剪、视频生成和视频优化的视频剪辑工具。目前的目标是通过人工智能技术减少视频剪辑的繁琐操作,让普通人也能简单实现专业剪辑人的水准!长远目标是让视频剪辑实现真正的AIGC!

multimodal
Yuan-ManX/ai-audio-datasets
未知语言 · 2025-07-08 多模态 模型 实验 Stars 958 周增 +0

AI 音频数据集(AI-ADS)🎵,包含语音、音乐和音效,可为 Generative AI、AIGC、AI 模型训练、智能音频工具开发及音频应用提供训练数据。AI Audio Datasets (AI-ADS) 🎵, including Speech, Music, and Sound Effects, which can provide training data for Generative AI, AIGC, AI model training, intelligent audio tool development, and audio applications.

poleHansen/baibaiAIGC
Python · 2026-05-15 多模态 工具 研究原型 Stars 945 周增 +7

baibaiAIGCbaibaiAIGC

wshzd/Awesome-AIGC
未知语言 · 2023-10-22 多模态 收藏榜 实验 Stars 873 周增 +0

AIGC资料汇总学习,持续更新......