研究库 开源仓库
Repositories · organized/repo_cards

仓库/Skill 库

236 个 · 多模态

排序 Stars 周增
teamchong/pxpipe
TypeScript · 2026-07-18 多模态 工具 研究原型 Stars 6425 周增 +140

通过将文本上下文渲染为图像,将 Fable 5 的 token 使用量降低cut Fable 5 token usage by rendering text context as images

multimodal
argmaxinc/argmax-oss-swift
Swift · 2026-08-06 多模态 应用 生产可用 Stars 6317 周增 +7

面向 Apple Silicon 的端侧语音 AI。On-device Speech AI for Apple Silicon

llm-infra
op7418/guizang-social-card-skill
HTML · 2026-07-01 多模态 工具 研究原型 Stars 5512 周增 +70

🪧 Claude Code / Codex Skill——生成小红书图文卡片与公众号 21:9+1:1 封面配对。编辑 × Swiss 视觉系统,28 套版式,10 种主题,单文件 HTML → PNG。小红书图文 + 公众号封面对🪧 Claude Code / Codex skill — generate Xiaohongshu carousels & WeChat 21:9+1:1 cover pairs. Editorial × Swiss visual systems, 28 layouts, 10 themes, single-file HTML → PNG. 小红书图文 + 公众号封面对

agentmultimodal
modstart-lib/aigcpanel
TypeScript · 2026-07-16 多模态 应用 生产可用 Stars 5434 周增 +14

AIGCPanel 是一个简单易用的一站式AI数字人系统,支持视频合成、声音合成、声音克隆,简化本地模型管理、一键导入和使用AI模型。

LiamGvchi/gc-minimal-zine-poster
未知语言 · 2026-08-09 多模态 工具 实验 Stars 5385 周增 +693

Codex skill,用于生成安静极简的 zine 风格编辑海报提示词与图像。Codex skill for generating quiet minimal zine-style editorial poster prompts and images.

multimodal
aigc-apps/sd-webui-EasyPhoto
Python · 2024-07-10 多模态 工具 生产可用 Stars 5153 周增 +0

📷 EasyPhoto | 你的智能 AI 照片生成器。📷 EasyPhoto | Your Smart AI Photo Generator.

diffusionstudio/lottie
TypeScript · 2026-07-25 多模态 应用 研究原型 Stars 5020 周增 +28

使用 Claude Code 或 Codex 生成可投产的 Lottie 动画Generate production-ready Lottie animations with Claude Code or Codex

engineering
Vincentwei1021/video-shotcraft
TypeScript · 2026-08-09 多模态 教程 研究原型 Stars 4535 周增 +350

面向 Claude Code 和 Codex 的 AI 视频 skill —— 基于 Remotion 制作电影级产品视频:含 152 张分镜配方卡、209 个动效预览,以及一套开箱即用的模板。AI video skill for Claude Code & Codex — cinematic product videos with Remotion: 152 shot recipe cards, 209 motion previews, a production-ready template

agentmultimodalengineering
EvolvingLMMs-Lab/lmms-eval
Python · 2026-08-29 多模态 工具 研究原型 Stars 4382 周增 +10

一个覆盖文本、图像、视频与音频任务的统一多模态评估工具包。One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks

multimodalevaluationllm-infra
hoainho/img2threejs
Python · 2026-07-25 多模态 工具 实验 Stars 4176 周增 +2051

将参考图像中的对象重建为纯代码、程序化、带质量门控、可动画化的 Three.js 模型。token 高效的图像到三维转换。Rebuild the object in a reference image as a code-only, procedural, quality-gated, animation-ready Three.js model. Token-efficient image-to-3D.

agentmultimodal
nexu-io/html-video
HTML · 2026-06-21 多模态 教程 研究原型 Stars 4160 周增 +21

面向编码 Agent 的程序化视频方案——在本地将 HTML 转视频。把 HTML、CSS 与数据渲染为真实 MP4,支持可插拔渲染引擎、21 套模板与 AI 配乐。Apache-2.0,无按次计费。Open Design 团队的官方项目。Programmatic video for coding agents — HTML to video on your laptop. Turn HTML, CSS & data into real MP4s with pluggable render engines, 21 templates, AI soundtrack. Apache-2.0, no per-render fees. An official project by the Open Design team.

agentmultimodal
WUBING2023/PaperSpine
Python · 2026-07-01 多模态 应用 研究原型 Stars 4041 周增 +105

PaperSpine 是以动机驱动的 Skill,用于研读高质量学术论文、构建论文核心论点,并通过证据感知蓝图、修订矩阵与 LaTeX 安全审计来重写稿件。PaperSpine is a motivation-driven skill for learning from strong academic papers, building a paper’s central argument, and rewriting manuscripts through evidence-aware blueprints, revision matrices, and LaTeX-safe audits.

multimodalengineering
digimata/quill
Swift · 2026-07-30 多模态 应用 生产可用 Stars 3763 周增 +56

极简的 macOS 录音 + 转写工具。Ultra-minimalist macOS recording + transcription.

PaddlePaddle/FastDeploy
Python · 2026-08-10 多模态 工具 研究原型 Stars 3703 周增 +7

基于 PaddlePaddle 的高性能 LLM 与 VLM 推理部署工具包。High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle

llm-infraengineering
dramaclaw/dramaclaw
TypeScript · 2026-08-11 多模态 应用 研究原型 Stars 3493 周增 -126

A general-purpose AIGC video engine: script to finished film in one pipeline — dramas, ads, product videos, otome games, and more. | 通用 AIGC 视频引擎 —— 从剧本到成片一条流水线,漫剧、广告、电商、乙游皆可

agentmultimodalengineering
facebookresearch/vggt-omega
Python · 2026-07-02 多模态 模型 研究原型 Stars 3471 周增 +63

[CVPR 2026 Oral] VGGT Omega[CVPR 2026 Oral] VGGT Omega

breezedeus/Pix2Text
Jupyter Notebook · 2026-02-07 多模态 工具 研究原型 Stars 3213 周增 +7

一款开源 Python3 工具,使用小模型识别图像中的版面、表格、数学公式(LaTeX)和文本,并转换为 Markdown 格式。Mathpix 的免费替代方案,可将视觉内容无缝转换为文本表示,支持 80+ 种语言。An Open-Source Python3 tool with SMALL models for recognizing layouts, tables, math formulas (LaTeX), and text in images, converting them into Markdown format. A free alternative to Mathpix, empowering seamless conversion of visual content into text-based representations. 80+ languages are supported.

multimodalengineering
MisoLabsAI/MisoTTS
Python · 2026-06-09 多模态 模型 研究原型 Stars 3131 周增 +70

Miso TTS:拥有 80 亿参数、表现力强的文本转语音模型。Miso TTS is an 8 billion, highly emotive text-to-speech model

HiLab-git/SSL4MIS
Python · 2025-06-07 多模态 收藏榜 研究原型 Stars 2671 周增 +0

半监督学习在医学图像分割中的应用,含文献综述与代码实现合集Semi Supervised Learning for Medical Image Segmentation, a collection of literature reviews and code implementations.

multimodal
LearnPrompt/LearnPrompt
MDX · 2026-08-10 多模态 应用 研究原型 Stars 2580 周增 +0

永久免费开源的 AIGC 课程, 目前已支持Claude Code,Codex,Hermes,OpenClaw,Obsidian,Prompt Engineering, ChatGPT, Midjourney, Runway, Stable Diffusion, AI数字人,AI声音&音乐,开源大模型

agentllm-infra
aigc-apps/EasyAnimate
Python · 2025-03-06 多模态 应用 实验 Stars 2266 周增 +0

📺 基于 Transformer Diffusion 的高分辨率长视频生成端到端解决方案。📺 An End-to-End Solution for High-Resolution and Long Video Generation Based on Transformer Diffusion

multimodal
aigc-apps/VideoX-Fun
Python · 2026-08-10 多模态 框架 生产可用 Stars 2192 周增 +0

📹 一个更灵活的框架,支持任意分辨率视频生成以及从图像生成视频。📹 A more flexible framework that can generate videos at any resolution and creates videos from images.

multimodal
kingyiusuen/image-to-latex
Python · 2022-10-04 多模态 工具 实验 Stars 2160 周增 +0

将 LaTeX 数学公式的图像转换为 LaTeX 代码。Convert images of LaTex math equations into LaTex code.

multimodalengineering
NVIDIA-AI-Blueprints/video-search-and-summarization
Python · 2026-10-09 多模态 应用 研究原型 Stars 1918 周增 +26

NVIDIA AI Blueprint for video search and summarization(VSS)是一个 GPU 加速参考架构,用于构建具备实时验证告警、视觉问答与自动报告能力的视频分析 Agent。VSS Blueprint 采用 NVIDIA Cosmos 等视觉语言模型(VLM)、NVIDIA Nemotron 等 LLM,并结合 RAG 与 NVIDIA NIM。NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts, visual Q&A, and automated reporting. The VSS Blueprint uses vision language models (VLMs) such as NVIDIA Cosmos, LLMs such as NVIDIA Nemotron, RAG, and NVIDIA NIMs.

agentragmultimodalllm-infra
szczyglis-dev/py-gpt
Python · 2026-09-03 多模态 应用 研究原型 Stars 1904 周增 +0

由 GPT-5、GPT-4、o1、o3、Gemini、Claude、Ollama、DeepSeek、Perplexity、Grok、Bielik 驱动的桌面 AI Assistant,支持 chat、vision、voice、RAG、图像与视频生成、agents、tools、MCP、plugins、语音合成与识别、网页搜索、记忆、预设、assistants 等。Linux、Windows、Mac。Desktop AI Assistant powered by GPT-5, GPT-4, o1, o3, Gemini, Claude, Ollama, DeepSeek, Perplexity, Grok, Bielik, chat, vision, voice, RAG, image and video generation, agents, tools, MCP, plugins, speech synthesis and recognition, web search, memory, presets, assistants,and more. Linux, Windows, Mac

agentragmultimodalllm-infra
all-in-aigc/aicover
TypeScript · 2025-01-24 多模态 应用 生产可用 Stars 1807 周增 +0

AI 封面生成器。ai cover generator

HITsz-TMG/VideoClaw
Python · 2026-07-17 多模态 框架 研究原型 Stars 1685 周增 +21

🚀 AI 全自动化视频生成员工 | Your First AIGC Coworker. Chat an Idea. Get a Film. 🦞

agentmultimodal
bytedance/Sa2VA
Python · 2026-09-08 多模态 框架 研究原型 Stars 1666 周增 +0

Pixel-LLM 代码库官方仓库:Sa2VA(T-PAMI-26)、SAMTok(CVPR-26)、VRT(Arxiv-25)、SaSaSa2VA(LSVOS 第一名方案)。Official Repo For Pixel-LLM Codebase: Sa2VA (T-PAMI-26), SAMTok (CVPR-26), VRT (Arxiv-25), SaSaSa2VA (1-st solution for LSVOS)

multimodalllm-infra
intentee/paddler
Rust · 2026-07-19 多模态 模型 研究原型 Stars 1651 周增 +7

开源 LLM/VLM 负载均衡器与服务平台,用于规模化自托管 LLM(和 VLM)🏓🦙 作为 llm-d、Docker Model Runner 等项目的替代方案,组件更少、部署更简单,基于 ggml 生态构建。支持 CPU 和 GPU。Open-source LLM/VLM load balancer and serving platform for self-hosting LLMs (and VLMs) at scale 🏓🦙 Alternative to projects like llm-d, Docker Model Runner, etc but with less moving parts and simple deployments built around ggml ecosystem. Runs on CPU and GPU.

llm-infraengineering
xLLM-AI/xllm
C++ · 2026-10-04 多模态 应用 研究原型 Stars 1588 周增 +3

面向 LLM、VLM、DiT 和 REC 模型的高性能推理引擎,针对多种 AI 加速器进行了优化。该项目托管于 OpenAtom 基金会。A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.

llm-infra
huangserva/skill-prompt-generator
Python · 2026-05-10 多模态 工具 实验 Stars 1451 周增 +14

这是一个基于Claude Skill的**AI人像Prompt生成系统**,能够从特征库中智能组合生成高质量的人像描述Prompt,并具备自动学习和库扩展能力。 核心能力: Prompt生成、特征提取、自动学习、智能审核、版本控制

waybarrios/vllm-mlx
Python · 2026-06-28 多模态 工具 研究原型 Stars 1447 周增 +7

兼容 OpenAI 和 Anthropic 协议的 Apple Silicon 服务端。可运行 LLM 和视觉语言模型(Llama、Qwen-VL、LLaVA),支持 continuous batching、MCP 工具调用与多模态。原生 MLX 后端,速度达 400+ tok/s,兼容 Claude Code。OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous batching, MCP tool calling, and multimodal support. Native MLX backend, 400+ tok/s. Works with Claude Code.

multimodalllm-infra
jd-opensource/xllm
C++ · 2026-07-06 多模态 应用 研究原型 Stars 1390 周增 +28

高性能推理引擎,支持 LLM、VLM、DiT 和 REC 模型,针对多种 AI 加速器优化A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators.

llm-infra
liyue-aigc/female-portrait-director
未知语言 · 2026-07-15 多模态 工具 实验 Stars 1336 周增 +21

用于引导和扩展详细 AI 女性肖像 prompt 的模块化 Codex Skill。A modular Codex Skill for directing and expanding detailed AI female portrait prompts.

multimodal
all-in-aigc/sorafm
TypeScript · 2024-08-15 多模态 应用 实验 Stars 1152 周增 +14

Sora.FM 推出的 Sora AI 视频生成器。Sora AI Video Generator by Sora.FM

multimodal
pierrenade/short-video-generator-AI
Python · 2026-09-07 多模态 应用 研究原型 Stars 1148 周增 +0

免费开源项目,将 youtube 视频转化为爆款短视频。集成高光检测、字幕、翻译与配音,一站式内容制作。Free open-source project designed for turning youtube-viedos into viral short videos. Highlight detection, subtitles, translation, voiceover, all in one for your content.

multimodal