Repositories · organized/repo_cards

仓库/Skill 库

167 个 · 多模态

排序 Stars 周增
huggingface/transformers
Python · 2026-08-03 多模态 模型 生产可用 Stars 163296 周增 +0

Transformers:面向文本、视觉、音频及多模态 SOTA 机器学习模型的模型定义框架,同时支持推理与训练。🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.

multimodalllm-infra
ruvnet/RuView
Rust · 2026-08-11 多模态 应用 生产可用 Stars 89443 周增 +1218

π RuView 将现成 WiFi 信号转化为实时空间智能、生命体征监测和存在检测,全程无需任何视频画面。π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video.

multimodalrisk
PaddlePaddle/PaddleOCR
Python · 2026-07-22 多模态 工具 生产可用 Stars 87395 周增 +238

将任意 PDF 或图片文档转换为结构化数据供 AI 使用。强大而轻量的 OCR 工具集,弥合图像/PDF 与 LLM 之间的鸿沟,支持 100+ 种语言。Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

ragmultimodalllm-infra
CompVis/stable-diffusion
Jupyter Notebook · 2024-06-18 多模态 模型 实验 Stars 73272 周增 +0

潜空间文本到图像扩散模型。A latent text-to-image diffusion model

multimodal
ggml-org/whisper.cpp
C++ · 2026-08-07 多模态 生产可用 Stars 52799 周增 +70

OpenAI Whisper 模型的 C/C++ 移植版本。Port of OpenAI's Whisper model in C/C++

llm-infra
mudler/LocalAI
Go · 2026-08-11 多模态 模型 生产可用 Stars 48383 周增 +84

LocalAI 是开源 AI 引擎,可在任何硬件上运行任意模型 —— LLMs、视觉、语音、图像、视频,无需 GPU。LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.

agentmultimodalllm-infra
oobabooga/textgen
Python · 2026-06-02 多模态 工具 生产可用 Stars 47519 周增 +0

开源桌面应用,面向本地 LLM。支持文本、视觉、tool-calling,提供 OpenAI/Anthropic 兼容 API。100% 隐私保护。Open-source desktop app for local LLMs. Text, vision, tool-calling, OpenAI/Anthropic-compatible API. 100% private.

multimodalllm-infra
coqui-ai/TTS
Python · 2024-08-16 多模态 应用 实验 Stars 45852 周增 +0

🐸💬 —— 一个历经研究与生产环境考验的 Text-to-Speech 深度学习工具包。🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production

engineering
suno-ai/bark
Jupyter Notebook · 2024-08-19 多模态 模型 实验 Stars 39217 周增 +0

🔊 文本提示生成式音频模型🔊 Text-Prompted Generative Audio Model

google-ai-edge/mediapipe
C++ · 2026-08-10 多模态 框架 生产可用 Stars 36566 周增 +49

面向实时和流媒体、跨平台可定制的 ML 解决方案。Cross-platform, customizable ML solutions for live and streaming media.

multimodalllm-infraengineering
ocrmypdf/OCRmyPDF
Python · 2026-08-03 多模态 工具 生产可用 Stars 34348 周增 +0

OCRmyPDF 为扫描的 PDF 文件添加 OCR 文本层,使其可被搜索OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched

multimodal
JaidedAI/EasyOCR
Python · 2025-12-05 多模态 生产可用 Stars 29861 周增 +0

开箱即用的 OCR,支持 80+ 种语言及所有主流书写系统,包括拉丁文、中文、阿拉伯文、天城文、西里尔文等。Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.

ragmultimodal
cjpais/Handy
Rust · 2026-08-03 多模态 应用 生产可用 Stars 28596 周增 +0

一款免费、开源且可扩展的语音转文字应用程序,完全离线运行。A free, open source, and extensible speech-to-text application that works completely offline.

mozilla/DeepSpeech
C++ · 2025-06-19 多模态 应用 实验 Stars 26770 周增 +0

DeepSpeech 是一个开源嵌入式(离线、端侧)语音转文字引擎,可在从 Raspberry Pi 4 到高性能 GPU 服务器的设备上实时运行。DeepSpeech is an open source embedded (offline, on-device) speech-to-text engine which can run in real time on devices ranging from a Raspberry Pi 4 to high power GPU servers.

Anil-matcha/Open-Generative-AI
JavaScript · 2026-08-10 多模态 框架 生产可用 Stars 26047 周增 +280

开源无限制的 AI 视频平台替代方案 —— 免费 AI 图像与视频生成工作室,内置 200+ 模型(Flux、Midjourney、Kling、Sora、Veo)。无内容过滤,自托管,MIT 许可。Unrestricted Open-source alternative to AI video platforms — Free AI image & video generation studio with 500+ models (Flux, Midjourney, Kling, Sora, Veo). No content filters. Self-hosted, MIT licensed.

multimodal
HumanSignal/labelImg
Python · 2024-06-07 多模态 工具 实验 Stars 25069 周增 +0

LabelImg 现已成为 Label Studio 社区的一部分。由 Tzutalin 创建的热门图像标注工具已不再积极开发,但你可以查看 Label Studio——这款开源数据标注工具支持图像、文本、超文本、音频、视频和时序数据。LabelImg is now part of the Label Studio community. The popular image annotation tool created by Tzutalin is no longer actively being developed, but you can check out Label Studio, the open source data labeling tool for images, text, hypertext, audio, video and time-series data.

multimodal
pytorch/examples
Python · 2025-09-01 多模态 教程 实验 Stars 23986 周增 +0

围绕 PyTorch 在视觉、文本、强化学习等领域的一组示例。A set of examples around pytorch in Vision, Text, Reinforcement Learning, etc.

multimodal
baidu/Unlimited-OCR
Python · 2026-07-29 多模态 模型 生产可用 Stars 23399 周增 +581

Unlimited OCR Works:迈入一键长文档解析的时代。Unlimited OCR Works: Welcome the Era of One-shot Long-horizon Parsing.

wandb/openui
TypeScript · 2026-08-05 多模态 应用 生产可用 Stars 22494 周增 +7

OpenUI 可让你凭想象力描述 UI,并实时看到渲染效果。OpenUI let's you describe UI using your imagination, then see it rendered live.

index-tts/index-tts
Python · 2026-07-14 多模态 应用 生产可用 Stars 22384 周增 +0

工业级可控、高效的零样本文本转语音系统An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

NVIDIA-NeMo/Speech
Python · 2026-08-11 多模态 框架 生产可用 Stars 18087 周增 +56

为研究人员和开发者构建的可扩展生成式 AI 框架,面向 LLM、多模态和语音 AI(自动语音识别与文字转语音)。A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)

multimodalllm-infra
BradyFU/Awesome-Multimodal-Large-Language-Models
未知语言 · 2026-08-14 多模态 收藏榜 生产可用 Stars 17977 周增 +14

✨✨多模态大语言模型最新进展。:sparkles::sparkles:Latest Advances on Multimodal Large Language Models

multimodalllm-infra
lukas-blecher/LaTeX-OCR
Python · 2025-01-18 多模态 工具 实验 Stars 16530 周增 +7

pix2tex:使用 ViT 将公式图像转换为 LaTeX 代码pix2tex: Using a ViT to convert images of equations into LaTeX code.

multimodalengineering
HBAI-Ltd/Toonflow-app
TypeScript · 2026-07-28 多模态 框架 生产可用 Stars 13701 周增 +126

Toonflow 是开源一站式 AI 短剧创作工具,将小说、剧本快速转化为动画短剧。集成 AI 编剧、智能分镜、角色与视频生成,跨平台桌面端轻量部署,助力创作者低成本批量产出视觉内容。Toonflow is an open-source AI tool that turns stories and scripts into animated short dramas. Features AI scriptwriting, storyboarding, character and video generation. A cross-platform desktop app for efficient content creation.

multimodalllm-infra
img2threejs/img2threejs
Python · 2026-08-10 多模态 工具 实验 Stars 10708 周增 +1330

将参考图像中的物体重建为纯代码、程序化、质量可控、可直接用于动画的 Three.js 模型。Token 高效的图像转 3D。Rebuild the object in a reference image as a code-only, procedural, quality-gated, animation-ready Three.js model. Token-efficient image-to-3D.

agentmultimodal
Acly/krita-ai-diffusion
Python · 2026-06-30 多模态 应用 生产可用 Stars 10449 周增 +21

在 Krita 中使用 AI 生成图像的精简界面。支持 Inpaint 与 Outpaint,可选文本提示,无需调参。Streamlined interface for generating images with AI in Krita. Inpaint and outpaint with optional text prompt, no tweaking required.

multimodal
AIGC-Audio/AudioGPT
Python · 2024-07-06 多模态 应用 实验 Stars 10171 周增 +0

AudioGPT:理解与生成语音、音乐、声音与说话人头像AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head

xorbitsai/inference
Python · 2026-08-11 多模态 应用 生产可用 Stars 9486 周增 +0

通过修改一行代码即可将 GPT 替换为任意 LLM。Xinference 让你在云端、本地或笔记本上运行开源、语音和多模态模型,全部通过统一的生产就绪推理 API。Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.

multimodalllm-infraengineering
XxHuberrr/Mineradio
JavaScript · 2026-07-28 多模态 应用 生产可用 Stars 9412 周增 +196

一款以电影镜头、粒子视觉和歌词舞台为核心的沉浸式音乐播放器。

oumi-ai/oumi
Python · 2026-08-11 多模态 应用 生产可用 Stars 9369 周增 +7

轻松微调、评估和部署 Gemma 4、Qwen3.5、Qwen3.6、gpt-oss、DeepSeek-R1 或任意开源 LLM / VLM。Easily fine-tune, evaluate and deploy Qwen, Gemma, or any open weight LLM!

llm-infraevaluation
Lightricks/LTX-2
Python · 2026-08-03 多模态 模型 生产可用 Stars 8560 周增 +14

LTX-2 音视频生成模型的官方 Python 推理与 LoRA 训练包。Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.

multimodalllm-infra
helloianneo/ian-xiaohei-illustrations
未知语言 · 2026-06-03 多模态 工具 生产可用 Stars 8494 周增 +119

中文小黑怪诞正文配图生成 Skill | 16:9 白底手绘 | 少量红橙蓝批注 | Codex Skill

agentmultimodal
oso95/scroll-world
JavaScript · 2026-07-29 多模态 应用 生产可用 Stars 7924 周增 +189

将任意品牌转化为可滚动 3D 世界落地页的 skillA skill that turn any brand into a scrollable 3D world landing page

open-mmlab/mmagic
Jupyter Notebook · 2024-08-06 多模态 模型 实验 Stars 7452 周增 +0

OpenMMLab 多模态高级生成与智能创作工具箱。释放魔力🪄:AIGC、易用 API、丰富模型库、扩散模型,支持文生图、图像/视频修复与增强等任务OpenMMLab Multimodal Advanced, Generative, and Intelligent Creation Toolbox. Unlock the magic 🪄: Generative-AI (AIGC), easy-to-use APIs, awsome model zoo, diffusion models, for text-to-image generation, image/video restoration/enhancement, etc.

multimodal
Moonvy/OpenPromptStudio
Vue · 2024-04-28 多模态 工具 生产可用 Stars 6648 周增 +21

🥣 AIGC 提示词可视化编辑器 | OPS | Open Prompt Studio

teamchong/pxpipe
TypeScript · 2026-07-18 多模态 工具 研究原型 Stars 6425 周增 +140

通过将文本上下文渲染为图像,将 Fable 5 的 token 使用量降低cut Fable 5 token usage by rendering text context as images

multimodal