研究库 开源仓库
Repositories · organized/repo_cards

仓库/Skill 库

17 个 · 多模态 · 模型

排序 Stars 周增
huggingface/transformers
Python · 2026-08-03 多模态 模型 生产可用 Stars 163296 周增 +0

Transformers:面向文本、视觉、音频及多模态 SOTA 机器学习模型的模型定义框架,同时支持推理与训练。🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.

multimodalllm-infra
CompVis/stable-diffusion
Jupyter Notebook · 2024-06-18 多模态 模型 实验 Stars 73272 周增 +0

潜空间文本到图像扩散模型。A latent text-to-image diffusion model

multimodal
mudler/LocalAI
Go · 2026-09-21 多模态 模型 生产可用 Stars 49205 周增 +137

LocalAI 是开源 AI 引擎,可在任何硬件上运行任意模型 —— LLMs、视觉、语音、图像、视频,无需 GPU。LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.

agentmultimodalllm-infra
suno-ai/bark
Jupyter Notebook · 2024-08-19 多模态 模型 实验 Stars 39217 周增 +0

🔊 文本提示生成式音频模型🔊 Text-Prompted Generative Audio Model

baidu/Unlimited-OCR
Python · 2026-07-29 多模态 模型 生产可用 Stars 23399 周增 +581

Unlimited OCR Works:迈入一键长文档解析的时代。Unlimited OCR Works: Welcome the Era of One-shot Long-horizon Parsing.

Lightricks/LTX-2
Python · 2026-08-03 多模态 模型 生产可用 Stars 8560 周增 +14

LTX-2 音视频生成模型的官方 Python 推理与 LoRA 训练包。Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.

multimodalllm-infra
open-mmlab/mmagic
Jupyter Notebook · 2024-08-06 多模态 模型 实验 Stars 7452 周增 +0

OpenMMLab 多模态高级生成与智能创作工具箱。释放魔力🪄:AIGC、易用 API、丰富模型库、扩散模型,支持文生图、图像/视频修复与增强等任务OpenMMLab Multimodal Advanced, Generative, and Intelligent Creation Toolbox. Unlock the magic 🪄: Generative-AI (AIGC), easy-to-use APIs, awsome model zoo, diffusion models, for text-to-image generation, image/video restoration/enhancement, etc.

multimodal
facebookresearch/vggt-omega
Python · 2026-07-02 多模态 模型 研究原型 Stars 3471 周增 +63

[CVPR 2026 Oral] VGGT Omega[CVPR 2026 Oral] VGGT Omega

MisoLabsAI/MisoTTS
Python · 2026-06-09 多模态 模型 研究原型 Stars 3131 周增 +70

Miso TTS:拥有 80 亿参数、表现力强的文本转语音模型。Miso TTS is an 8 billion, highly emotive text-to-speech model

intentee/paddler
Rust · 2026-07-19 多模态 模型 研究原型 Stars 1651 周增 +7

开源 LLM/VLM 负载均衡器与服务平台,用于规模化自托管 LLM(和 VLM)🏓🦙 作为 llm-d、Docker Model Runner 等项目的替代方案,组件更少、部署更简单,基于 ggml 生态构建。支持 CPU 和 GPU。Open-source LLM/VLM load balancer and serving platform for self-hosting LLMs (and VLMs) at scale 🏓🦙 Alternative to projects like llm-d, Docker Model Runner, etc but with less moving parts and simple deployments built around ggml ecosystem. Runs on CPU and GPU.

llm-infraengineering
Yuan-ManX/ai-audio-datasets
未知语言 · 2025-07-08 多模态 模型 实验 Stars 958 周增 +0

AI 音频数据集(AI-ADS)🎵,包含语音、音乐和音效,可为 Generative AI、AIGC、AI 模型训练、智能音频工具开发及音频应用提供训练数据。AI Audio Datasets (AI-ADS) 🎵, including Speech, Music, and Sound Effects, which can provide training data for Generative AI, AIGC, AI model training, intelligent audio tool development, and audio applications.

modelscope/richdreamer
Python · 2024-09-27 多模态 模型 实验 Stars 478 周增 +0

[CVPR2024 (Highlight)] RichDreamer:一种可泛化的法线-深度扩散模型,用于生成细节丰富的文本到 3D 内容。Live Demo:https://modelscope.cn/studios/Damo_XR_Lab/3D_AIGC[CVPR2024 (Highlight)] RichDreamer: A Generalizable Normal-Depth Diffusion Model for Detail Richness in Text-to-3D. Live Demo:https://modelscope.cn/studios/Damo_XR_Lab/3D_AIGC

liyupi/ai-model-world
TypeScript · 2026-09-20 多模态 模型 实验 Stars 144 周增 +140

AI 大模型世界,把 556 个大模型拟人化成像素小人的可视化站点。进来就能看到此刻谁最聪明、谁最会写代码、谁最便宜、谁刚发布,往下是国内与国外分区的厂商广场、完整的发布时间线和多维排行榜。搜索认模型名、厂商和能力,输入「多模态」会直接列出全部多模态模型。数据取自 Epoch AI、models.dev、LiveBench 与 Hugging Face,每小时自动同步,所有文案由真实数据生成,不调用任何 LLM。Next.js 静态导出,零后端。

evaluationllm-infra
ModelTC/Minimax-H3-Turbo
Python · 2026-08-11 多模态 模型 实验 Stars 91 周增 +0

将 Minimax-H3 蒸馏为 4 步。Distill Minimax-H3 into 4 steps

kenhayward/Diariz
C# · 2026-08-11 多模态 模型 实验 Stars 3 周增 +0

自托管、多用户转录平台:录制或上传音频。支持说话人标记、带时间戳的转录、跨录音识别说话人、摘要、提取行动项,并可使用自有 OpenAI 兼容 LLM 与转录内容对话。你的音频、你的服务器、你的模型。已在笔记本 RTX4070、台式机 RTX3090 与 RTX5090 上测试。Self-hosted, multi-user transcription platform: record or upload audio. Speaker-labeled, timestamped transcripts, Recognize speakers across recordings, Summarize, extract action items and chat over your transcripts with your own OpenAI-compatible LLM. Your Audio, your Server, your Model. Tested on Laptop RTX4070, Desktop RTX3090 and RTX5090

ragdatabasellm-infra
choxos/NMAViz
JavaScript · 2026-09-07 多模态 模型 实验 Stars 1 周增 +0

阅读网络 meta 分析的证据结构。在浏览器中拟合图论模型并以八种方式绘制:精度几何、试验、证据流、贡献、带符号的研究级重构、不一致性的 Hodge 分解、扩散与弹簧图。Read the evidence structure of a network meta-analysis. Fits the graph-theoretical model in your browser and draws it eight ways: precision geometry, trials, evidence flow, contributions, a signed study-level reconstruction, a Hodge split of inconsistency, diffusion and springs.

Erikalaylafajri15/MOSS-VL
未知语言 · 2026-09-22 多模态 模型 实验 Stars 0 周增 +0

使用 vision-language model 处理视觉与文本数据,执行多模态推理与图像理解任务。Process visual and textual data with this vision-language model for multimodal reasoning and image understanding tasks.

multimodalllm-infra