Repositories · organized/repo_cards

仓库/Skill 库

73 个 · 多模态 · 应用

排序 Stars 周增
ruvnet/RuView
Rust · 2026-08-11 多模态 应用 生产可用 Stars 89443 周增 +1218

π RuView 将现成 WiFi 信号转化为实时空间智能、生命体征监测和存在检测,全程无需任何视频画面。π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video.

multimodalrisk
coqui-ai/TTS
Python · 2024-08-16 多模态 应用 实验 Stars 45852 周增 +0

🐸💬 —— 一个历经研究与生产环境考验的 Text-to-Speech 深度学习工具包。🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production

engineering
cjpais/Handy
Rust · 2026-08-03 多模态 应用 生产可用 Stars 28596 周增 +0

一款免费、开源且可扩展的语音转文字应用程序,完全离线运行。A free, open source, and extensible speech-to-text application that works completely offline.

mozilla/DeepSpeech
C++ · 2025-06-19 多模态 应用 实验 Stars 26770 周增 +0

DeepSpeech 是一个开源嵌入式(离线、端侧)语音转文字引擎,可在从 Raspberry Pi 4 到高性能 GPU 服务器的设备上实时运行。DeepSpeech is an open source embedded (offline, on-device) speech-to-text engine which can run in real time on devices ranging from a Raspberry Pi 4 to high power GPU servers.

wandb/openui
TypeScript · 2026-08-05 多模态 应用 生产可用 Stars 22494 周增 +7

OpenUI 可让你凭想象力描述 UI,并实时看到渲染效果。OpenUI let's you describe UI using your imagination, then see it rendered live.

index-tts/index-tts
Python · 2026-07-14 多模态 应用 生产可用 Stars 22384 周增 +0

工业级可控、高效的零样本文本转语音系统An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Acly/krita-ai-diffusion
Python · 2026-06-30 多模态 应用 生产可用 Stars 10449 周增 +21

在 Krita 中使用 AI 生成图像的精简界面。支持 Inpaint 与 Outpaint,可选文本提示,无需调参。Streamlined interface for generating images with AI in Krita. Inpaint and outpaint with optional text prompt, no tweaking required.

multimodal
AIGC-Audio/AudioGPT
Python · 2024-07-06 多模态 应用 实验 Stars 10171 周增 +0

AudioGPT:理解与生成语音、音乐、声音与说话人头像AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head

xorbitsai/inference
Python · 2026-08-11 多模态 应用 生产可用 Stars 9486 周增 +0

通过修改一行代码即可将 GPT 替换为任意 LLM。Xinference 让你在云端、本地或笔记本上运行开源、语音和多模态模型,全部通过统一的生产就绪推理 API。Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.

multimodalllm-infraengineering
XxHuberrr/Mineradio
JavaScript · 2026-07-28 多模态 应用 生产可用 Stars 9412 周增 +196

一款以电影镜头、粒子视觉和歌词舞台为核心的沉浸式音乐播放器。

oumi-ai/oumi
Python · 2026-08-11 多模态 应用 生产可用 Stars 9369 周增 +7

轻松微调、评估和部署 Gemma 4、Qwen3.5、Qwen3.6、gpt-oss、DeepSeek-R1 或任意开源 LLM / VLM。Easily fine-tune, evaluate and deploy Qwen, Gemma, or any open weight LLM!

llm-infraevaluation
oso95/scroll-world
JavaScript · 2026-07-29 多模态 应用 生产可用 Stars 7924 周增 +189

将任意品牌转化为可滚动 3D 世界落地页的 skillA skill that turn any brand into a scrollable 3D world landing page

argmaxinc/argmax-oss-swift
Swift · 2026-08-06 多模态 应用 生产可用 Stars 6317 周增 +7

面向 Apple Silicon 的端侧语音 AI。On-device Speech AI for Apple Silicon

llm-infra
modstart-lib/aigcpanel
TypeScript · 2026-07-16 多模态 应用 生产可用 Stars 5434 周增 +14

AIGCPanel 是一个简单易用的一站式AI数字人系统,支持视频合成、声音合成、声音克隆,简化本地模型管理、一键导入和使用AI模型。

diffusionstudio/lottie
TypeScript · 2026-07-25 多模态 应用 研究原型 Stars 5020 周增 +28

使用 Claude Code 或 Codex 生成可投产的 Lottie 动画Generate production-ready Lottie animations with Claude Code or Codex

engineering
WUBING2023/PaperSpine
Python · 2026-07-01 多模态 应用 研究原型 Stars 4041 周增 +105

PaperSpine 是以动机驱动的 Skill,用于研读高质量学术论文、构建论文核心论点,并通过证据感知蓝图、修订矩阵与 LaTeX 安全审计来重写稿件。PaperSpine is a motivation-driven skill for learning from strong academic papers, building a paper’s central argument, and rewriting manuscripts through evidence-aware blueprints, revision matrices, and LaTeX-safe audits.

multimodalengineering
digimata/quill
Swift · 2026-07-30 多模态 应用 生产可用 Stars 3763 周增 +56

极简的 macOS 录音 + 转写工具。Ultra-minimalist macOS recording + transcription.

dramaclaw/dramaclaw
TypeScript · 2026-08-11 多模态 应用 研究原型 Stars 3493 周增 -126

A general-purpose AIGC video engine: script to finished film in one pipeline — dramas, ads, product videos, otome games, and more. | 通用 AIGC 视频引擎 —— 从剧本到成片一条流水线,漫剧、广告、电商、乙游皆可

agentmultimodalengineering
LearnPrompt/LearnPrompt
MDX · 2026-08-10 多模态 应用 研究原型 Stars 2580 周增 +0

永久免费开源的 AIGC 课程, 目前已支持Claude Code,Codex,Hermes,OpenClaw,Obsidian,Prompt Engineering, ChatGPT, Midjourney, Runway, Stable Diffusion, AI数字人,AI声音&音乐,开源大模型

agentllm-infra
aigc-apps/EasyAnimate
Python · 2025-03-06 多模态 应用 实验 Stars 2266 周增 +0

📺 基于 Transformer Diffusion 的高分辨率长视频生成端到端解决方案。📺 An End-to-End Solution for High-Resolution and Long Video Generation Based on Transformer Diffusion

multimodal
all-in-aigc/aicover
TypeScript · 2025-01-24 多模态 应用 生产可用 Stars 1807 周增 +0

AI 封面生成器。ai cover generator

NVIDIA-AI-Blueprints/video-search-and-summarization
C++ · 2026-08-11 多模态 应用 研究原型 Stars 1790 周增 +7

NVIDIA AI Blueprint for video search and summarization(VSS)是一个 GPU 加速参考架构,用于构建具备实时验证告警、视觉问答与自动报告能力的视频分析 Agent。VSS Blueprint 采用 NVIDIA Cosmos 等视觉语言模型(VLM)、NVIDIA Nemotron 等 LLM,并结合 RAG 与 NVIDIA NIM。NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts, visual Q&A, and automated reporting. The VSS Blueprint uses vision language models (VLMs) such as NVIDIA Cosmos, LLMs such as NVIDIA Nemotron, RAG, and NVIDIA NIMs.

agentragmultimodalllm-infra
xLLM-AI/xllm
C++ · 2026-08-12 多模态 应用 研究原型 Stars 1516 周增 +0

面向 LLM、VLM、DiT 和 REC 模型的高性能推理引擎,针对多种 AI 加速器进行了优化。该项目托管于 OpenAtom 基金会。A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.

llm-infra
jd-opensource/xllm
C++ · 2026-07-06 多模态 应用 研究原型 Stars 1390 周增 +28

高性能推理引擎,支持 LLM、VLM、DiT 和 REC 模型,针对多种 AI 加速器优化A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators.

llm-infra
all-in-aigc/sorafm
TypeScript · 2024-08-15 多模态 应用 实验 Stars 1152 周增 +14

Sora.FM 推出的 Sora AI 视频生成器。Sora AI Video Generator by Sora.FM

multimodal
TypeTale/TypeTale
未知语言 · 2026-01-06 多模态 应用 生产可用 Stars 776 周增 +14

字字动画 - 完全免费的AIGC视频生成软件,主要用于AI短剧,AI电影,小说推文

Small-tailqwq/dsh-deep-whale
TypeScript · 2026-08-15 多模态 应用 研究原型 Stars 669 周增 +0

DSH Web 鲸鱼娘皮肤系列(深海女仆工坊 maid-atelier)——CC BY-NC-SA 4.0

Alain00/blobatar
TypeScript · 2026-08-20 多模态 应用 研究原型 Stars 599 周增 +0
all-in-aigc/melodisco
TypeScript · 2024-09-16 多模态 应用 生产可用 Stars 514 周增 +0

AI Music PlayerAI Music Player

Moonlit-Pages/AIGC-Detector-Rewriter-Skill
未知语言 · 2026-06-19 多模态 应用 研究原型 Stars 488 周增 +0

一个面向英文学术写作的保守型 AIGC 检测器指导的论文改写 Skill。支持 Turnitin AI、CNKI AIGC、最小化编辑修订、保留学术要素、定性/定量路由,以及逐章降低 AI 写作风险,且不宣称绕过检测器。A conservative AIGC detector-informed thesis rewriting skill for English and Chinese academic writing. Supports Turnitin AI, CNKI AIGC, minimal-edit revision, protected academic elements, qualitative/quantitative routing, and chapter-by-chapter AI-writing risk reduction without detector-bypass claims.

multimodalriskengineering
modu-ai/cowork-plugins
HTML · 2026-06-19 多模态 应用 实验 Stars 267 周增 +7

所有人可用的 AI(MoAI)——面向韩语实务领域的 Claude Cowork 与 Claude Code AI harness 与插件市场。覆盖商业计划书、税务、法律、HR、营销、电商、BI、内容等领域,提供 Skill、Agent 与工作流。支持韩语 B2B 场景与办公文档(HWPX/DOCX/XLSX/PPTX/PDF)及 AI 多模态生成(图像/视频/语音)。内置 AI 痕迹审核与韩语 humanize-korean모두의 AI (MoAI) — Claude Cowork & Claude Code 한국 실무 도메인 AI 하네스·플러그인 마켓플레이스. 사업계획서·세무·법률·HR·마케팅·커머스·BI·콘텐츠 도메인 스킬·에이전트·워크플로우. Korean B2B + office docs (HWPX/DOCX/XLSX/PPTX/PDF) + AI media (image/video/voice). AI-slop 검수 + humanize-korean 내장.

agentmultimodal
VisionVerse/RemoteSensing-Restoration-Survey
Python · 2026-08-09 多模态 应用 实验 Stars 231 周增 +0

[ISPRS 2026] 遥感图像去雾:进展、挑战与前景的系统综述[ISPRS 2026] Remote Sensing Image Dehazing: A Systematic Review of Progress, Challenges, and Prospects

multimodal
tritant/ComfyUI_MiniMax_H3_Extender
Python · 2026-08-21 多模态 应用 实验 Stars 129 周增 +0

ComfyUI 节点,为 MiniMax H3 而设计,可串联多个视频片段,具备运动上下文、磁盘缓存、动态图像参考、音频参考支持以及最终视频/音频的无缝解码ComfyUI node for MiniMax H3 that chains multiple video clips with motion context, disk caching, dynamic image references, audio reference support, and seamless final video/audio decoding.

multimodal
kasturikhanke/generative-loaders
TypeScript · 2026-08-12 多模态 应用 实验 Stars 111 周增 +0

面向生成式接口的无障碍 React 加载状态:流式文本、内联活动指示与图像生成。Accessible React loading states for generative interfaces: streamed text, inline activity, and image generation.

multimodal
TarunTomar122/better-voice
Swift · 2026-08-24 多模态 应用 实验 Stars 103 周增 +0

语音听写,同时结合你所指屏幕位置的上下文。Voice dictation with the screen context you point at.

jaredrhod/barehands
HTML · 2026-08-17 多模态 应用 实验 Stars 103 周增 +0

用双手操控屏幕——基于摄像头的免穿戴、免手柄手势追踪界面,让你的 AI 直接响应动作Move things on your screen with your bare hands. A webcam-powered, hand-tracked interface for your AI. No headset. No controllers.