研究库 开源仓库
Repositories · organized/repo_cards

仓库/Skill 库

112 个 · 多模态 · 应用

排序 Stars 周增
oboroge0/hayamimi
Python · 2026-08-31 多模态 应用 实验 Stars 313 周增 +91

早耳——仅依赖 CPU 的实时多语种语音转文字。支持实时字幕、浏览器仪表盘、说话人标签与翻译,无需 GPU,无需云端。早耳 - Real-time multilingual speech-to-text on CPU only. Live subtitles, browser dashboard, speaker labels, translation. No GPU, no cloud.

timoncool/YuE2-Studio
TypeScript · 2026-09-28 多模态 应用 实验 Stars 275 周增 +371

本地 AI 歌曲生成器,支持可编辑乐谱 —— 在你的 GPU 上运行 YuE2:带人声的全曲、五线谱、翻唱与精确回放。原生 Windows 应用,无需 Python,安装包自带自动更新。Local AI song generator with an editable score — YuE2 on your GPU: full songs with vocals, sheet music, covers, exact replay. Native Windows app, no Python, installer with auto-update.

modu-ai/cowork-plugins
HTML · 2026-06-19 多模态 应用 实验 Stars 267 周增 +7

所有人可用的 AI(MoAI)——面向韩语实务领域的 Claude Cowork 与 Claude Code AI harness 与插件市场。覆盖商业计划书、税务、法律、HR、营销、电商、BI、内容等领域,提供 Skill、Agent 与工作流。支持韩语 B2B 场景与办公文档(HWPX/DOCX/XLSX/PPTX/PDF)及 AI 多模态生成(图像/视频/语音)。内置 AI 痕迹审核与韩语 humanize-korean모두의 AI (MoAI) — Claude Cowork & Claude Code 한국 실무 도메인 AI 하네스·플러그인 마켓플레이스. 사업계획서·세무·법률·HR·마케팅·커머스·BI·콘텐츠 도메인 스킬·에이전트·워크플로우. Korean B2B + office docs (HWPX/DOCX/XLSX/PPTX/PDF) + AI media (image/video/voice). AI-slop 검수 + humanize-korean 내장.

agentmultimodal
soirihiroka/shrimply
Rust · 2026-08-31 多模态 应用 实验 Stars 264 周增 +0

你是说这条视频是一只虾做的?you're telling me a shrimp made this video?

multimodal
gatewai-dev/framefields
TypeScript · 2026-10-06 多模态 应用 实验 Stars 251 周增 +0

代码优先的视频,原生渲染于 WebGPUCode-first video, rendered natively on WebGPU.

multimodal
Merserk/dlss5-visual-enhancer
Python · 2026-09-02 多模态 应用 实验 Stars 237 周增 +0

DLSS 5 神经视频与图像增强器DLSS 5 Neural Video & Image Enhancer

multimodal
VisionVerse/RemoteSensing-Restoration-Survey
Python · 2026-08-09 多模态 应用 实验 Stars 231 周增 +0

[ISPRS 2026] 遥感图像去雾:进展、挑战与前景的系统综述[ISPRS 2026] Remote Sensing Image Dehazing: A Systematic Review of Progress, Challenges, and Prospects

multimodal
razorback16/openjev
Python · 2026-09-21 多模态 应用 实验 Stars 225 周增 +0

基于 DiffusionGemma 构建的开放、Jev 兼容的 System One 决策服务器。Open, Jev-compatible System One decision server on DiffusionGemma

jankeesvw/omarchy-meeting-recorder
Rust · 2026-09-25 多模态 应用 实验 Stars 216 周增 +525

在 Omarchy 上录制会议:将麦克风与电脑音频分别录为两条音轨,在本地机器上转录,包含说话人、章节与播放器。Record meetings on Omarchy: mic and computer audio as two tracks, transcribed on your own machine, with speakers, chapters and a player.

opencoredev/bg0
TypeScript · 2026-09-16 多模态 应用 实验 Stars 199 周增 +0

私密、无限制的背景移除,在浏览器中即可运行。Private, unlimited background removal that runs in your browser.

sqzw-x/amane
Python · 2026-08-27 多模态 应用 实验 Stars 197 周增 +0

AI 时代的私人影库

bridge-mind/bridgeclip
TypeScript · 2026-09-25 多模态 应用 实验 Stars 195 周增 +910

BridgeMind 开源的 AI 视频剪辑桌面应用Open-source AI video clipping desktop app by BridgeMind

multimodal
telepath-computer/television
TypeScript · 2026-10-07 多模态 应用 实验 Stars 191 周增 +0

个人 Agent 缺失的 GUI。Television 为你和你的 Agent 提供用于创建和处理工件的可视化空间。The missing GUI for personal agents. Television gives you and your agent a visual space for creating and working with artifacts.

agentmultimodal
gulelmatthews/Polymarket-Perpetual-Bot
Python · 2026-09-27 多模态 应用 实验 Stars 183 周增 +763

面向 Polymarket 预测市场的交易机器人——浏览 CLOB 市场、在终端查看订单簿、运行套利检测、流动性提供与跨市场套利策略,支持模拟交易和风险限额。教育性开源工具包——不构成投资建议。非官方社区项目,与 Polymarket 无关。Polymarket trading bot for prediction markets — browse CLOB markets, watch the order book in the terminal, run edge detection, liquidity provision and cross-market arbitrage strategies with paper trading and risk limits. Educational open-source toolkit — not financial advice. Unofficial community project, not affiliated with Polymarket.

agentragmultimodalrisk
int64ago/vistep
TypeScript · 2026-09-12 多模态 应用 实验 Stars 183 周增 +0

用 AI 可视化每一步——双语可视化讲解、交互式模型与同步旁白。Visualize Every Step with AI — bilingual visual explanations, interactive models and synchronized narration.

Atomicx7/Duo-animation
Kotlin · 2026-09-12 多模态 应用 实验 Stars 159 周增 +0
blixvip/MotionClone
Python · 2026-09-14 多模态 应用 实验 Stars 152 周增 +0

使用 Codex + ChatGPT 将参考视频转换为可编辑的动态图形;支持对比、定制并导出 MP4 或 HyperFrames 项目;提供本地 Windows 应用与在线工作台。Turn reference videos into editable motion graphics with Codex + ChatGPT. Compare, customize, and export MP4s or HyperFrames projects. Local Windows app + online studio.

agentmultimodal
gillesgoetsch/OmacVM
Shell · 2026-10-07 多模态 应用 实验 Stars 150 周增 +70

在 Mac 上的虚拟机中运行 Omarchy(OmacVM.app、UTM、VMware Fusion 或 Parallels),获得原生般的体验:Omarchy 栏位于刘海旁,支持触控板手势与类 macOS 滚动,复用 Mac 的 Wi-Fi、音频和键盘。一条命令即可构建和切换功能。Omarchy in a VM on your Mac (OmacVM.app, UTM, VMware Fusion or Parallels), feeling native: Omarchy's bar beside the notch, trackpad gestures, macOS-like scrolling, the Mac's Wi-Fi, audio and keys in Omarchy. One command to build and switch features.

tritant/ComfyUI_MiniMax_H3_Extender
Python · 2026-08-21 多模态 应用 实验 Stars 129 周增 +0

ComfyUI 节点,为 MiniMax H3 而设计,可串联多个视频片段,具备运动上下文、磁盘缓存、动态图像参考、音频参考支持以及最终视频/音频的无缝解码ComfyUI node for MiniMax H3 that chains multiple video clips with motion context, disk caching, dynamic image references, audio reference support, and seamless final video/audio decoding.

multimodal
kasturikhanke/generative-loaders
TypeScript · 2026-08-12 多模态 应用 实验 Stars 111 周增 +0

面向生成式接口的无障碍 React 加载状态:流式文本、内联活动指示与图像生成。Accessible React loading states for generative interfaces: streamed text, inline activity, and image generation.

multimodal
TarunTomar122/better-voice
Swift · 2026-08-24 多模态 应用 实验 Stars 103 周增 +0

语音听写,同时结合你所指屏幕位置的上下文。Voice dictation with the screen context you point at.

jaredrhod/barehands
HTML · 2026-08-17 多模态 应用 实验 Stars 103 周增 +0

用双手操控屏幕——基于摄像头的免穿戴、免手柄手势追踪界面,让你的 AI 直接响应动作Move things on your screen with your bare hands. A webcam-powered, hand-tracked interface for your AI. No headset. No controllers.

lxj5820/dsh-boot-animation
JavaScript · 2026-10-04 多模态 应用 实验 Stars 102 周增 +0

DSH 插件:将内核启动页替换为全窗口视频片段,随后渐变过渡到应用。中文说明见 MANUAL.md。DSH plugin: replaces the kernel boot page with a full-window video clip, then dissolves into the app. 中文说明见 MANUAL.md

multimodal
canberk7/ema-lightning
Python · 2026-10-07 多模态 应用 实验 Stars 101 周增 +0

轻量、快速且准确的土耳其语 TTS。860 万参数,在 Freya-TR-Eval 上 WER 为 0.92%,单 GPU 上首段音频延迟约 4 ms,速度达 1.300× 实时。支持音频流式输出、单 GPU 多调用者批处理,并可在 GPU 或 CPU 上离线运行。Tiny, fast and accurate Turkish text-to-speech. 8.6M parameters, 0.92% WER on Freya-TR-Eval, first audio in ~4 ms and 1,300× real time on one GPU. Streams audio, batches many callers on one GPU, and runs offline on a GPU or CPU.

WenyuChiou/academic-writing-skills
Python · 2026-10-09 多模态 应用 实验 Stars 88 周增 +14

面向严谨学术论文写作、修订与投稿的 Claude Code skill。跨领域通用,支持按论文设置期刊覆盖规则。Claude Code skill for rigorous academic paper writing, revision, and submission. Field-agnostic with per-paper journal overrides.

multimodal
stabgan/openrouter-mcp-multimodal
TypeScript · 2026-08-22 多模态 应用 实验 Stars 78 周增 +0

OpenRouter 的 MCP server——与 300+ LLM(Claude、Gemini、GPT)对话,分析图像/音频/视频,生成图像/语音/音乐/视频(Veo 3.1、Sora、Seedance、Wan),支持 Claude Desktop、Cursor、Kiro、VS CodeMCP server for OpenRouter — chat with 300+ LLMs (Claude, Gemini, GPT), analyze images / audio / video, generate images / speech / music / video (Veo 3.1, Sora, Seedance, Wan) from Claude Desktop, Cursor, Kiro, VS Code.

agentmultimodalllm-infra
karuvanan/MiniMax-H3-Director-Cut-Studio
Python · 2026-08-29 多模态 应用 实验 Stars 77 周增 +0

受 Premiere 启发、面向 MiniMax H3 Ref2VA 的 PySide6 导演工作室,集成 AI 分镜规划、语义媒体增强、时间线 prompt 调和,以及通过 ComfyUI 实现的镜头感知长视频渲染。Premiere-inspired PySide6 director studio for MiniMax H3 Ref2VA with AI shot planning, semantic media enrichment, timeline prompt reconciliation and shot-aware long-video rendering through ComfyUI.

multimodal
A-Box-of-Tools/website
HTML · 2026-08-24 多模态 应用 实验 Stars 74 周增 +0

https://abox.tools/ 的源码——一个面向图像、视频、音频、PDF 与文本的小型 Web 工具集。文件永不离开本机,因为不存在任何能将它们外发的代码路径。Source for https://abox.tools/ — a box of small web tools for images, video, audio, PDFs and text. Your files never leave your machine, because there is no code path that could send them anywhere.

multimodal
adunext/adu-motion-video
JavaScript · 2026-10-03 多模态 应用 实验 Stars 71 周增 +0

阿杜与 Opus 5.5 精选制作的动画动效模板 Skill。选模板和风格,用 Codex / Claude Code 把口播与素材剪成视频;提供 11,028 个精选 Lottie 动画素材 API。

OneMana-Soft/OneCamp-fe
TypeScript · 2026-09-30 多模态 应用 实验 Stars 66 周增 +0

OneCamp:自托管一体化工作空间(聊天、任务、视频通话、文档、日历与 AI)。OneCamp: self-hosted all-in-one workspace (chat, tasks, video calls, docs, calendar & AI)

agentmultimodal
Vaquill-AI/open-india-law
TypeScript · 2026-08-26 多模态 应用 实验 Stars 65 周增 +0

开放、结构化的印度一手法律数据:包含最高法院与全部 25 所高等法院的 3250 万判决分块、110 万条立法条款,以及构建数据集的爬虫;采用 CC BY 4.0 许可。Open, structured Indian primary law: 32.5M judgment chunks from the Supreme Court and all 25 High Courts, 1.1M legislation provisions, and the scrapers that build it. CC BY 4.0.

multimodal
ArtemPavlov1994/polymarket-prediction-bot
Python · 2026-09-24 多模态 应用 实验 Stars 62 周增 +7

面向 Polymarket 预测市场的交易机器人——可在终端浏览 CLOB 市场、查看订单簿,并执行 edge detection、流动性提供与跨市场套利策略,支持 paper trading 与风险限额。开源教育工具——不构成投资建议。非官方社区项目,与 Polymarket 无关。Polymarket trading bot for prediction markets — browse CLOB markets, watch the order book in the terminal, run edge detection, liquidity provision and cross-market arbitrage strategies with paper trading and risk limits. Educational open-source toolkit — not financial advice. Unofficial community project, not affiliated with Polymarket.

agentragmultimodalrisk
coll3879xx-cyber/dola-render-gateway
Python · 2026-09-06 多模态 应用 实验 Stars 59 周增 +0

高性能视频生成网关与会话协调器。High-performance video generation gateway and session coordinator

multimodal
Jalen-Brunson/ComfyUI-MiniMax-H3-PDD-Acc
Python · 2026-08-30 多模态 应用 实验 Stars 56 周增 +0

ComfyUI 中的官方 MiniMax-H3 8 步 PDD Acc LoRA(alibaba-pai):LoRA + 并行解码头库,基于已训练 sigmas 的 euler 调度,8 步生成音视频Official MiniMax-H3 8-step PDD Acc LoRAs (alibaba-pai) in ComfyUI: LoRA + parallel-decoding head bank, euler on trained sigmas, audio+video in 8 steps

multimodal
abdelkhalk93-star/audiocraft-finetune-lab
HTML · 2026-08-30 多模态 应用 实验 Stars 55 周增 +0

MusicGen Pro Trainer 2026:终极 AI 音乐工作室指南MusicGen Pro Trainer 2026: The Ultimate AI Music Studio Guide

3laxy0/Warhound-Vision-Ultra
HTML · 2026-08-30 多模态 应用 实验 Stars 55 周增 +0

高级 WARDOGS 辅助:ESP、自瞄与载具战斗增强 2026Advanced WARDOGS Hacks: ESP, Aimbot & Vehicle Combat Enhancements 2026