研究库 开源仓库
Repositories · organized/repo_cards

仓库/Skill 库

236 个 · 多模态

排序 Stars 周增
WenyuChiou/academic-writing-skills
Python · 2026-10-09 多模态 应用 实验 Stars 88 周增 +14

面向严谨学术论文写作、修订与投稿的 Claude Code skill。跨领域通用,支持按论文设置期刊覆盖规则。Claude Code skill for rigorous academic paper writing, revision, and submission. Field-agnostic with per-paper journal overrides.

multimodal
filliptm/ComfyUI-FL-YuE2
Python · 2026-09-15 多模态 工具 生产可用 Stars 86 周增 +0

YuE2 音乐生成,以及面向 ComfyUI 的可编辑钢琴卷帘界面。YuE2 music generation and an editable piano roll for ComfyUI

Ashlixy17/PCB_lightgraph_Portable
HTML · 2026-08-23 多模态 工具 实验 Stars 81 周增 +0

PCB_lightgraph_portable 是一款点击即可运行的 PCB 图片智能分层与图纸导出的html工具。

stabgan/openrouter-mcp-multimodal
TypeScript · 2026-08-22 多模态 应用 实验 Stars 78 周增 +0

OpenRouter 的 MCP server——与 300+ LLM(Claude、Gemini、GPT)对话,分析图像/音频/视频,生成图像/语音/音乐/视频(Veo 3.1、Sora、Seedance、Wan),支持 Claude Desktop、Cursor、Kiro、VS CodeMCP server for OpenRouter — chat with 300+ LLMs (Claude, Gemini, GPT), analyze images / audio / video, generate images / speech / music / video (Veo 3.1, Sora, Seedance, Wan) from Claude Desktop, Cursor, Kiro, VS Code.

agentmultimodalllm-infra
ttthanh2044/voxdub
Python · 2026-08-12 多模态 工具 实验 Stars 77 周增 +0
karuvanan/MiniMax-H3-Director-Cut-Studio
Python · 2026-08-29 多模态 应用 实验 Stars 77 周增 +0

受 Premiere 启发、面向 MiniMax H3 Ref2VA 的 PySide6 导演工作室,集成 AI 分镜规划、语义媒体增强、时间线 prompt 调和,以及通过 ComfyUI 实现的镜头感知长视频渲染。Premiere-inspired PySide6 director studio for MiniMax H3 Ref2VA with AI shot planning, semantic media enrichment, timeline prompt reconciliation and shot-aware long-video rendering through ComfyUI.

multimodal
A-Box-of-Tools/website
HTML · 2026-08-24 多模态 应用 实验 Stars 74 周增 +0

https://abox.tools/ 的源码——一个面向图像、视频、音频、PDF 与文本的小型 Web 工具集。文件永不离开本机,因为不存在任何能将它们外发的代码路径。Source for https://abox.tools/ — a box of small web tools for images, video, audio, PDFs and text. Your files never leave your machine, because there is no code path that could send them anywhere.

multimodal
tig3rmast3r/OFXR-Bridge
C++ · 2026-09-07 多模态 工具 实验 Stars 73 周增 +0

基于 Optical Flow 的 VR 帧生成。Frame generation for VR using Optical Flow

adunext/adu-motion-video
JavaScript · 2026-10-03 多模态 应用 实验 Stars 71 周增 +0

阿杜与 Opus 5.5 精选制作的动画动效模板 Skill。选模板和风格,用 Codex / Claude Code 把口播与素材剪成视频;提供 11,028 个精选 Lottie 动画素材 API。

arrival-space/splat.js
JavaScript · 2026-08-25 多模态 库 实验 Stars 67 周增 +112
ayandexyz/glance-linux
Python · 2026-09-22 多模态 工具 实验 Stars 67 周增 +0

面向 Linux 的真活体检测人脸解锁:一个 PAM 模块、一个守护进程,以及一个 Omarchy 锁屏指示器。Face unlock for Linux with real liveness detection: a PAM module, a daemon, and an Omarchy lock screen indicator

OneMana-Soft/OneCamp-fe
TypeScript · 2026-09-30 多模态 应用 实验 Stars 66 周增 +0

OneCamp:自托管一体化工作空间(聊天、任务、视频通话、文档、日历与 AI)。OneCamp: self-hosted all-in-one workspace (chat, tasks, video calls, docs, calendar & AI)

agentmultimodal
Vaquill-AI/open-india-law
TypeScript · 2026-08-26 多模态 应用 实验 Stars 65 周增 +0

开放、结构化的印度一手法律数据:包含最高法院与全部 25 所高等法院的 3250 万判决分块、110 万条立法条款,以及构建数据集的爬虫;采用 CC BY 4.0 许可。Open, structured Indian primary law: 32.5M judgment chunks from the Supreme Court and all 25 High Courts, 1.1M legislation provisions, and the scrapers that build it. CC BY 4.0.

multimodal
quanru/doubao-say
Python · 2026-09-20 多模态 工具 实验 Stars 64 周增 +4

由豆包驱动的语音输入工具,适用于 Linux、Wayland 和 OmarchyDoubao-powered voice input for Linux, Wayland and Omarchy

awdqwdasdg/Comfyui-Spectrum-Qwen2.1
Python · 2026-09-22 多模态 工具 实验 Stars 63 周增 +0

字面意义上的 Slop。Literal Slop.

ArtemPavlov1994/polymarket-prediction-bot
Python · 2026-09-24 多模态 应用 实验 Stars 62 周增 +7

面向 Polymarket 预测市场的交易机器人——可在终端浏览 CLOB 市场、查看订单簿,并执行 edge detection、流动性提供与跨市场套利策略,支持 paper trading 与风险限额。开源教育工具——不构成投资建议。非官方社区项目,与 Polymarket 无关。Polymarket trading bot for prediction markets — browse CLOB markets, watch the order book in the terminal, run edge detection, liquidity provision and cross-market arbitrage strategies with paper trading and risk limits. Educational open-source toolkit — not financial advice. Unofficial community project, not affiliated with Polymarket.

agentragmultimodalrisk
jlucasmcrell/ComfyUI-H3-Multishot
Python · 2026-08-11 多模态 工具 实验 Stars 62 周增 +0
JGRFW/comfyui-AICG3D
JavaScript · 2026-09-17 多模态 工具 实验 Stars 60 周增 +0

把 ComfyUI-MiniMaxH3-Easy 与 Goohai-MiniMax-H3_Integration 合并成一个统一的 MiniMax H3 创作工作台(非官方)

multimodal
coll3879xx-cyber/dola-render-gateway
Python · 2026-09-06 多模态 应用 实验 Stars 59 周增 +0

高性能视频生成网关与会话协调器。High-performance video generation gateway and session coordinator

multimodal
horvitzs/Interactive_Segmentation_Models
未知语言 · 2018-05-18 多模态 收藏榜 研究原型 Stars 58 周增 +0

交互式标注/分割的文献综述Literature Review for Interactive annotation/segmentation

multimodal
bhouston/three-dlss-nr
TypeScript · 2026-10-05 多模态 工具 实验 Stars 58 周增 +0

NVIDIA DLSS 5 神经渲染网络开源复现 OpenDLSS-NR 的 Three.js(TSL / WebGPU)移植版。A Three.js (TSL / WebGPU) port of OpenDLSS-NR, the open-source reimplementation of NVIDIA's DLSS 5 neural rendering network

youkely/awesome-visual-localization
未知语言 · 2022-07-12 多模态 收藏榜 研究原型 Stars 57 周增 +0

视觉定位的文献综述Literature review of visual localization.

ragmultimodal
siddzzzz/Music-Decoder
TypeScript · 2026-09-04 多模态 工具 实验 Stars 57 周增 +0
Jalen-Brunson/ComfyUI-MiniMax-H3-PDD-Acc
Python · 2026-08-30 多模态 应用 实验 Stars 56 周增 +0

ComfyUI 中的官方 MiniMax-H3 8 步 PDD Acc LoRA(alibaba-pai):LoRA + 并行解码头库,基于已训练 sigmas 的 euler 调度,8 步生成音视频Official MiniMax-H3 8-step PDD Acc LoRAs (alibaba-pai) in ComfyUI: LoRA + parallel-decoding head bank, euler on trained sigmas, audio+video in 8 steps

multimodal
abdelkhalk93-star/audiocraft-finetune-lab
HTML · 2026-08-30 多模态 应用 实验 Stars 55 周增 +0

MusicGen Pro Trainer 2026:终极 AI 音乐工作室指南MusicGen Pro Trainer 2026: The Ultimate AI Music Studio Guide

3laxy0/Warhound-Vision-Ultra
HTML · 2026-08-30 多模态 应用 实验 Stars 55 周增 +0

高级 WARDOGS 辅助:ESP、自瞄与载具战斗增强 2026Advanced WARDOGS Hacks: ESP, Aimbot & Vehicle Combat Enhancements 2026

matlowai/ComfyUI-MAINodes
Python · 2026-08-14 多模态 应用 实验 Stars 54 周增 +0

MatlowAI 的 MiniMax-H3 ComfyUI 节点:Contact-Sheet diffusion + Motion Lab(针对快速运动的 test-time 去绳状畸变)MatlowAI's MiniMax-H3 ComfyUI nodes: Contact-Sheet diffusion + Motion Lab (test-time de-roping of fast motion)

YZCU/OOTB
C · 2025-07-11 多模态 评测集 实验 Stars 53 周增 +0

[ISPRS 2024] 卫星视频单目标跟踪:系统综述与定向目标跟踪基准[ISPRS 2024] Satellite Video Single Object Tracking: A Systematic Review and An Oriented Object Tracking Benchmark

multimodalevaluation
Cuongyd196/cit-voice-studio
未知语言 · 2026-09-28 多模态 工具 生产可用 Stars 53 周增 +0

CIT Voice Studio —— 将越南语文本转为语音,完全在本地运行。CIT Voice Stuido - Chuyển văn bản tiếng Việt thành giọng nói, chạy hoàn toàn trên máy bạn.

zlab-princeton/VisionFoundry
Python · 2026-09-08 多模态 应用 实验 Stars 52 周增 +0

VisionFoundry:使用合成图像教会 VLM 视觉感知。VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images

multimodalllm-infra
itrcz/calab
TypeScript · 2026-10-02 多模态 应用 实验 Stars 52 周增 +0

受 Discord、Telegram 和 Zoom 启发的开源团队聊天和视频会议工具。OpenSource team chats and video conferences inspired with Discord, Telegram & Zoom

multimodal
binghe1980/PromptLens
未知语言 · 2026-09-08 多模态 工具 生产可用 Stars 51 周增 +0

一款把任意图片转换为专业生图提示词的 Chrome / Edge 浏览器插件。支持网页右键识图、本地上传与剪贴板粘贴,生成中文、英文和 JSON 提示词,并提供历史记录与收藏。

multimodal
xmutfyh/dsh-plugin-writing-guard
JavaScript · 2026-10-05 多模态 应用 实验 Stars 45 周增 +3

面向 AI 辅助研究的科学写作与文档完整性守护——论证经济性、确定性完整性、安全的 Word 编辑。本地 · 确定性 · 零 LLM。Scientific writing & document integrity guard for AI-assisted research — argument economy, deterministic integrity, safe Word editing. Local · Deterministic · Zero LLM.

llm-infra
xulihang/ImageTrans_plugins
B4X · 2026-08-11 多模态 库 实验 Stars 36 周增 +0

ImageTrans 插件集合。Plugins for ImageTrans

multimodalllm-infra
Syh1906/openai-compatible-imagegen
JavaScript · 2026-08-21 多模态 应用 实验 Stars 34 周增 +0

独立的 Skill 与 Codex 插件,支持 OpenAI 兼容的图像生成、编辑、批处理工作流、QA 以及聚焦画布编辑。Standalone Skill and Codex Plugin for OpenAI-compatible image generation, editing, batch workflows, QA, and focused canvas editing.

agentmultimodal
yuanzh0u/embodied-learning
HTML · 2026-09-28 多模态 应用 实验 Stars 7 周增 +0

证据优先的具身 AI 研究枢纽,覆盖 VLA、世界模型、多模态感知、文献综述、声明审计以及可复用的 Agent skillEvidence-first Embodied AI research hub for VLA, world models, multimodal sensing, literature reviews, claim audits, and reusable agent skills.

agentmultimodal