研究库 开源仓库
Repositories · organized/repo_cards

仓库/Skill 库

90 个 · 多模态 · 快速增长

排序 Stars 周增
thiagotigaz/ocr-it
JavaScript · 2026-08-25 多模态 工具 生产可用 Stars 234 周增 +175

Chrome 扩展:固定一个屏幕区域后,可用快捷键翻阅分页文档。OCR 通过内置 Tesseract 100% 离线运行。Chrome extension: pin a screen region once, then hotkey your way through a paginated document. OCR runs 100% offline via bundled Tesseract.

EverettFish/holo-card-studio
Python · 2026-09-07 多模态 工具 实验 Stars 231 周增 +0

将用户的描述或上传的参考素材转化为一张成品的、可编辑的 Blender 卡片以及一个交互式 Three.js 页面。保留所要求的主题、风格、字体与目标场景。本 skill 仅包含代码与文本;生成的艺术作品属于用户的输出项目。Turn the user's description or uploaded reference into a finished, editable Blender card and an interactive Three.js page. Preserve the requested subject, style, typography and destination. This skill contains code and text only; generated artwork belongs in the user's output project.

mizorewww/course2md
Rust · 2026-09-03 多模态 工具 实验 Stars 227 周增 +224

将 YouTube、Bilibili 或本地课程/会议录像转换为带幻灯片配图的 Markdown 与 HTML 讲义。Turn YouTube, Bilibili, or local course/meeting recordings into slide-illustrated Markdown and HTML lecture notes.

razorback16/openjev
Python · 2026-09-21 多模态 应用 实验 Stars 225 周增 +0

基于 DiffusionGemma 构建的开放、Jev 兼容的 System One 决策服务器。Open, Jev-compatible System One decision server on DiffusionGemma

programasweights/claudish
Python · 2026-08-27 多模态 工具 实验 Stars 224 周增 +0

通过小型 ProgramAsWeights 函数实现英语与 Claudish 之间的相互翻译Translate between English and Claudish with tiny ProgramAsWeights functions.

jankeesvw/omarchy-meeting-recorder
Rust · 2026-09-25 多模态 应用 实验 Stars 216 周增 +525

在 Omarchy 上录制会议:将麦克风与电脑音频分别录为两条音轨,在本地机器上转录,包含说话人、章节与播放器。Record meetings on Omarchy: mic and computer audio as two tracks, transcribed on your own machine, with speakers, chapters and a player.

opencoredev/bg0
TypeScript · 2026-09-16 多模态 应用 实验 Stars 199 周增 +0

私密、无限制的背景移除,在浏览器中即可运行。Private, unlimited background removal that runs in your browser.

sqzw-x/amane
Python · 2026-08-27 多模态 应用 实验 Stars 197 周增 +0

AI 时代的私人影库

bridge-mind/bridgeclip
TypeScript · 2026-09-25 多模态 应用 实验 Stars 195 周增 +910

BridgeMind 开源的 AI 视频剪辑桌面应用Open-source AI video clipping desktop app by BridgeMind

multimodal
telepath-computer/television
TypeScript · 2026-10-07 多模态 应用 实验 Stars 191 周增 +0

个人 Agent 缺失的 GUI。Television 为你和你的 Agent 提供用于创建和处理工件的可视化空间。The missing GUI for personal agents. Television gives you and your agent a visual space for creating and working with artifacts.

agentmultimodal
gulelmatthews/Polymarket-Perpetual-Bot
Python · 2026-09-27 多模态 应用 实验 Stars 183 周增 +763

面向 Polymarket 预测市场的交易机器人——浏览 CLOB 市场、在终端查看订单簿、运行套利检测、流动性提供与跨市场套利策略,支持模拟交易和风险限额。教育性开源工具包——不构成投资建议。非官方社区项目,与 Polymarket 无关。Polymarket trading bot for prediction markets — browse CLOB markets, watch the order book in the terminal, run edge detection, liquidity provision and cross-market arbitrage strategies with paper trading and risk limits. Educational open-source toolkit — not financial advice. Unofficial community project, not affiliated with Polymarket.

agentragmultimodalrisk
int64ago/vistep
TypeScript · 2026-09-12 多模态 应用 实验 Stars 183 周增 +0

用 AI 可视化每一步——双语可视化讲解、交互式模型与同步旁白。Visualize Every Step with AI — bilingual visual explanations, interactive models and synchronized narration.

LynnReal-AI/LynnReal-Omni
Python · 2026-09-18 多模态 框架 实验 Stars 182 周增 +0

LynnReal-Omni 将文生视频、图生视频、人体与手部姿态引导生成、结构控制、全模态参考生成、风格迁移、视频编辑、受损视频修复以及流式长视频生成统一到一个框架内,且全部支持四步快速生成。LynnReal-Omni brings text-to-video, image-to-video, human- and hand-pose guided generation, structural control, omni-reference generation, style transfer, video editing, degraded-video restoration and streaming long-video generation into a single framework, all at four-step fast generation.

multimodalevaluation
lihaoyun6/ComfyUI-H3VAE_TRT
Python · 2026-09-06 多模态 工具 实验 Stars 180 周增 +0

在 ComfyUI 中运行 MiniMax-H3 VAE 的 ONNX/TRT 版本,速度最高提升至 1.7 倍。Running ONNX/TRT version of the MiniMax-H3 VAE in ComfyUI, which increase speed by up to 1.7x

guanmo-ai/awesome-ai-motion
JavaScript · 2026-10-02 多模态 收藏榜 实验 Stars 160 周增 +0

精选 AI 动画与视频:作品封面、可播放案例、作者原始提示词与来源。Curated AI motion & video with original prompts. Focused on Claude Opus 5.5.

multimodal
Atomicx7/Duo-animation
Kotlin · 2026-09-12 多模态 应用 实验 Stars 159 周增 +0
AlLHHH/ALH-Pro
C# · 2026-09-05 多模态 工具 生产可用 Stars 155 周增 +0

图片视频等增强工具

blixvip/MotionClone
Python · 2026-09-14 多模态 应用 实验 Stars 152 周增 +0

使用 Codex + ChatGPT 将参考视频转换为可编辑的动态图形;支持对比、定制并导出 MP4 或 HyperFrames 项目;提供本地 Windows 应用与在线工作台。Turn reference videos into editable motion graphics with Codex + ChatGPT. Compare, customize, and export MP4s or HyperFrames projects. Local Windows app + online studio.

agentmultimodal
gillesgoetsch/OmacVM
Shell · 2026-10-07 多模态 应用 实验 Stars 150 周增 +70

在 Mac 上的虚拟机中运行 Omarchy(OmacVM.app、UTM、VMware Fusion 或 Parallels),获得原生般的体验:Omarchy 栏位于刘海旁,支持触控板手势与类 macOS 滚动,复用 Mac 的 Wi-Fi、音频和键盘。一条命令即可构建和切换功能。Omarchy in a VM on your Mac (OmacVM.app, UTM, VMware Fusion or Parallels), feeling native: Omarchy's bar beside the notch, trackpad gestures, macOS-like scrolling, the Mac's Wi-Fi, audio and keys in Omarchy. One command to build and switch features.

liyupi/ai-model-world
TypeScript · 2026-09-20 多模态 模型 实验 Stars 144 周增 +140

AI 大模型世界,把 556 个大模型拟人化成像素小人的可视化站点。进来就能看到此刻谁最聪明、谁最会写代码、谁最便宜、谁刚发布,往下是国内与国外分区的厂商广场、完整的发布时间线和多维排行榜。搜索认模型名、厂商和能力,输入「多模态」会直接列出全部多模态模型。数据取自 Epoch AI、models.dev、LiveBench 与 Hugging Face,每小时自动同步,所有文案由真实数据生成,不调用任何 LLM。Next.js 静态导出,零后端。

evaluationllm-infra
tritant/ComfyUI_MiniMax_H3_Extender
Python · 2026-08-21 多模态 应用 实验 Stars 129 周增 +0

ComfyUI 节点,为 MiniMax H3 而设计,可串联多个视频片段,具备运动上下文、磁盘缓存、动态图像参考、音频参考支持以及最终视频/音频的无缝解码ComfyUI node for MiniMax H3 that chains multiple video clips with motion context, disk caching, dynamic image references, audio reference support, and seamless final video/audio decoding.

multimodal
GENEXIS-AI/gpt-image-skill
JavaScript · 2026-08-28 多模态 工具 生产可用 Stars 126 周增 +0

使用 ChatGPT 订阅在 Codex 或 Claude Code 中生成 GPT 图像,无需 Images API。Generate GPT images from Codex or Claude Code using a ChatGPT subscription, without the Images API.

agentmultimodal
kasturikhanke/generative-loaders
TypeScript · 2026-08-12 多模态 应用 实验 Stars 111 周增 +0

面向生成式接口的无障碍 React 加载状态:流式文本、内联活动指示与图像生成。Accessible React loading states for generative interfaces: streamed text, inline activity, and image generation.

multimodal
facok/comfyui-SelfLift
Python · 2026-09-12 多模态 工具 实验 Stars 106 周增 +124
TarunTomar122/better-voice
Swift · 2026-08-24 多模态 应用 实验 Stars 103 周增 +0

语音听写,同时结合你所指屏幕位置的上下文。Voice dictation with the screen context you point at.

jaredrhod/barehands
HTML · 2026-08-17 多模态 应用 实验 Stars 103 周增 +0

用双手操控屏幕——基于摄像头的免穿戴、免手柄手势追踪界面,让你的 AI 直接响应动作Move things on your screen with your bare hands. A webcam-powered, hand-tracked interface for your AI. No headset. No controllers.

lxj5820/dsh-boot-animation
JavaScript · 2026-10-04 多模态 应用 实验 Stars 102 周增 +0

DSH 插件:将内核启动页替换为全窗口视频片段,随后渐变过渡到应用。中文说明见 MANUAL.md。DSH plugin: replaces the kernel boot page with a full-window video clip, then dissolves into the app. 中文说明见 MANUAL.md

multimodal
canberk7/ema-lightning
Python · 2026-10-07 多模态 应用 实验 Stars 101 周增 +0

轻量、快速且准确的土耳其语 TTS。860 万参数,在 Freya-TR-Eval 上 WER 为 0.92%,单 GPU 上首段音频延迟约 4 ms,速度达 1.300× 实时。支持音频流式输出、单 GPU 多调用者批处理,并可在 GPU 或 CPU 上离线运行。Tiny, fast and accurate Turkish text-to-speech. 8.6M parameters, 0.92% WER on Freya-TR-Eval, first audio in ~4 ms and 1,300× real time on one GPU. Streams audio, batches many callers on one GPU, and runs offline on a GPU or CPU.

ModelTC/Minimax-H3-Turbo
Python · 2026-08-11 多模态 模型 实验 Stars 91 周增 +0

将 Minimax-H3 蒸馏为 4 步。Distill Minimax-H3 into 4 steps

filliptm/ComfyUI-FL-YuE2
Python · 2026-09-15 多模态 工具 生产可用 Stars 86 周增 +0

YuE2 音乐生成,以及面向 ComfyUI 的可编辑钢琴卷帘界面。YuE2 music generation and an editable piano roll for ComfyUI

Ashlixy17/PCB_lightgraph_Portable
HTML · 2026-08-23 多模态 工具 实验 Stars 81 周增 +0

PCB_lightgraph_portable 是一款点击即可运行的 PCB 图片智能分层与图纸导出的html工具。

ttthanh2044/voxdub
Python · 2026-08-12 多模态 工具 实验 Stars 77 周增 +0
karuvanan/MiniMax-H3-Director-Cut-Studio
Python · 2026-08-29 多模态 应用 实验 Stars 77 周增 +0

受 Premiere 启发、面向 MiniMax H3 Ref2VA 的 PySide6 导演工作室,集成 AI 分镜规划、语义媒体增强、时间线 prompt 调和,以及通过 ComfyUI 实现的镜头感知长视频渲染。Premiere-inspired PySide6 director studio for MiniMax H3 Ref2VA with AI shot planning, semantic media enrichment, timeline prompt reconciliation and shot-aware long-video rendering through ComfyUI.

multimodal
A-Box-of-Tools/website
HTML · 2026-08-24 多模态 应用 实验 Stars 74 周增 +0

https://abox.tools/ 的源码——一个面向图像、视频、音频、PDF 与文本的小型 Web 工具集。文件永不离开本机,因为不存在任何能将它们外发的代码路径。Source for https://abox.tools/ — a box of small web tools for images, video, audio, PDFs and text. Your files never leave your machine, because there is no code path that could send them anywhere.

multimodal
tig3rmast3r/OFXR-Bridge
C++ · 2026-09-07 多模态 工具 实验 Stars 73 周增 +0

基于 Optical Flow 的 VR 帧生成。Frame generation for VR using Optical Flow

adunext/adu-motion-video
JavaScript · 2026-10-03 多模态 应用 实验 Stars 71 周增 +0

阿杜与 Opus 5.5 精选制作的动画动效模板 Skill。选模板和风格,用 Codex / Claude Code 把口播与素材剪成视频;提供 11,028 个精选 Lottie 动画素材 API。