研究库 开源仓库
Repositories · organized/repo_cards

仓库/Skill 库

236 个 · 多模态

排序 Stars 周增
gatewai-dev/framefields
TypeScript · 2026-10-06 多模态 应用 实验 Stars 251 周增 +0

代码优先的视频,原生渲染于 WebGPUCode-first video, rendered natively on WebGPU.

multimodal
Merserk/dlss5-visual-enhancer
Python · 2026-09-02 多模态 应用 实验 Stars 237 周增 +0

DLSS 5 神经视频与图像增强器DLSS 5 Neural Video & Image Enhancer

multimodal
thiagotigaz/ocr-it
JavaScript · 2026-08-25 多模态 工具 生产可用 Stars 234 周增 +175

Chrome 扩展:固定一个屏幕区域后,可用快捷键翻阅分页文档。OCR 通过内置 Tesseract 100% 离线运行。Chrome extension: pin a screen region once, then hotkey your way through a paginated document. OCR runs 100% offline via bundled Tesseract.

VisionVerse/RemoteSensing-Restoration-Survey
Python · 2026-08-09 多模态 应用 实验 Stars 231 周增 +0

[ISPRS 2026] 遥感图像去雾:进展、挑战与前景的系统综述[ISPRS 2026] Remote Sensing Image Dehazing: A Systematic Review of Progress, Challenges, and Prospects

multimodal
EverettFish/holo-card-studio
Python · 2026-09-07 多模态 工具 实验 Stars 231 周增 +0

将用户的描述或上传的参考素材转化为一张成品的、可编辑的 Blender 卡片以及一个交互式 Three.js 页面。保留所要求的主题、风格、字体与目标场景。本 skill 仅包含代码与文本;生成的艺术作品属于用户的输出项目。Turn the user's description or uploaded reference into a finished, editable Blender card and an interactive Three.js page. Preserve the requested subject, style, typography and destination. This skill contains code and text only; generated artwork belongs in the user's output project.

mizorewww/course2md
Rust · 2026-09-03 多模态 工具 实验 Stars 227 周增 +224

将 YouTube、Bilibili 或本地课程/会议录像转换为带幻灯片配图的 Markdown 与 HTML 讲义。Turn YouTube, Bilibili, or local course/meeting recordings into slide-illustrated Markdown and HTML lecture notes.

razorback16/openjev
Python · 2026-09-21 多模态 应用 实验 Stars 225 周增 +0

基于 DiffusionGemma 构建的开放、Jev 兼容的 System One 决策服务器。Open, Jev-compatible System One decision server on DiffusionGemma

programasweights/claudish
Python · 2026-08-27 多模态 工具 实验 Stars 224 周增 +0

通过小型 ProgramAsWeights 函数实现英语与 Claudish 之间的相互翻译Translate between English and Claudish with tiny ProgramAsWeights functions.

jankeesvw/omarchy-meeting-recorder
Rust · 2026-09-25 多模态 应用 实验 Stars 216 周增 +525

在 Omarchy 上录制会议:将麦克风与电脑音频分别录为两条音轨,在本地机器上转录,包含说话人、章节与播放器。Record meetings on Omarchy: mic and computer audio as two tracks, transcribed on your own machine, with speakers, chapters and a player.

opencoredev/bg0
TypeScript · 2026-09-16 多模态 应用 实验 Stars 199 周增 +0

私密、无限制的背景移除,在浏览器中即可运行。Private, unlimited background removal that runs in your browser.

sqzw-x/amane
Python · 2026-08-27 多模态 应用 实验 Stars 197 周增 +0

AI 时代的私人影库

bridge-mind/bridgeclip
TypeScript · 2026-09-25 多模态 应用 实验 Stars 195 周增 +910

BridgeMind 开源的 AI 视频剪辑桌面应用Open-source AI video clipping desktop app by BridgeMind

multimodal
telepath-computer/television
TypeScript · 2026-10-07 多模态 应用 实验 Stars 191 周增 +0

个人 Agent 缺失的 GUI。Television 为你和你的 Agent 提供用于创建和处理工件的可视化空间。The missing GUI for personal agents. Television gives you and your agent a visual space for creating and working with artifacts.

agentmultimodal
LJungang/Awesome-Video-Reasoning-Landscape
Python · 2026-09-17 多模态 收藏榜 实验 Stars 191 周增 +0

🔥一份关于最新视频推理任务、范式与基准的开源综述。🔥An open-source survey of the latest video reasoning tasks, paradigms, and benchmarks.

multimodalevaluationllm-infra
gulelmatthews/Polymarket-Perpetual-Bot
Python · 2026-09-27 多模态 应用 实验 Stars 183 周增 +763

面向 Polymarket 预测市场的交易机器人——浏览 CLOB 市场、在终端查看订单簿、运行套利检测、流动性提供与跨市场套利策略,支持模拟交易和风险限额。教育性开源工具包——不构成投资建议。非官方社区项目,与 Polymarket 无关。Polymarket trading bot for prediction markets — browse CLOB markets, watch the order book in the terminal, run edge detection, liquidity provision and cross-market arbitrage strategies with paper trading and risk limits. Educational open-source toolkit — not financial advice. Unofficial community project, not affiliated with Polymarket.

agentragmultimodalrisk
int64ago/vistep
TypeScript · 2026-09-12 多模态 应用 实验 Stars 183 周增 +0

用 AI 可视化每一步——双语可视化讲解、交互式模型与同步旁白。Visualize Every Step with AI — bilingual visual explanations, interactive models and synchronized narration.

LynnReal-AI/LynnReal-Omni
Python · 2026-09-18 多模态 框架 实验 Stars 182 周增 +0

LynnReal-Omni 将文生视频、图生视频、人体与手部姿态引导生成、结构控制、全模态参考生成、风格迁移、视频编辑、受损视频修复以及流式长视频生成统一到一个框架内,且全部支持四步快速生成。LynnReal-Omni brings text-to-video, image-to-video, human- and hand-pose guided generation, structural control, omni-reference generation, style transfer, video editing, degraded-video restoration and streaming long-video generation into a single framework, all at four-step fast generation.

multimodalevaluation
lihaoyun6/ComfyUI-H3VAE_TRT
Python · 2026-09-06 多模态 工具 实验 Stars 180 周增 +0

在 ComfyUI 中运行 MiniMax-H3 VAE 的 ONNX/TRT 版本,速度最高提升至 1.7 倍。Running ONNX/TRT version of the MiniMax-H3 VAE in ComfyUI, which increase speed by up to 1.7x

guanmo-ai/awesome-ai-motion
JavaScript · 2026-10-02 多模态 收藏榜 实验 Stars 160 周增 +0

精选 AI 动画与视频:作品封面、可播放案例、作者原始提示词与来源。Curated AI motion & video with original prompts. Focused on Claude Opus 5.5.

multimodal
hubertjb/dl-eeg-review
Python · 2020-02-12 多模态 收藏榜 研究原型 Stars 159 周增 +0

深度学习与 EEG 系统文献综述的补充材料Supplementary material for systematic literature review on deep learning and EEG.

Atomicx7/Duo-animation
Kotlin · 2026-09-12 多模态 应用 实验 Stars 159 周增 +0
huangserva/servasyy_skills
Python · 2026-02-05 多模态 工具 实验 Stars 155 周增 +0

AI驱动的多媒体内容生产skills集合:document-writer(写作)、illustration-generator(配图)、ppt-generator(PPT风格)、podcast-generator(TTS)、remoti on-dev(视频制作)、twitter-crawler(推文爬取)、markdown-illustrator(Markdown配图)、comic-generator(漫画生成)、media-downloader(媒体下载)、tts-script-generator(TTS脚本)、md-t o-pdf(文档转换)、wechat-formatter(微信格式化)、humanizer-zh(中文人性化)、shared-lib(核心API库)

AlLHHH/ALH-Pro
C# · 2026-09-05 多模态 工具 生产可用 Stars 155 周增 +0

图片视频等增强工具

blixvip/MotionClone
Python · 2026-09-14 多模态 应用 实验 Stars 152 周增 +0

使用 Codex + ChatGPT 将参考视频转换为可编辑的动态图形;支持对比、定制并导出 MP4 或 HyperFrames 项目;提供本地 Windows 应用与在线工作台。Turn reference videos into editable motion graphics with Codex + ChatGPT. Compare, customize, and export MP4s or HyperFrames projects. Local Windows app + online studio.

agentmultimodal
gillesgoetsch/OmacVM
Shell · 2026-10-07 多模态 应用 实验 Stars 150 周增 +70

在 Mac 上的虚拟机中运行 Omarchy(OmacVM.app、UTM、VMware Fusion 或 Parallels),获得原生般的体验:Omarchy 栏位于刘海旁,支持触控板手势与类 macOS 滚动,复用 Mac 的 Wi-Fi、音频和键盘。一条命令即可构建和切换功能。Omarchy in a VM on your Mac (OmacVM.app, UTM, VMware Fusion or Parallels), feeling native: Omarchy's bar beside the notch, trackpad gestures, macOS-like scrolling, the Mac's Wi-Fi, audio and keys in Omarchy. One command to build and switch features.

liyupi/ai-model-world
TypeScript · 2026-09-20 多模态 模型 实验 Stars 144 周增 +140

AI 大模型世界,把 556 个大模型拟人化成像素小人的可视化站点。进来就能看到此刻谁最聪明、谁最会写代码、谁最便宜、谁刚发布,往下是国内与国外分区的厂商广场、完整的发布时间线和多维排行榜。搜索认模型名、厂商和能力,输入「多模态」会直接列出全部多模态模型。数据取自 Epoch AI、models.dev、LiveBench 与 Hugging Face,每小时自动同步,所有文案由真实数据生成,不调用任何 LLM。Next.js 静态导出,零后端。

evaluationllm-infra
tritant/ComfyUI_MiniMax_H3_Extender
Python · 2026-08-21 多模态 应用 实验 Stars 129 周增 +0

ComfyUI 节点,为 MiniMax H3 而设计,可串联多个视频片段,具备运动上下文、磁盘缓存、动态图像参考、音频参考支持以及最终视频/音频的无缝解码ComfyUI node for MiniMax H3 that chains multiple video clips with motion context, disk caching, dynamic image references, audio reference support, and seamless final video/audio decoding.

multimodal
GENEXIS-AI/gpt-image-skill
JavaScript · 2026-08-28 多模态 工具 生产可用 Stars 126 周增 +0

使用 ChatGPT 订阅在 Codex 或 Claude Code 中生成 GPT 图像,无需 Images API。Generate GPT images from Codex or Claude Code using a ChatGPT subscription, without the Images API.

agentmultimodal
kasturikhanke/generative-loaders
TypeScript · 2026-08-12 多模态 应用 实验 Stars 111 周增 +0

面向生成式接口的无障碍 React 加载状态:流式文本、内联活动指示与图像生成。Accessible React loading states for generative interfaces: streamed text, inline activity, and image generation.

multimodal
facok/comfyui-SelfLift
Python · 2026-09-12 多模态 工具 实验 Stars 106 周增 +124
TarunTomar122/better-voice
Swift · 2026-08-24 多模态 应用 实验 Stars 103 周增 +0

语音听写,同时结合你所指屏幕位置的上下文。Voice dictation with the screen context you point at.

jaredrhod/barehands
HTML · 2026-08-17 多模态 应用 实验 Stars 103 周增 +0

用双手操控屏幕——基于摄像头的免穿戴、免手柄手势追踪界面,让你的 AI 直接响应动作Move things on your screen with your bare hands. A webcam-powered, hand-tracked interface for your AI. No headset. No controllers.

lxj5820/dsh-boot-animation
JavaScript · 2026-10-04 多模态 应用 实验 Stars 102 周增 +0

DSH 插件:将内核启动页替换为全窗口视频片段,随后渐变过渡到应用。中文说明见 MANUAL.md。DSH plugin: replaces the kernel boot page with a full-window video clip, then dissolves into the app. 中文说明见 MANUAL.md

multimodal
canberk7/ema-lightning
Python · 2026-10-07 多模态 应用 实验 Stars 101 周增 +0

轻量、快速且准确的土耳其语 TTS。860 万参数,在 Freya-TR-Eval 上 WER 为 0.92%,单 GPU 上首段音频延迟约 4 ms,速度达 1.300× 实时。支持音频流式输出、单 GPU 多调用者批处理,并可在 GPU 或 CPU 上离线运行。Tiny, fast and accurate Turkish text-to-speech. 8.6M parameters, 0.92% WER on Freya-TR-Eval, first audio in ~4 ms and 1,300× real time on one GPU. Streams audio, batches many callers on one GPU, and runs offline on a GPU or CPU.

synlp/RRG-Review
TeX · 2025-05-17 多模态 收藏榜 研究原型 Stars 100 周增 +0

综述论文"A Systematic Review of Deep Learning-based Research on Radiology Report Generation"的官方 GitHub 仓库The official GitHub repository of the survey paper "A Systematic Review of Deep Learning-based Research on Radiology Report Generation".

ModelTC/Minimax-H3-Turbo
Python · 2026-08-11 多模态 模型 实验 Stars 91 周增 +0

将 Minimax-H3 蒸馏为 4 步。Distill Minimax-H3 into 4 steps