研究库 开源仓库
Repositories · organized/repo_cards

仓库/Skill 库

281 个

排序 Stars 周增
LynnReal-AI/LynnReal-Omni
Python · 2026-09-18 多模态 框架 实验 Stars 182 周增 +0

LynnReal-Omni 将文生视频、图生视频、人体与手部姿态引导生成、结构控制、全模态参考生成、风格迁移、视频编辑、受损视频修复以及流式长视频生成统一到一个框架内,且全部支持四步快速生成。LynnReal-Omni brings text-to-video, image-to-video, human- and hand-pose guided generation, structural control, omni-reference generation, style transfer, video editing, degraded-video restoration and streaming long-video generation into a single framework, all at four-step fast generation.

multimodalevaluation
vericontext/vibeframe
TypeScript · 2026-10-04 Agent 智能体 工具 实验 Stars 173 周增 +0

前沿 AI 视频生成用于编程 Agent — Seedance、Runway、Veo、Kling 使用你自己的 API key,受硬性成本上限控制。CLI + MCP。Frontier AI video generation for coding agents - Seedance, Runway, Veo, Kling on your own keys, behind a hard cost cap. CLI + MCP.

agentmultimodal
kodelyx/flow-agent
JavaScript · 2026-09-25 Agent 智能体 工具 实验 Stars 172 周增 +0

⚡ Google Flow 的 CLI 工具包——Nano Banana Pro 图像、Omni Flash 视频、MCP v2 与 OpenAI API。⚡ CLI toolkit for Google Flow — Nano Banana Pro images, Omni Flash videos, MCP v2 & OpenAI API.

multimodal
codecentric/c4-genai-suite
TypeScript · 2026-08-20 工程化 工具 实验 Stars 172 周增 +0

c4 GenAI Suitec4 GenAI Suite

agentragmultimodalllm-infra
sisinflab/adversarial-recommender-systems-survey
未知语言 · 2021-03-03 安全与风险 收藏榜 研究原型 Stars 166 周增 +0

本综述目标有二:(i) 综述对抗机器学习(AML)在推荐系统(RS)安全上的最新进展,即攻击与防御推荐模型;(ii) 展示 AML 在生成对抗网络(GAN)生成应用中的成功应用,得益于其学习(高维)数据分布的能力。本文对发表于主流 RS 与 ML 期刊会议的 74 篇文献进行了详尽综述,可作为 RS 社区在推荐系统安全及利用 GAN 提升生成模型质量方面的参考The goal of this survey is two-fold: (i) to present recent advances on adversarial machine learning (AML) for the security of RS (i.e., attacking and defense recommendation models), (ii) to show another successful application of AML in generative adversarial networks (GANs) for generative applications, thanks to their ability for learning (high-dimensional) data distributions. In this survey, we provide an exhaustive literature review of 74 articles published in major RS and ML journals and conferences. This review serves as a reference for the RS community, working on the security of RS or on generative models using GANs to improve their quality.

multimodalrisk
guanmo-ai/awesome-ai-motion
JavaScript · 2026-10-02 多模态 收藏榜 实验 Stars 160 周增 +0

精选 AI 动画与视频:作品封面、可播放案例、作者原始提示词与来源。Curated AI motion & video with original prompts. Focused on Claude Opus 5.5.

multimodal
elysia395/dsh-wallpaper-engine
JavaScript · 2026-08-23 工程化 应用 实验 Stars 159 周增 +150

把本机 Wallpaper Engine 的壁纸变成 DSH 网页界面的背景:Video 动态播放、Web 以 iframe 加载、Scene 壁纸提取主纹理作为静态帧;iOS 液态玻璃设置窗口(配色 / 玻璃颜色 / 透明度)、内容分级与类型过滤、自定义壁纸上传、紧凑 CD 架布局、黑胶唱片展示、隐藏 / 恢复、倍速 / 翻转与自动轮播。感谢 Jerry 维护 macOS 版。

multimodal
blixvip/MotionClone
Python · 2026-09-14 多模态 应用 实验 Stars 152 周增 +0

使用 Codex + ChatGPT 将参考视频转换为可编辑的动态图形;支持对比、定制并导出 MP4 或 HyperFrames 项目;提供本地 Windows 应用与在线工作台。Turn reference videos into editable motion graphics with Codex + ChatGPT. Compare, customize, and export MP4s or HyperFrames projects. Local Windows app + online studio.

agentmultimodal
xiayu1987/noobot
JavaScript · 2026-09-13 Agent 智能体 模型 实验 Stars 150 周增 +0

最省钱的自托管 AI Agent 工作空间,集成 tool calling、MCP、多模型路由、沙箱执行、多 Agent 工作流,以及由 LLM 编写并通过 Three.js 渲染的 3D 角色动画。Cheapest Money-Saving Self-hosted AI agent workspace with tool calling, MCP, multi-model routing, sandboxed execution, multi-agent workflows, and LLM-authored 3D character animation rendered with Three.js.

agentmultimodalllm-infra
Hebbian-Robotics/hflow
Python · 2026-08-25 工程化 库 实验 Stars 142 周增 +98

用于机器人和 Physical AI 的多模态数据质量、处理、enrichment 与 curation pipelines 的开源 SDK。Open source SDK for building multimodal data-quality, processing, enrichment, and curation pipelines for robotics and Physical AI.

multimodalengineering
tritant/ComfyUI_MiniMax_H3_Extender
Python · 2026-08-21 多模态 应用 实验 Stars 129 周增 +0

ComfyUI 节点,为 MiniMax H3 而设计,可串联多个视频片段,具备运动上下文、磁盘缓存、动态图像参考、音频参考支持以及最终视频/音频的无缝解码ComfyUI node for MiniMax H3 that chains multiple video clips with motion context, disk caching, dynamic image references, audio reference support, and seamless final video/audio decoding.

multimodal
SysAdminDoc/HushFacebook
Java · 2026-10-01 工程化 工具 实验 Stars 126 周增 +98

Hushfacebook v0.5.0: 针对 Facebook 580.0.0.51.74 的 51 个补丁,配合 Morphe 使用。可隐藏广告与推荐帖、保存视频、调节播放与 Reels,并提供专注的 Marketplace 模式。Hushfacebook v0.5.0: 51 patches for Facebook 580.0.0.51.74 with Morphe. Hide ads and suggested posts, save videos, tune playback and Reels, and use a focused Marketplace mode.

multimodal
GENEXIS-AI/gpt-image-skill
JavaScript · 2026-08-28 多模态 工具 生产可用 Stars 126 周增 +0

使用 ChatGPT 订阅在 Codex 或 Claude Code 中生成 GPT 图像,无需 Images API。Generate GPT images from Codex or Claude Code using a ChatGPT subscription, without the Images API.

agentmultimodal
MysticGSI/mysticgsi
Shell · 2026-09-26 工程化 工具 生产可用 Stars 125 周增 +0

用于从(几乎)任意原厂固件构建通用系统镜像(GSI)的工具。A tool to build a Generic System Image from (almost) any stock firmware

multimodal
kasturikhanke/generative-loaders
TypeScript · 2026-08-12 多模态 应用 实验 Stars 111 周增 +0

面向生成式接口的无障碍 React 加载状态:流式文本、内联活动指示与图像生成。Accessible React loading states for generative interfaces: streamed text, inline activity, and image generation.

multimodal
kunpengtalk/OmniStudio
TypeScript · 2026-09-16 LLM 基础设施 应用 实验 Stars 108 周增 +114

OmniStudio 是一个本地大模型一体化桌面工作台,集模型市集下载、llama.cpp/vLLM/SGLang 三引擎推理管理,以及对话、语音合成、ASR语音识别、图片生成、视频生成、OCR 等多种大模型应用于一体,全程本地优先。

multimodalllm-infra
lxj5820/dsh-boot-animation
JavaScript · 2026-10-04 多模态 应用 实验 Stars 102 周增 +0

DSH 插件:将内核启动页替换为全窗口视频片段,随后渐变过渡到应用。中文说明见 MANUAL.md。DSH plugin: replaces the kernel boot page with a full-window video clip, then dissolves into the app. 中文说明见 MANUAL.md

multimodal
howseen-ai/claude-motion-design
Python · 2026-09-29 Agent 智能体 工具 实验 Stars 96 周增 +0

一项 Claude Code skill,可通过纯代码制作动态设计视频:HTML + Playwright + ffmpeg。无需 After Effects,无需 Remotion 授权。A Claude Code skill to make motion design videos in pure code: HTML + Playwright + ffmpeg. No After Effects, no Remotion license.

multimodal
WenyuChiou/academic-writing-skills
Python · 2026-10-09 多模态 应用 实验 Stars 88 周增 +14

面向严谨学术论文写作、修订与投稿的 Claude Code skill。跨领域通用,支持按论文设置期刊覆盖规则。Claude Code skill for rigorous academic paper writing, revision, and submission. Field-agnostic with per-paper journal overrides.

multimodal
Bike4Mind/bike4mind
TypeScript · 2026-10-09 Agent 智能体 模型 实验 Stars 88 周增 +0

开源核心的 AI 工作台——覆盖任何模型的 notebooks、Agent、RAG、语音和图像:OpenAI、Anthropic、Google、xAI,或通过 Ollama/vLLM 的本地模型。BSL 1.1,两年后自动转换为 Apache-2.0。当别人的 AI 停止运行时,你的 AI 仍在运行。The open-core AI workbench — notebooks, agents, RAG, voice, and images across any model: OpenAI, Anthropic, Google, xAI, or local via Ollama/vLLM. BSL 1.1, auto-converting to Apache-2.0 on a two-year clock. Your AI keeps running when theirs doesn't.

agentragmultimodalllm-infra
leeguooooo/iphone-use
Rust · 2026-10-07 Agent 智能体 工具 实验 Stars 83 周增 +0

开源的 AI Agent 真实 iPhone 控制方案:屏幕文本、点击/输入/滑动、自定义 XCTest runner、CLI + HTTP API + MCP、浏览器接管与可回放流程。Rust 编写,自托管于 macOS。Open-source real iPhone control for AI agents: screen text, tap/type/swipe, custom XCTest runner, CLI + HTTP API + MCP, browser takeover and replayable flows. Rust, self-hosted on macOS.

agentragmultimodal
a-r-d/PureJsImage
TypeScript · 2026-08-11 工程化 库 实验 Stars 82 周增 +0

面向 Node.js、浏览器与 serverless 运行时的纯 TypeScript 低内存、零依赖图像编解码与处理库。具备有界解码/缩放管线,并提供深度 TIFF、科学栅格、GeoTIFF、OME-TIFF 与全切片成像支持。Low-memory, zero-dependency image codecs and processing in pure TypeScript for Node.js, browsers, and serverless runtimes. Bounded decode/resize pipelines plus deep TIFF, scientific raster, GeoTIFF, OME-TIFF, and whole-slide support.

multimodalengineering
artokun/comfyui-mcp-panel
JavaScript · 2026-08-12 Agent 智能体 应用 实验 Stars 79 周增 +0

面向 ComfyUI 的本地优先侧边栏 AI Agent——运行在你自己的 Claude 或 ChatGPT 订阅上(无需 API key,不产生额外 LLM 成本)。通过自然语言驱动你的实时图:编辑、工作流与安装。作为 comfyui-mcp 的面板 UI,comfyui-mcp 是 ComfyUI 的 Agent 原生控制平面。The local-first sidebar AI agent for ComfyUI — runs on your own Claude OR ChatGPT subscription (no API keys, no extra LLM costs). Drives your live graph: edits, workflows & installs in natural language. The panel UI for comfyui-mcp, the agent-native control plane for ComfyUI.

agentmultimodalllm-infra
stabgan/openrouter-mcp-multimodal
TypeScript · 2026-08-22 多模态 应用 实验 Stars 78 周增 +0

OpenRouter 的 MCP server——与 300+ LLM(Claude、Gemini、GPT)对话,分析图像/音频/视频,生成图像/语音/音乐/视频(Veo 3.1、Sora、Seedance、Wan),支持 Claude Desktop、Cursor、Kiro、VS CodeMCP server for OpenRouter — chat with 300+ LLMs (Claude, Gemini, GPT), analyze images / audio / video, generate images / speech / music / video (Veo 3.1, Sora, Seedance, Wan) from Claude Desktop, Cursor, Kiro, VS Code.

agentmultimodalllm-infra
Totoro-qaq/dsh-plugin-bridge
JavaScript · 2026-08-22 评测基准 模型 实验 Stars 77 周增 +0

DeepSeek Harness 插件,支持跨预设会话迁移的可预览;固定 schema 的交接保留状态、源模型意图与未处理图像,原始会话保持不变DeepSeek Harness plugin for previewable cross-preset session migration. Fixed-schema handoffs preserve state, source-model intent, and unresolved images; the original session stays untouched.

multimodal
karuvanan/MiniMax-H3-Director-Cut-Studio
Python · 2026-08-29 多模态 应用 实验 Stars 77 周增 +0

受 Premiere 启发、面向 MiniMax H3 Ref2VA 的 PySide6 导演工作室,集成 AI 分镜规划、语义媒体增强、时间线 prompt 调和,以及通过 ComfyUI 实现的镜头感知长视频渲染。Premiere-inspired PySide6 director studio for MiniMax H3 Ref2VA with AI shot planning, semantic media enrichment, timeline prompt reconciliation and shot-aware long-video rendering through ComfyUI.

multimodal
A-Box-of-Tools/website
HTML · 2026-08-24 多模态 应用 实验 Stars 74 周增 +0

https://abox.tools/ 的源码——一个面向图像、视频、音频、PDF 与文本的小型 Web 工具集。文件永不离开本机,因为不存在任何能将它们外发的代码路径。Source for https://abox.tools/ — a box of small web tools for images, video, audio, PDFs and text. Your files never leave your machine, because there is no code path that could send them anywhere.

multimodal
tonyd2wild/GLM-5.3-Flash-NVFP4-DFlash2-2x-DGX-Spark
Python · 2026-08-28 LLM 基础设施 应用 实验 Stars 69 周增 +0

2× NVIDIA DGX Spark 上的 GLM-5.3-Flash (NVFP4)——vLLM TP2,262K 上下文,MTP。全球首发部署方案:发现并修复 7 个 day-0 bug,附 sm121 镜像补丁、探针与完整报告。GLM-5.3-Flash (NVFP4) on 2x NVIDIA DGX Spark - vLLM TP2, 262K context, MTP. World-first deploy recipe: 7 day-0 bugs found and fixed, patched sm121 image, probes and full report.

multimodalllm-infra
genieincodebottle/aiml-companion
Jupyter Notebook · 2026-09-07 Agent 智能体 框架 实验 Stars 69 周增 +0

AI-ML Companion:一个通过"看 AI 运作"来学习 AI 与 ML 的交互式平台 — 28 条学习路径、400+ 模块、实时可视化 — 另有 20+ 端到端真实项目,覆盖经典 ML、深度学习、计算机视觉到生产级 LLM 与多 Agent 系统。AI-ML Companion: an interactive platform to learn AI & ML by watching it work - 28 tracks, 400+ modules, live visualizations - plus 20+ end-to-end real world projects from classical ML, DL, Computer Vision to production LLM & multi-agent systems.

agentragmultimodalllm-infra
yxxbc/downvid
TypeScript · 2026-08-24 工程化 工具 生产可用 Stars 67 周增 +0

现代化的开源视频下载工具 — 一键下载全网无水印高清视频

multimodal
Agents365-ai/365-skills
Python · 2026-10-03 Agent 智能体 应用 实验 Stars 66 周增 +63

日常使用的 Agent skills。Agent skills for daily use

agentmultimodal
OneMana-Soft/OneCamp-fe
TypeScript · 2026-09-30 多模态 应用 实验 Stars 66 周增 +0

OneCamp:自托管一体化工作空间(聊天、任务、视频通话、文档、日历与 AI)。OneCamp: self-hosted all-in-one workspace (chat, tasks, video calls, docs, calendar & AI)

agentmultimodal
Vaquill-AI/open-india-law
TypeScript · 2026-08-26 多模态 应用 实验 Stars 65 周增 +0

开放、结构化的印度一手法律数据:包含最高法院与全部 25 所高等法院的 3250 万判决分块、110 万条立法条款,以及构建数据集的爬虫;采用 CC BY 4.0 许可。Open, structured Indian primary law: 32.5M judgment chunks from the Supreme Court and all 25 High Courts, 1.1M legislation provisions, and the scrapers that build it. CC BY 4.0.

multimodal
luobosibing2/dsh-jev-plugin
JavaScript · 2026-10-02 Agent 智能体 应用 实验 Stars 63 周增 +0

原生 DeepSeek Harness (DSH) 插件,将 TypeSafe Jev 集成为 System One 决策层,用于 Agent 选择、监督、纠错与审批。Native DeepSeek Harness (DSH) plugin integrating TypeSafe Jev as a System One decision layer for agent selection, supervision, corrections, and approvals.

agentmultimodal
ArtemPavlov1994/polymarket-prediction-bot
Python · 2026-09-24 多模态 应用 实验 Stars 62 周增 +7

面向 Polymarket 预测市场的交易机器人——可在终端浏览 CLOB 市场、查看订单簿,并执行 edge detection、流动性提供与跨市场套利策略,支持 paper trading 与风险限额。开源教育工具——不构成投资建议。非官方社区项目,与 Polymarket 无关。Polymarket trading bot for prediction markets — browse CLOB markets, watch the order book in the terminal, run edge detection, liquidity provision and cross-market arbitrage strategies with paper trading and risk limits. Educational open-source toolkit — not financial advice. Unofficial community project, not affiliated with Polymarket.

agentragmultimodalrisk
JGRFW/comfyui-AICG3D
JavaScript · 2026-09-17 多模态 工具 实验 Stars 60 周增 +0

把 ComfyUI-MiniMaxH3-Easy 与 Goohai-MiniMax-H3_Integration 合并成一个统一的 MiniMax H3 创作工作台(非官方)

multimodal