Repositories · organized/repo_cards

仓库/Skill 库

206 个

排序 Stars 周增
hao-ai-lab/FastVideo
Python · 2026-08-11 工程化 框架 研究原型 Stars 3936 周增 +7

用于加速视频生成的统一推理与后训练框架。A unified inference and post-training framework for accelerated video generation.

multimodalllm-infra
s1dashu/ip-as-logo-skill
未知语言 · 2026-08-22 Agent 智能体 应用 研究原型 Stars 3727 周增 +4646

一项紧凑的 Agent Skill,用于生成高度简化、圆润且带有微妙新拟物化风格的 IP 吉祥物 Logo。A compact Agent Skill for highly simplified, rounded, subtly neo-skeuomorphic IP mascot logos.

agentmultimodal
dramaclaw/dramaclaw
TypeScript · 2026-08-11 多模态 应用 研究原型 Stars 3493 周增 -126

A general-purpose AIGC video engine: script to finished film in one pipeline — dramas, ads, product videos, otome games, and more. | 通用 AIGC 视频引擎 —— 从剧本到成片一条流水线,漫剧、广告、电商、乙游皆可

agentmultimodalengineering
breezedeus/Pix2Text
Jupyter Notebook · 2026-02-07 多模态 工具 研究原型 Stars 3213 周增 +7

一款开源 Python3 工具,使用小模型识别图像中的版面、表格、数学公式(LaTeX)和文本,并转换为 Markdown 格式。Mathpix 的免费替代方案,可将视觉内容无缝转换为文本表示,支持 80+ 种语言。An Open-Source Python3 tool with SMALL models for recognizing layouts, tables, math formulas (LaTeX), and text in images, converting them into Markdown format. A free alternative to Mathpix, empowering seamless conversion of visual content into text-based representations. 80+ languages are supported.

multimodalengineering
openvinotoolkit/openvino_notebooks
Jupyter Notebook · 2026-08-10 LLM 基础设施 教程 研究原型 Stars 3195 周增 +0

📚 OpenVINO™ 的 Jupyter notebook 教程。📚 Jupyter notebook tutorials for OpenVINO™

multimodalllm-infra
HiLab-git/SSL4MIS
Python · 2025-06-07 多模态 收藏榜 研究原型 Stars 2671 周增 +0

半监督学习在医学图像分割中的应用,含文献综述与代码实现合集Semi Supervised Learning for Medical Image Segmentation, a collection of literature reviews and code implementations.

multimodal
genieincodebottle/generative-ai
Jupyter Notebook · 2026-08-05 工程化 收藏榜 研究原型 Stars 2593 周增 +14

生成式AI综合资源,包含详细路线图、项目、用例、面试准备与编码练习Comprehensive resources on Generative AI, including a detailed roadmap, projects, use cases, interview preparation, and coding preparation.

agentragmultimodalevaluation
roboflow/inference
Python · 2026-08-10 LLM 基础设施 应用 研究原型 Stars 2411 周增 +7

将任意电脑或边缘设备打造为计算机视觉项目的指挥中心Turn any computer or edge device into a command center for your computer vision projects.

agentmultimodalllm-infraengineering
aigc-apps/EasyAnimate
Python · 2025-03-06 多模态 应用 实验 Stars 2266 周增 +0

📺 基于 Transformer Diffusion 的高分辨率长视频生成端到端解决方案。📺 An End-to-End Solution for High-Resolution and Long Video Generation Based on Transformer Diffusion

multimodal
aigc-apps/VideoX-Fun
Python · 2026-08-10 多模态 框架 生产可用 Stars 2192 周增 +0

📹 一个更灵活的框架,支持任意分辨率视频生成以及从图像生成视频。📹 A more flexible framework that can generate videos at any resolution and creates videos from images.

multimodal
kingyiusuen/image-to-latex
Python · 2022-10-04 多模态 工具 实验 Stars 2160 周增 +0

将 LaTeX 数学公式的图像转换为 LaTeX 代码。Convert images of LaTex math equations into LaTex code.

multimodalengineering
NVIDIA-AI-Blueprints/video-search-and-summarization
C++ · 2026-08-11 多模态 应用 研究原型 Stars 1790 周增 +7

NVIDIA AI Blueprint for video search and summarization(VSS)是一个 GPU 加速参考架构,用于构建具备实时验证告警、视觉问答与自动报告能力的视频分析 Agent。VSS Blueprint 采用 NVIDIA Cosmos 等视觉语言模型(VLM)、NVIDIA Nemotron 等 LLM,并结合 RAG 与 NVIDIA NIM。NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts, visual Q&A, and automated reporting. The VSS Blueprint uses vision language models (VLMs) such as NVIDIA Cosmos, LLMs such as NVIDIA Nemotron, RAG, and NVIDIA NIMs.

agentragmultimodalllm-infra
hymie122/RAG-Survey
未知语言 · 2024-08-20 RAG 检索增强 收藏榜 实验 Stars 1790 周增 +0

收录 AIGC 领域 RAG 的优秀论文。我们在论文 "Retrieval-Augmented Generation for AI-Generated Content: A Survey" 中提出了 RAG 基础、增强与应用分类法。Collecting awesome papers of RAG for AIGC. We propose a taxonomy of RAG foundations, enhancements, and applications in paper "Retrieval-Augmented Generation for AI-Generated Content: A Survey".

ragmultimodalllm-infra
jau123/MeiGen-AI-Design-MCP
TypeScript · 2026-08-05 Agent 智能体 研究原型 Stars 1689 周增 +14

支持GPT Image 2、Seedance与ComfyUI,配备1,400+提示词库、精心设计的hooks以及多任务编排系统Supports GPT Image 2, Seedance & ComfyUI, with a 1,400+ prompt library, carefully crafted hooks and a multi-task orchestration system

multimodal
HITsz-TMG/VideoClaw
Python · 2026-07-17 多模态 框架 研究原型 Stars 1685 周增 +21

🚀 AI 全自动化视频生成员工 | Your First AIGC Coworker. Chat an Idea. Get a Film. 🦞

agentmultimodal
Devin-AXIS/iPolloWork
TypeScript · 2026-07-20 Agent 智能体 应用 研究原型 Stars 1452 周增 +266

下一代源码可用的 Codex 与 Claude Code 替代品——一个本地优先、可自托管的 Agent 工作空间,覆盖代码、办公文档、可编辑设计、演示文稿、网站与视频。用 AI 构建后,可像 PowerPoint 一样轻松编辑文字、图片、配色、布局与场景。Next-gen, source-available alternative to Codex and Claude Code — one local-first, self-hostable agent workspace for code, office work, editable design, presentations, websites, and video. Build with AI, then edit text, images, colors, layouts, and scenes as easily as PowerPoint.

agentmultimodalllm-infra
waybarrios/vllm-mlx
Python · 2026-06-28 多模态 工具 研究原型 Stars 1447 周增 +7

兼容 OpenAI 和 Anthropic 协议的 Apple Silicon 服务端。可运行 LLM 和视觉语言模型(Llama、Qwen-VL、LLaVA),支持 continuous batching、MCP 工具调用与多模态。原生 MLX 后端,速度达 400+ tok/s,兼容 Claude Code。OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous batching, MCP tool calling, and multimodal support. Native MLX backend, 400+ tok/s. Works with Claude Code.

multimodalllm-infra
kimsungwhee/apple-docs-mcp
TypeScript · 2026-03-17 Agent 智能体 教程 研究原型 Stars 1347 周增 +7

面向Apple开发者文档的MCP服务器——在Claude、Cursor及AI助手中检索iOS/macOS/SwiftUI/UIKit文档、WWDC视频、Swift/Objective-C API及代码示例MCP server for Apple Developer Documentation - Search iOS/macOS/SwiftUI/UIKit docs, WWDC videos, Swift/Objective-C APIs & code examples in Claude, Cursor & AI assistants

multimodal
liyue-aigc/female-portrait-director
未知语言 · 2026-07-15 多模态 工具 实验 Stars 1336 周增 +21

用于引导和扩展详细 AI 女性肖像 prompt 的模块化 Codex Skill。A modular Codex Skill for directing and expanding detailed AI female portrait prompts.

multimodal
yzfly/douyin-mcp-server
HTML · 2026-07-02 Agent 智能体 应用 研究原型 Stars 1246 周增 +0

提取抖音无水印视频链接,视频文案,douyin-mcp-server,mcp,claude skill,支持龙虾

agentmultimodalllm-infra
gyoridavid/short-video-maker
TypeScript · 2025-06-21 Agent 智能体 模型 实验 Stars 1225 周增 +0

使用 Model Context Protocol (MCP) 和 REST API 为 TikTok、Instagram Reels 和 YouTube Shorts 创建短视频。Creates short videos for TikTok, Instagram Reels, and YouTube Shorts using the Model Context Protocol (MCP) and a REST API.

multimodal
all-in-aigc/sorafm
TypeScript · 2024-08-15 多模态 应用 实验 Stars 1152 周增 +14

Sora.FM 推出的 Sora AI 视频生成器。Sora AI Video Generator by Sora.FM

multimodal
FutureUniant/Tailor
Python · 2025-06-03 多模态 工具 生产可用 Stars 1107 周增 +0

Tailor是一款视频智能裁剪、视频生成和视频优化的视频剪辑工具。目前的目标是通过人工智能技术减少视频剪辑的繁琐操作,让普通人也能简单实现专业剪辑人的水准!长远目标是让视频剪辑实现真正的AIGC!

multimodal
ATH-MaaS/Pixelle-MCP
Python · 2025-12-17 Agent 智能体 应用 实验 Stars 1100 周增 +7

基于 ComfyUI + MCP + LLM 的开源多模态 AIGC 解决方案,https://pixelle.aiAn Open-Source Multimodal AIGC Solution based on ComfyUI + MCP + LLM https://pixelle.ai

multimodalllm-infra
MAC-AutoML/MindPipe
Python · 2026-08-11 LLM 基础设施 框架 研究原型 Stars 1013 周增 +0

面向 LLM 和 LVLM 的强大模型压缩框架,适配 NVIDIA GPU 和华为昇腾 NPUA powerful model compression framework for LLMs and LVLMs, adapted for NVIDIA GPUs and Huawei Ascend NPUs.

multimodalllm-infraevaluationengineering
RevoltDevScript/Revolt-Script
未知语言 · 2026-07-20 安全与风险 工具 实验 Stars 887 周增 +21

Revolt 是排名第一的 Edgenuity 自动化工具,提供自动答题、自动写作文、自动跳课、humanization 驱动的 AI 写作、自动词汇、自动日志等功能。这是一款半挂机脚本,可轻松应对测验、考试、测试、项目和作文。支持桌面端和移动端。访问 https://revolt.ly 即可开始使用。Revolt is the #1 Edgenuity automation tool featuring Auto Quiz, Auto Essay, Auto Advance, AI-powered writing with humanization, Auto Vocabulary, Auto Journal, and more. A semi-AFK script that handles quizzes, tests, exams, projects, and essays with ease. Compatible with desktop and mobile. Visit https://revolt.ly to get started.

multimodal
ysr666/dsh-vision-router
JavaScript · 2026-08-19 多模态 工具 实验 Stars 841 周增 +763

为纯文本 DeepSeek Harness Agent 装上眼睛:内置免费视觉链路(无需 key)+ 像素级视觉工具(问答、定位、裁剪、像素 diff、取色、OCR、SVG 矢量化、抠图、截图)。一条命令安装,无需 Python,图像回合与普通工具调用回合一致。Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.

agentmultimodal
coderonion/awesome-llm-and-aigc
未知语言 · 2025-08-01 多模态 收藏榜 实验 Stars 811 周增 +0

🚀🚀🚀 收录关于 LLM、VLM、VLA、AIGC 及相关数据集与应用的一些优秀开源项目合集。🚀🚀🚀A collection of some awesome public projects about Large Language Model(LLM), Vision Language Model(VLM), Vision Language Action(VLA), AI Generated Content(AIGC), the related Datasets and Applications.

multimodalllm-infra
lidge-jun/ima2-gen
TypeScript · 2026-08-14 Agent 智能体 应用 研究原型 Stars 675 周增 +0

本地优先的视觉生成运行时与工作台,面向人类与编程 agent,跨多提供商提供可复现的图像与视频工作流。Local-first visual generation runtime and studio for people and coding agents, with reproducible image and video workflows across multiple providers.

agentmultimodal
Kobaayyy/Awesome-CVPR2026-CVPR2025-ICCV2025-CVPR2024-ECCV2026-ECCV2024-AIGC
未知语言 · 2026-08-05 多模态 收藏榜 研究原型 Stars 674 周增 +0

CVPR2026/CVPR2025/ICCV2025/CVPR2024/ECCV2026/ECCV2024 AIGC 方向的论文与代码合集。A Collection of Papers and Codes for CVPR2026/CVPR2025/ICCV2025/CVPR2024/ECCV2026/ECCV2024 AIGC

multimodalllm-infra
henrydaum/second-brain
Python · 2026-08-15 Agent 智能体 框架 研究原型 Stars 599 周增 +0

Second Brain 是一个 Agentic 框架,作为操作系统运行,利用本地文件智能、工作流自动化和 LLM 来完成任务,并通过多种模态和消息平台进行通信。Second Brain is an agentic framework that acts as an operating system, using local file intelligence, workflow automation, and LLMs to complete tasks and communicate over multiple modalities and messaging platforms.

agentragmultimodaldatabase
Moonlit-Pages/AIGC-Detector-Rewriter-Skill
未知语言 · 2026-06-19 多模态 应用 研究原型 Stars 488 周增 +0

一个面向英文学术写作的保守型 AIGC 检测器指导的论文改写 Skill。支持 Turnitin AI、CNKI AIGC、最小化编辑修订、保留学术要素、定性/定量路由,以及逐章降低 AI 写作风险,且不宣称绕过检测器。A conservative AIGC detector-informed thesis rewriting skill for English and Chinese academic writing. Supports Turnitin AI, CNKI AIGC, minimal-edit revision, protected academic elements, qualitative/quantitative routing, and chapter-by-chapter AI-writing risk reduction without detector-bypass claims.

multimodalriskengineering
alonw0/web-asset-generator
Python · 2026-01-28 多模态 工具 实验 Stars 482 周增 +0

Claude skill,可基于 logo、文字或 emoji 生成 favicon、应用图标与社交媒体图片。支持 emoji 推荐、校验与框架自动集成。Claude skill to generate favicons, app icons, and social media images from logos, text, or emojis. Supports emoji suggestions, validation, and framework auto-integration.

multimodal
lycohana/BiliSum
Python · 2026-08-15 多模态 研究原型 Stars 471 周增 +0

为 Bilibili、YouTube 及本地视频提供 AI 视频摘要和知识库.AI video summarizer and knowledge base for Bilibili, YouTube and local videos.

ragmultimodal
westlake-repl/Recommendation-Systems-without-Explicit-ID-Features-A-Literature-Review
未知语言 · 2024-08-12 工程化 收藏榜 研究原型 Stars 368 周增 +0

预训练基础推荐模型论文清单Paper List of Pre-trained Foundation Recommender Models

multimodalllm-infra
zyrexdz/cyberleek-leak-research
未知语言 · 2026-08-20 安全与风险 工具 实验 Stars 318 周增 +0

全面解析 2026 年 8 月 GTA 6 玩法泄露事件:技术核查、暗网溯源、地图细节与诈骗警示Full breakdown of the August 2026 GTA 6 gameplay leaks, technical checks, dark web trail, map details, and scam warning.

multimodal