用于加速视频生成的统一推理与后训练框架。A unified inference and post-training framework for accelerated video generation.
仓库/Skill 库
206 个
一项紧凑的 Agent Skill,用于生成高度简化、圆润且带有微妙新拟物化风格的 IP 吉祥物 Logo。A compact Agent Skill for highly simplified, rounded, subtly neo-skeuomorphic IP mascot logos.
A general-purpose AIGC video engine: script to finished film in one pipeline — dramas, ads, product videos, otome games, and more. | 通用 AIGC 视频引擎 —— 从剧本到成片一条流水线,漫剧、广告、电商、乙游皆可
一款开源 Python3 工具,使用小模型识别图像中的版面、表格、数学公式(LaTeX)和文本,并转换为 Markdown 格式。Mathpix 的免费替代方案,可将视觉内容无缝转换为文本表示,支持 80+ 种语言。An Open-Source Python3 tool with SMALL models for recognizing layouts, tables, math formulas (LaTeX), and text in images, converting them into Markdown format. A free alternative to Mathpix, empowering seamless conversion of visual content into text-based representations. 80+ languages are supported.
📚 OpenVINO™ 的 Jupyter notebook 教程。📚 Jupyter notebook tutorials for OpenVINO™
半监督学习在医学图像分割中的应用,含文献综述与代码实现合集Semi Supervised Learning for Medical Image Segmentation, a collection of literature reviews and code implementations.
生成式AI综合资源,包含详细路线图、项目、用例、面试准备与编码练习Comprehensive resources on Generative AI, including a detailed roadmap, projects, use cases, interview preparation, and coding preparation.
将任意电脑或边缘设备打造为计算机视觉项目的指挥中心Turn any computer or edge device into a command center for your computer vision projects.
📺 基于 Transformer Diffusion 的高分辨率长视频生成端到端解决方案。📺 An End-to-End Solution for High-Resolution and Long Video Generation Based on Transformer Diffusion
📹 一个更灵活的框架,支持任意分辨率视频生成以及从图像生成视频。📹 A more flexible framework that can generate videos at any resolution and creates videos from images.
将 LaTeX 数学公式的图像转换为 LaTeX 代码。Convert images of LaTex math equations into LaTex code.
NVIDIA AI Blueprint for video search and summarization(VSS)是一个 GPU 加速参考架构,用于构建具备实时验证告警、视觉问答与自动报告能力的视频分析 Agent。VSS Blueprint 采用 NVIDIA Cosmos 等视觉语言模型(VLM)、NVIDIA Nemotron 等 LLM,并结合 RAG 与 NVIDIA NIM。NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts, visual Q&A, and automated reporting. The VSS Blueprint uses vision language models (VLMs) such as NVIDIA Cosmos, LLMs such as NVIDIA Nemotron, RAG, and NVIDIA NIMs.
收录 AIGC 领域 RAG 的优秀论文。我们在论文 "Retrieval-Augmented Generation for AI-Generated Content: A Survey" 中提出了 RAG 基础、增强与应用分类法。Collecting awesome papers of RAG for AIGC. We propose a taxonomy of RAG foundations, enhancements, and applications in paper "Retrieval-Augmented Generation for AI-Generated Content: A Survey".
支持GPT Image 2、Seedance与ComfyUI,配备1,400+提示词库、精心设计的hooks以及多任务编排系统Supports GPT Image 2, Seedance & ComfyUI, with a 1,400+ prompt library, carefully crafted hooks and a multi-task orchestration system
🚀 AI 全自动化视频生成员工 | Your First AIGC Coworker. Chat an Idea. Get a Film. 🦞
下一代源码可用的 Codex 与 Claude Code 替代品——一个本地优先、可自托管的 Agent 工作空间,覆盖代码、办公文档、可编辑设计、演示文稿、网站与视频。用 AI 构建后,可像 PowerPoint 一样轻松编辑文字、图片、配色、布局与场景。Next-gen, source-available alternative to Codex and Claude Code — one local-first, self-hostable agent workspace for code, office work, editable design, presentations, websites, and video. Build with AI, then edit text, images, colors, layouts, and scenes as easily as PowerPoint.
兼容 OpenAI 和 Anthropic 协议的 Apple Silicon 服务端。可运行 LLM 和视觉语言模型(Llama、Qwen-VL、LLaVA),支持 continuous batching、MCP 工具调用与多模态。原生 MLX 后端,速度达 400+ tok/s,兼容 Claude Code。OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous batching, MCP tool calling, and multimodal support. Native MLX backend, 400+ tok/s. Works with Claude Code.
面向Apple开发者文档的MCP服务器——在Claude、Cursor及AI助手中检索iOS/macOS/SwiftUI/UIKit文档、WWDC视频、Swift/Objective-C API及代码示例MCP server for Apple Developer Documentation - Search iOS/macOS/SwiftUI/UIKit docs, WWDC videos, Swift/Objective-C APIs & code examples in Claude, Cursor & AI assistants
用于引导和扩展详细 AI 女性肖像 prompt 的模块化 Codex Skill。A modular Codex Skill for directing and expanding detailed AI female portrait prompts.
提取抖音无水印视频链接,视频文案,douyin-mcp-server,mcp,claude skill,支持龙虾
使用 Model Context Protocol (MCP) 和 REST API 为 TikTok、Instagram Reels 和 YouTube Shorts 创建短视频。Creates short videos for TikTok, Instagram Reels, and YouTube Shorts using the Model Context Protocol (MCP) and a REST API.
Tailor是一款视频智能裁剪、视频生成和视频优化的视频剪辑工具。目前的目标是通过人工智能技术减少视频剪辑的繁琐操作,让普通人也能简单实现专业剪辑人的水准!长远目标是让视频剪辑实现真正的AIGC!
基于 ComfyUI + MCP + LLM 的开源多模态 AIGC 解决方案,https://pixelle.aiAn Open-Source Multimodal AIGC Solution based on ComfyUI + MCP + LLM https://pixelle.ai
面向 LLM 和 LVLM 的强大模型压缩框架,适配 NVIDIA GPU 和华为昇腾 NPUA powerful model compression framework for LLMs and LVLMs, adapted for NVIDIA GPUs and Huawei Ascend NPUs.
Revolt 是排名第一的 Edgenuity 自动化工具,提供自动答题、自动写作文、自动跳课、humanization 驱动的 AI 写作、自动词汇、自动日志等功能。这是一款半挂机脚本,可轻松应对测验、考试、测试、项目和作文。支持桌面端和移动端。访问 https://revolt.ly 即可开始使用。Revolt is the #1 Edgenuity automation tool featuring Auto Quiz, Auto Essay, Auto Advance, AI-powered writing with humanization, Auto Vocabulary, Auto Journal, and more. A semi-AFK script that handles quizzes, tests, exams, projects, and essays with ease. Compatible with desktop and mobile. Visit https://revolt.ly to get started.
为纯文本 DeepSeek Harness Agent 装上眼睛:内置免费视觉链路(无需 key)+ 像素级视觉工具(问答、定位、裁剪、像素 diff、取色、OCR、SVG 矢量化、抠图、截图)。一条命令安装,无需 Python,图像回合与普通工具调用回合一致。Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.
🚀🚀🚀 收录关于 LLM、VLM、VLA、AIGC 及相关数据集与应用的一些优秀开源项目合集。🚀🚀🚀A collection of some awesome public projects about Large Language Model(LLM), Vision Language Model(VLM), Vision Language Action(VLA), AI Generated Content(AIGC), the related Datasets and Applications.
本地优先的视觉生成运行时与工作台,面向人类与编程 agent,跨多提供商提供可复现的图像与视频工作流。Local-first visual generation runtime and studio for people and coding agents, with reproducible image and video workflows across multiple providers.
CVPR2026/CVPR2025/ICCV2025/CVPR2024/ECCV2026/ECCV2024 AIGC 方向的论文与代码合集。A Collection of Papers and Codes for CVPR2026/CVPR2025/ICCV2025/CVPR2024/ECCV2026/ECCV2024 AIGC
Second Brain 是一个 Agentic 框架,作为操作系统运行,利用本地文件智能、工作流自动化和 LLM 来完成任务,并通过多种模态和消息平台进行通信。Second Brain is an agentic framework that acts as an operating system, using local file intelligence, workflow automation, and LLMs to complete tasks and communicate over multiple modalities and messaging platforms.
一个面向英文学术写作的保守型 AIGC 检测器指导的论文改写 Skill。支持 Turnitin AI、CNKI AIGC、最小化编辑修订、保留学术要素、定性/定量路由,以及逐章降低 AI 写作风险,且不宣称绕过检测器。A conservative AIGC detector-informed thesis rewriting skill for English and Chinese academic writing. Supports Turnitin AI, CNKI AIGC, minimal-edit revision, protected academic elements, qualitative/quantitative routing, and chapter-by-chapter AI-writing risk reduction without detector-bypass claims.
Claude skill,可基于 logo、文字或 emoji 生成 favicon、应用图标与社交媒体图片。支持 emoji 推荐、校验与框架自动集成。Claude skill to generate favicons, app icons, and social media images from logos, text, or emojis. Supports emoji suggestions, validation, and framework auto-integration.
为 Bilibili、YouTube 及本地视频提供 AI 视频摘要和知识库.AI video summarizer and knowledge base for Bilibili, YouTube and local videos.
预训练基础推荐模型论文清单Paper List of Pre-trained Foundation Recommender Models
全面解析 2026 年 8 月 GTA 6 玩法泄露事件:技术核查、暗网溯源、地图细节与诈骗警示Full breakdown of the August 2026 GTA 6 gameplay leaks, technical checks, dark web trail, map details, and scam warning.