研究库 开源仓库
Repositories · organized/repo_cards

仓库/Skill 库

281 个

排序 Stars 周增
nexu-io/html-video
HTML · 2026-06-21 多模态 教程 研究原型 Stars 4160 周增 +21

面向编码 Agent 的程序化视频方案——在本地将 HTML 转视频。把 HTML、CSS 与数据渲染为真实 MP4,支持可插拔渲染引擎、21 套模板与 AI 配乐。Apache-2.0,无按次计费。Open Design 团队的官方项目。Programmatic video for coding agents — HTML to video on your laptop. Turn HTML, CSS & data into real MP4s with pluggable render engines, 21 templates, AI soundtrack. Apache-2.0, no per-render fees. An official project by the Open Design team.

agentmultimodal
IDEA-CCNL/Fengshenbang-LM
Python · 2026-06-08 LLM 基础设施 模型 生产可用 Stars 4126 周增 +7

Fengshenbang-LM(封神榜大模型)是IDEA研究院认知计算与自然语言研究中心主导的大模型开源体系,成为中文AIGC和认知智能的基础设施。

multimodal
WUBING2023/PaperSpine
Python · 2026-07-01 多模态 应用 研究原型 Stars 4041 周增 +105

PaperSpine 是以动机驱动的 Skill,用于研读高质量学术论文、构建论文核心论点,并通过证据感知蓝图、修订矩阵与 LaTeX 安全审计来重写稿件。PaperSpine is a motivation-driven skill for learning from strong academic papers, building a paper’s central argument, and rewriting manuscripts through evidence-aware blueprints, revision matrices, and LaTeX-safe audits.

multimodalengineering
hao-ai-lab/FastVideo
Python · 2026-08-11 工程化 框架 研究原型 Stars 3936 周增 +7

用于加速视频生成的统一推理与后训练框架。A unified inference and post-training framework for accelerated video generation.

multimodalllm-infra
s1dashu/ip-as-logo-skill
未知语言 · 2026-08-22 Agent 智能体 应用 研究原型 Stars 3727 周增 +4646

一项紧凑的 Agent Skill,用于生成高度简化、圆润且带有微妙新拟物化风格的 IP 吉祥物 Logo。A compact Agent Skill for highly simplified, rounded, subtly neo-skeuomorphic IP mascot logos.

agentmultimodal
dramaclaw/dramaclaw
TypeScript · 2026-08-11 多模态 应用 研究原型 Stars 3493 周增 -126

A general-purpose AIGC video engine: script to finished film in one pipeline — dramas, ads, product videos, otome games, and more. | 通用 AIGC 视频引擎 —— 从剧本到成片一条流水线,漫剧、广告、电商、乙游皆可

agentmultimodalengineering
NVIDIA/skills
Python · 2026-09-23 Agent 智能体 应用 研究原型 Stars 3416 周增 +83

NVIDIA 产品的 Agent Skills — 安装到 Claude Code、Codex 和其他编码 Agent 中,端到端运行 Physical AI、机器人、仿真、CUDA 和 RAG 工作流。Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end.

agentragmultimodalllm-infra
Niko1221/Strata
C++ · 2026-10-01 LLM 基础设施 应用 研究原型 Stars 3310 周增 +4375

在任意消费级硬件上运行 Qwen3.8-Flash-Next:Windows / Linux 一键安装。Strata 推理引擎,本地 localhost 提供 OpenAI / Anthropic API,支持可选图像输入。Qwen3.8-Flash-Next on any consumer hardware: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input.

multimodalllm-infra
breezedeus/Pix2Text
Jupyter Notebook · 2026-02-07 多模态 工具 研究原型 Stars 3213 周增 +7

一款开源 Python3 工具,使用小模型识别图像中的版面、表格、数学公式(LaTeX)和文本,并转换为 Markdown 格式。Mathpix 的免费替代方案,可将视觉内容无缝转换为文本表示,支持 80+ 种语言。An Open-Source Python3 tool with SMALL models for recognizing layouts, tables, math formulas (LaTeX), and text in images, converting them into Markdown format. A free alternative to Mathpix, empowering seamless conversion of visual content into text-based representations. 80+ languages are supported.

multimodalengineering
openvinotoolkit/openvino_notebooks
Jupyter Notebook · 2026-08-10 LLM 基础设施 教程 研究原型 Stars 3195 周增 +0

📚 OpenVINO™ 的 Jupyter notebook 教程。📚 Jupyter notebook tutorials for OpenVINO™

multimodalllm-infra
HiLab-git/SSL4MIS
Python · 2025-06-07 多模态 收藏榜 研究原型 Stars 2671 周增 +0

半监督学习在医学图像分割中的应用,含文献综述与代码实现合集Semi Supervised Learning for Medical Image Segmentation, a collection of literature reviews and code implementations.

multimodal
genieincodebottle/generative-ai
Jupyter Notebook · 2026-08-05 工程化 收藏榜 研究原型 Stars 2593 周增 +14

生成式AI综合资源,包含详细路线图、项目、用例、面试准备与编码练习Comprehensive resources on Generative AI, including a detailed roadmap, projects, use cases, interview preparation, and coding preparation.

agentragmultimodalevaluation
roboflow/inference
Python · 2026-08-10 LLM 基础设施 应用 研究原型 Stars 2411 周增 +7

将任意电脑或边缘设备打造为计算机视觉项目的指挥中心Turn any computer or edge device into a command center for your computer vision projects.

agentmultimodalllm-infraengineering
aigc-apps/EasyAnimate
Python · 2025-03-06 多模态 应用 实验 Stars 2266 周增 +0

📺 基于 Transformer Diffusion 的高分辨率长视频生成端到端解决方案。📺 An End-to-End Solution for High-Resolution and Long Video Generation Based on Transformer Diffusion

multimodal
NVIDIA-NeMo/DataDesigner
Python · 2026-09-21 LLM 基础设施 工具 生产可用 Stars 2264 周增 +0

🎨 NeMo Data Designer:从零或从种子数据生成高质量合成数据。🎨 NeMo Data Designer: Generate high-quality synthetic data from scratch or from seed data.

agentmultimodalllm-infra
aigc-apps/VideoX-Fun
Python · 2026-08-10 多模态 框架 生产可用 Stars 2192 周增 +0

📹 一个更灵活的框架,支持任意分辨率视频生成以及从图像生成视频。📹 A more flexible framework that can generate videos at any resolution and creates videos from images.

multimodal
kingyiusuen/image-to-latex
Python · 2022-10-04 多模态 工具 实验 Stars 2160 周增 +0

将 LaTeX 数学公式的图像转换为 LaTeX 代码。Convert images of LaTex math equations into LaTex code.

multimodalengineering
NVIDIA-AI-Blueprints/video-search-and-summarization
Python · 2026-10-09 多模态 应用 研究原型 Stars 1918 周增 +26

NVIDIA AI Blueprint for video search and summarization(VSS)是一个 GPU 加速参考架构,用于构建具备实时验证告警、视觉问答与自动报告能力的视频分析 Agent。VSS Blueprint 采用 NVIDIA Cosmos 等视觉语言模型(VLM)、NVIDIA Nemotron 等 LLM,并结合 RAG 与 NVIDIA NIM。NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts, visual Q&A, and automated reporting. The VSS Blueprint uses vision language models (VLMs) such as NVIDIA Cosmos, LLMs such as NVIDIA Nemotron, RAG, and NVIDIA NIMs.

agentragmultimodalllm-infra
szczyglis-dev/py-gpt
Python · 2026-09-03 多模态 应用 研究原型 Stars 1904 周增 +0

由 GPT-5、GPT-4、o1、o3、Gemini、Claude、Ollama、DeepSeek、Perplexity、Grok、Bielik 驱动的桌面 AI Assistant,支持 chat、vision、voice、RAG、图像与视频生成、agents、tools、MCP、plugins、语音合成与识别、网页搜索、记忆、预设、assistants 等。Linux、Windows、Mac。Desktop AI Assistant powered by GPT-5, GPT-4, o1, o3, Gemini, Claude, Ollama, DeepSeek, Perplexity, Grok, Bielik, chat, vision, voice, RAG, image and video generation, agents, tools, MCP, plugins, speech synthesis and recognition, web search, memory, presets, assistants,and more. Linux, Windows, Mac

agentragmultimodalllm-infra
hymie122/RAG-Survey
未知语言 · 2024-08-20 RAG 检索增强 收藏榜 实验 Stars 1790 周增 +0

收录 AIGC 领域 RAG 的优秀论文。我们在论文 "Retrieval-Augmented Generation for AI-Generated Content: A Survey" 中提出了 RAG 基础、增强与应用分类法。Collecting awesome papers of RAG for AIGC. We propose a taxonomy of RAG foundations, enhancements, and applications in paper "Retrieval-Augmented Generation for AI-Generated Content: A Survey".

ragmultimodalllm-infra
jau123/MeiGen-AI-Design-MCP
TypeScript · 2026-08-05 Agent 智能体 库 研究原型 Stars 1689 周增 +14

支持GPT Image 2、Seedance与ComfyUI,配备1,400+提示词库、精心设计的hooks以及多任务编排系统Supports GPT Image 2, Seedance & ComfyUI, with a 1,400+ prompt library, carefully crafted hooks and a multi-task orchestration system

multimodal
HITsz-TMG/VideoClaw
Python · 2026-07-17 多模态 框架 研究原型 Stars 1685 周增 +21

🚀 AI 全自动化视频生成员工 | Your First AIGC Coworker. Chat an Idea. Get a Film. 🦞

agentmultimodal
bytedance/Sa2VA
Python · 2026-09-08 多模态 框架 研究原型 Stars 1666 周增 +0

Pixel-LLM 代码库官方仓库:Sa2VA(T-PAMI-26)、SAMTok(CVPR-26)、VRT(Arxiv-25)、SaSaSa2VA(LSVOS 第一名方案)。Official Repo For Pixel-LLM Codebase: Sa2VA (T-PAMI-26), SAMTok (CVPR-26), VRT (Arxiv-25), SaSaSa2VA (1-st solution for LSVOS)

multimodalllm-infra
pixeltable/pixeltable
Python · 2026-10-07 Agent 智能体 应用 研究原型 Stars 1638 周增 +0

后端 Agent 构建工具 —— 多模态数据库、编排与服务,一个文件搞定The backend agents build with - Multimodal database, orchestration, and serving in one file

agentragmultimodalllm-infra
limecloud/lime
TypeScript · 2026-10-04 Agent 智能体 应用 研究原型 Stars 1482 周增 +0

用于编程、文件、终端、工具、研究、内容、多模态工作和多 Agent 工作流的全栈 AI Agent。Full-stack AI agent for coding, files, terminals, tools, research, content, multimodal work, and multi-agent workflows.

agentmultimodal
Devin-AXIS/iPolloWork
TypeScript · 2026-07-20 Agent 智能体 应用 研究原型 Stars 1452 周增 +266

下一代源码可用的 Codex 与 Claude Code 替代品——一个本地优先、可自托管的 Agent 工作空间,覆盖代码、办公文档、可编辑设计、演示文稿、网站与视频。用 AI 构建后,可像 PowerPoint 一样轻松编辑文字、图片、配色、布局与场景。Next-gen, source-available alternative to Codex and Claude Code — one local-first, self-hostable agent workspace for code, office work, editable design, presentations, websites, and video. Build with AI, then edit text, images, colors, layouts, and scenes as easily as PowerPoint.

agentmultimodalllm-infra
waybarrios/vllm-mlx
Python · 2026-06-28 多模态 工具 研究原型 Stars 1447 周增 +7

兼容 OpenAI 和 Anthropic 协议的 Apple Silicon 服务端。可运行 LLM 和视觉语言模型(Llama、Qwen-VL、LLaVA),支持 continuous batching、MCP 工具调用与多模态。原生 MLX 后端,速度达 400+ tok/s,兼容 Claude Code。OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous batching, MCP tool calling, and multimodal support. Native MLX backend, 400+ tok/s. Works with Claude Code.

multimodalllm-infra
kimsungwhee/apple-docs-mcp
TypeScript · 2026-03-17 Agent 智能体 教程 研究原型 Stars 1347 周增 +7

面向Apple开发者文档的MCP服务器——在Claude、Cursor及AI助手中检索iOS/macOS/SwiftUI/UIKit文档、WWDC视频、Swift/Objective-C API及代码示例MCP server for Apple Developer Documentation - Search iOS/macOS/SwiftUI/UIKit docs, WWDC videos, Swift/Objective-C APIs & code examples in Claude, Cursor & AI assistants

multimodal
liyue-aigc/female-portrait-director
未知语言 · 2026-07-15 多模态 工具 实验 Stars 1336 周增 +21

用于引导和扩展详细 AI 女性肖像 prompt 的模块化 Codex Skill。A modular Codex Skill for directing and expanding detailed AI female portrait prompts.

multimodal
yzfly/douyin-mcp-server
HTML · 2026-07-02 Agent 智能体 应用 研究原型 Stars 1246 周增 +0

提取抖音无水印视频链接,视频文案,douyin-mcp-server,mcp,claude skill,支持龙虾

agentmultimodalllm-infra
gyoridavid/short-video-maker
TypeScript · 2025-06-21 Agent 智能体 模型 实验 Stars 1225 周增 +0

使用 Model Context Protocol (MCP) 和 REST API 为 TikTok、Instagram Reels 和 YouTube Shorts 创建短视频。Creates short videos for TikTok, Instagram Reels, and YouTube Shorts using the Model Context Protocol (MCP) and a REST API.

multimodal
all-in-aigc/sorafm
TypeScript · 2024-08-15 多模态 应用 实验 Stars 1152 周增 +14

Sora.FM 推出的 Sora AI 视频生成器。Sora AI Video Generator by Sora.FM

multimodal
pierrenade/short-video-generator-AI
Python · 2026-09-07 多模态 应用 研究原型 Stars 1148 周增 +0

免费开源项目,将 youtube 视频转化为爆款短视频。集成高光检测、字幕、翻译与配音,一站式内容制作。Free open-source project designed for turning youtube-viedos into viral short videos. Highlight detection, subtitles, translation, voiceover, all in one for your content.

multimodal
FutureUniant/Tailor
Python · 2025-06-03 多模态 工具 生产可用 Stars 1107 周增 +0

Tailor是一款视频智能裁剪、视频生成和视频优化的视频剪辑工具。目前的目标是通过人工智能技术减少视频剪辑的繁琐操作,让普通人也能简单实现专业剪辑人的水准!长远目标是让视频剪辑实现真正的AIGC!

multimodal
ATH-MaaS/Pixelle-MCP
Python · 2025-12-17 Agent 智能体 应用 实验 Stars 1100 周增 +7

基于 ComfyUI + MCP + LLM 的开源多模态 AIGC 解决方案,https://pixelle.aiAn Open-Source Multimodal AIGC Solution based on ComfyUI + MCP + LLM https://pixelle.ai

multimodalllm-infra
Vincentwei1021/anything2explainer
TypeScript · 2026-09-12 多模态 应用 研究原型 Stars 1040 周增 +0

输入主题,输出带解说的解释视频。一款 Claude Code / Codex Skill,可将任意主题转化为黑底动态图形解释视频,配备 TTS 配音、字幕和章节进度条。支持中文或英文;每一帧均通过 Remotion 以代码方式绘制。Topic in, narrated explainer video out. A Claude Code / Codex skill that turns any topic into a black-canvas motion-graphics explainer video with TTS voiceover, subtitles and a chapter progress bar. Chinese or English; every frame drawn in code with Remotion.

agentmultimodal