OCRmyPDF 为扫描的 PDF 文件添加 OCR 文本层,使其可被搜索OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
仓库/Skill 库
78 个 · 多模态 · 学术写作
pix2tex:使用 ViT 将公式图像转换为 LaTeX 代码pix2tex: Using a ViT to convert images of equations into LaTeX code.
AudioGPT:理解与生成语音、音乐、声音与说话人头像AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
OpenMMLab 多模态高级生成与智能创作工具箱。释放魔力🪄:AIGC、易用 API、丰富模型库、扩散模型,支持文生图、图像/视频修复与增强等任务OpenMMLab Multimodal Advanced, Generative, and Intelligent Creation Toolbox. Unlock the magic 🪄: Generative-AI (AIGC), easy-to-use APIs, awsome model zoo, diffusion models, for text-to-image generation, image/video restoration/enhancement, etc.
AIGCPanel 是一个简单易用的一站式AI数字人系统,支持视频合成、声音合成、声音克隆,简化本地模型管理、一键导入和使用AI模型。
📷 EasyPhoto | 你的智能 AI 照片生成器。📷 EasyPhoto | Your Smart AI Photo Generator.
A general-purpose AIGC video engine: script to finished film in one pipeline — dramas, ads, product videos, otome games, and more. | 通用 AIGC 视频引擎 —— 从剧本到成片一条流水线,漫剧、广告、电商、乙游皆可
一款开源 Python3 工具,使用小模型识别图像中的版面、表格、数学公式(LaTeX)和文本,并转换为 Markdown 格式。Mathpix 的免费替代方案,可将视觉内容无缝转换为文本表示,支持 80+ 种语言。An Open-Source Python3 tool with SMALL models for recognizing layouts, tables, math formulas (LaTeX), and text in images, converting them into Markdown format. A free alternative to Mathpix, empowering seamless conversion of visual content into text-based representations. 80+ languages are supported.
半监督学习在医学图像分割中的应用,含文献综述与代码实现合集Semi Supervised Learning for Medical Image Segmentation, a collection of literature reviews and code implementations.
永久免费开源的 AIGC 课程, 目前已支持Claude Code,Codex,Hermes,OpenClaw,Obsidian,Prompt Engineering, ChatGPT, Midjourney, Runway, Stable Diffusion, AI数字人,AI声音&音乐,开源大模型
📺 基于 Transformer Diffusion 的高分辨率长视频生成端到端解决方案。📺 An End-to-End Solution for High-Resolution and Long Video Generation Based on Transformer Diffusion
📹 一个更灵活的框架,支持任意分辨率视频生成以及从图像生成视频。📹 A more flexible framework that can generate videos at any resolution and creates videos from images.
将 LaTeX 数学公式的图像转换为 LaTeX 代码。Convert images of LaTex math equations into LaTex code.
🚀 AI 全自动化视频生成员工 | Your First AIGC Coworker. Chat an Idea. Get a Film. 🦞
这是一个基于Claude Skill的**AI人像Prompt生成系统**,能够从特征库中智能组合生成高质量的人像描述Prompt,并具备自动学习和库扩展能力。 核心能力: Prompt生成、特征提取、自动学习、智能审核、版本控制
用于引导和扩展详细 AI 女性肖像 prompt 的模块化 Codex Skill。A modular Codex Skill for directing and expanding detailed AI female portrait prompts.
Tailor是一款视频智能裁剪、视频生成和视频优化的视频剪辑工具。目前的目标是通过人工智能技术减少视频剪辑的繁琐操作,让普通人也能简单实现专业剪辑人的水准!长远目标是让视频剪辑实现真正的AIGC!
AI 音频数据集(AI-ADS)🎵,包含语音、音乐和音效,可为 Generative AI、AIGC、AI 模型训练、智能音频工具开发及音频应用提供训练数据。AI Audio Datasets (AI-ADS) 🎵, including Speech, Music, and Sound Effects, which can provide training data for Generative AI, AIGC, AI model training, intelligent audio tool development, and audio applications.
🚀🚀🚀 收录关于 LLM、VLM、VLA、AIGC 及相关数据集与应用的一些优秀开源项目合集。🚀🚀🚀A collection of some awesome public projects about Large Language Model(LLM), Vision Language Model(VLM), Vision Language Action(VLA), AI Generated Content(AIGC), the related Datasets and Applications.
CVPR2026/CVPR2025/ICCV2025/CVPR2024/ECCV2026/ECCV2024 AIGC 方向的论文与代码合集。A Collection of Papers and Codes for CVPR2026/CVPR2025/ICCV2025/CVPR2024/ECCV2026/ECCV2024 AIGC
生成干净的 2D 游戏精灵图与动画图集——组件行流水线:状态行、alpha 清理、帧提取、运行时图集。Codex/Claude skill。Generate clean 2D game sprites & animation atlases — component-row pipeline: state rows, alpha cleanup, frame extraction, runtime atlases. Codex/Claude skill.
让幻灯片摆脱 AI 味的 Claude skill。品牌 DNA × 20 条编码化设计原则。A Claude skill for slides that don't look like AI made them. Brand DNA × 20 codified design principles.
你的个人 AIGC 工厂。任意图像,任意视频,以 Comfy 方式实现。©️Your own personal AIGC Factory. Any picture. Any reel. The Comfy way. ©️
一个面向英文学术写作的保守型 AIGC 检测器指导的论文改写 Skill。支持 Turnitin AI、CNKI AIGC、最小化编辑修订、保留学术要素、定性/定量路由,以及逐章降低 AI 写作风险,且不宣称绕过检测器。A conservative AIGC detector-informed thesis rewriting skill for English and Chinese academic writing. Supports Turnitin AI, CNKI AIGC, minimal-edit revision, protected academic elements, qualitative/quantitative routing, and chapter-by-chapter AI-writing risk reduction without detector-bypass claims.
Claude skill,可基于 logo、文字或 emoji 生成 favicon、应用图标与社交媒体图片。支持 emoji 推荐、校验与框架自动集成。Claude skill to generate favicons, app icons, and social media images from logos, text, or emojis. Supports emoji suggestions, validation, and framework auto-integration.
[CVPR2024 (Highlight)] RichDreamer:一种可泛化的法线-深度扩散模型,用于生成细节丰富的文本到 3D 内容。Live Demo:https://modelscope.cn/studios/Damo_XR_Lab/3D_AIGC[CVPR2024 (Highlight)] RichDreamer: A Generalizable Normal-Depth Diffusion Model for Detail Richness in Text-to-3D. Live Demo:https://modelscope.cn/studios/Damo_XR_Lab/3D_AIGC