将任意 PDF 或图片文档转换为结构化数据供 AI 使用。强大而轻量的 OCR 工具集,弥合图像/PDF 与 LLM 之间的鸿沟,支持 100+ 种语言。Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
仓库/Skill 库
33 个 · 多模态 · 工具 · 快速增长
将参考图像中的物体重建为纯代码、程序化、质量可控、可直接用于动画的 Three.js 模型。Token 高效的图像转 3D。Rebuild the object in a reference image as a code-only, procedural, quality-gated, animation-ready Three.js model. Token-efficient image-to-3D.
中文小黑怪诞正文配图生成 Skill | 16:9 白底手绘 | 少量红橙蓝批注 | Codex Skill
通过将文本上下文渲染为图像,将 Fable 5 的 token 使用量降低cut Fable 5 token usage by rendering text context as images
🪧 Claude Code / Codex Skill——生成小红书图文卡片与公众号 21:9+1:1 封面配对。编辑 × Swiss 视觉系统,28 套版式,10 种主题,单文件 HTML → PNG。小红书图文 + 公众号封面对🪧 Claude Code / Codex skill — generate Xiaohongshu carousels & WeChat 21:9+1:1 cover pairs. Editorial × Swiss visual systems, 28 layouts, 10 themes, single-file HTML → PNG. 小红书图文 + 公众号封面对
Codex skill,用于生成安静极简的 zine 风格编辑海报提示词与图像。Codex skill for generating quiet minimal zine-style editorial poster prompts and images.
将参考图像中的对象重建为纯代码、程序化、带质量门控、可动画化的 Three.js 模型。token 高效的图像到三维转换。Rebuild the object in a reference image as a code-only, procedural, quality-gated, animation-ready Three.js model. Token-efficient image-to-3D.
为纯文本 DeepSeek Harness Agent 装上眼睛:内置免费视觉链路(无需 key)+ 像素级视觉工具(问答、定位、裁剪、像素 diff、取色、OCR、SVG 矢量化、抠图、截图)。一条命令安装,无需 Python,图像回合与普通工具调用回合一致。Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.
DLSS 5 神经渲染,实时运行在你的整个 Windows 桌面上。支持 RTX 30/40/50,12 种语言,用户预设,带音频录制,单窗口模式。DLSS 5 Neural Rendering on your whole Windows desktop, in real time. RTX 30/40/50, 12 languages, user presets, recording with audio, one-window mode.
Learning notes and tooling skills for AI video - AI 视频相关的学习与工具 skill
Chrome 扩展:固定一个屏幕区域后,可用快捷键翻阅分页文档。OCR 通过内置 Tesseract 100% 离线运行。Chrome extension: pin a screen region once, then hotkey your way through a paginated document. OCR runs 100% offline via bundled Tesseract.
将用户的描述或上传的参考素材转化为一张成品的、可编辑的 Blender 卡片以及一个交互式 Three.js 页面。保留所要求的主题、风格、字体与目标场景。本 skill 仅包含代码与文本;生成的艺术作品属于用户的输出项目。Turn the user's description or uploaded reference into a finished, editable Blender card and an interactive Three.js page. Preserve the requested subject, style, typography and destination. This skill contains code and text only; generated artwork belongs in the user's output project.
将 YouTube、Bilibili 或本地课程/会议录像转换为带幻灯片配图的 Markdown 与 HTML 讲义。Turn YouTube, Bilibili, or local course/meeting recordings into slide-illustrated Markdown and HTML lecture notes.
通过小型 ProgramAsWeights 函数实现英语与 Claudish 之间的相互翻译Translate between English and Claudish with tiny ProgramAsWeights functions.
在 ComfyUI 中运行 MiniMax-H3 VAE 的 ONNX/TRT 版本,速度最高提升至 1.7 倍。Running ONNX/TRT version of the MiniMax-H3 VAE in ComfyUI, which increase speed by up to 1.7x
使用 ChatGPT 订阅在 Codex 或 Claude Code 中生成 GPT 图像,无需 Images API。Generate GPT images from Codex or Claude Code using a ChatGPT subscription, without the Images API.
YuE2 音乐生成,以及面向 ComfyUI 的可编辑钢琴卷帘界面。YuE2 music generation and an editable piano roll for ComfyUI
PCB_lightgraph_portable 是一款点击即可运行的 PCB 图片智能分层与图纸导出的html工具。
基于 Optical Flow 的 VR 帧生成。Frame generation for VR using Optical Flow
面向 Linux 的真活体检测人脸解锁:一个 PAM 模块、一个守护进程,以及一个 Omarchy 锁屏指示器。Face unlock for Linux with real liveness detection: a PAM module, a daemon, and an Omarchy lock screen indicator
由豆包驱动的语音输入工具,适用于 Linux、Wayland 和 OmarchyDoubao-powered voice input for Linux, Wayland and Omarchy
把 ComfyUI-MiniMaxH3-Easy 与 Goohai-MiniMax-H3_Integration 合并成一个统一的 MiniMax H3 创作工作台(非官方)
NVIDIA DLSS 5 神经渲染网络开源复现 OpenDLSS-NR 的 Three.js(TSL / WebGPU)移植版。A Three.js (TSL / WebGPU) port of OpenDLSS-NR, the open-source reimplementation of NVIDIA's DLSS 5 neural rendering network
CIT Voice Studio —— 将越南语文本转为语音,完全在本地运行。CIT Voice Stuido - Chuyển văn bản tiếng Việt thành giọng nói, chạy hoàn toàn trên máy bạn.
一款把任意图片转换为专业生图提示词的 Chrome / Edge 浏览器插件。支持网页右键识图、本地上传与剪贴板粘贴,生成中文、英文和 JSON 提示词,并提供历史记录与收藏。