为纯文本 DeepSeek Harness Agent 装上眼睛:内置免费视觉链路(无需 key)+ 像素级视觉工具(问答、定位、裁剪、像素 diff、取色、OCR、SVG 矢量化、抠图、截图)。一条命令安装,无需 Python,图像回合与普通工具调用回合一致。Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.
仓库/Skill 库
167 个 · 多模态
🚀🚀🚀 收录关于 LLM、VLM、VLA、AIGC 及相关数据集与应用的一些优秀开源项目合集。🚀🚀🚀A collection of some awesome public projects about Large Language Model(LLM), Vision Language Model(VLM), Vision Language Action(VLA), AI Generated Content(AIGC), the related Datasets and Applications.
CVPR2026/CVPR2025/ICCV2025/CVPR2024/ECCV2026/ECCV2024 AIGC 方向的论文与代码合集。A Collection of Papers and Codes for CVPR2026/CVPR2025/ICCV2025/CVPR2024/ECCV2026/ECCV2024 AIGC
生成干净的 2D 游戏精灵图与动画图集——组件行流水线:状态行、alpha 清理、帧提取、运行时图集。Codex/Claude skill。Generate clean 2D game sprites & animation atlases — component-row pipeline: state rows, alpha cleanup, frame extraction, runtime atlases. Codex/Claude skill.
让幻灯片摆脱 AI 味的 Claude skill。品牌 DNA × 20 条编码化设计原则。A Claude skill for slides that don't look like AI made them. Brand DNA × 20 codified design principles.
你的个人 AIGC 工厂。任意图像,任意视频,以 Comfy 方式实现。©️Your own personal AIGC Factory. Any picture. Any reel. The Comfy way. ©️
一个面向英文学术写作的保守型 AIGC 检测器指导的论文改写 Skill。支持 Turnitin AI、CNKI AIGC、最小化编辑修订、保留学术要素、定性/定量路由,以及逐章降低 AI 写作风险,且不宣称绕过检测器。A conservative AIGC detector-informed thesis rewriting skill for English and Chinese academic writing. Supports Turnitin AI, CNKI AIGC, minimal-edit revision, protected academic elements, qualitative/quantitative routing, and chapter-by-chapter AI-writing risk reduction without detector-bypass claims.
Claude skill,可基于 logo、文字或 emoji 生成 favicon、应用图标与社交媒体图片。支持 emoji 推荐、校验与框架自动集成。Claude skill to generate favicons, app icons, and social media images from logos, text, or emojis. Supports emoji suggestions, validation, and framework auto-integration.
[CVPR2024 (Highlight)] RichDreamer:一种可泛化的法线-深度扩散模型,用于生成细节丰富的文本到 3D 内容。Live Demo:https://modelscope.cn/studios/Damo_XR_Lab/3D_AIGC[CVPR2024 (Highlight)] RichDreamer: A Generalizable Normal-Depth Diffusion Model for Detail Richness in Text-to-3D. Live Demo:https://modelscope.cn/studios/Damo_XR_Lab/3D_AIGC
为 Bilibili、YouTube 及本地视频提供 AI 视频摘要和知识库.AI video summarizer and knowledge base for Bilibili, YouTube and local videos.
决定你的 AI 编码 Agent 阅读哪些内容,并保留凭证。每次裁剪都可逐字节恢复,从不杜撰结果,所有数据从你自己的语料库回放。Decides what your AI coding agent reads, and keeps receipts. Every cut is recoverable byte for byte, it never invents a result, and the numbers are replayed from your own corpus.
所有人可用的 AI(MoAI)——面向韩语实务领域的 Claude Cowork 与 Claude Code AI harness 与插件市场。覆盖商业计划书、税务、法律、HR、营销、电商、BI、内容等领域,提供 Skill、Agent 与工作流。支持韩语 B2B 场景与办公文档(HWPX/DOCX/XLSX/PPTX/PDF)及 AI 多模态生成(图像/视频/语音)。内置 AI 痕迹审核与韩语 humanize-korean모두의 AI (MoAI) — Claude Cowork & Claude Code 한국 실무 도메인 AI 하네스·플러그인 마켓플레이스. 사업계획서·세무·법률·HR·마케팅·커머스·BI·콘텐츠 도메인 스킬·에이전트·워크플로우. Korean B2B + office docs (HWPX/DOCX/XLSX/PPTX/PDF) + AI media (image/video/voice). AI-slop 검수 + humanize-korean 내장.
Chrome 扩展:固定一个屏幕区域后,可用快捷键翻阅分页文档。OCR 通过内置 Tesseract 100% 离线运行。Chrome extension: pin a screen region once, then hotkey your way through a paginated document. OCR runs 100% offline via bundled Tesseract.
[ISPRS 2026] 遥感图像去雾:进展、挑战与前景的系统综述[ISPRS 2026] Remote Sensing Image Dehazing: A Systematic Review of Progress, Challenges, and Prospects
深度学习与 EEG 系统文献综述的补充材料Supplementary material for systematic literature review on deep learning and EEG.
AI驱动的多媒体内容生产skills集合:document-writer(写作)、illustration-generator(配图)、ppt-generator(PPT风格)、podcast-generator(TTS)、remoti on-dev(视频制作)、twitter-crawler(推文爬取)、markdown-illustrator(Markdown配图)、comic-generator(漫画生成)、media-downloader(媒体下载)、tts-script-generator(TTS脚本)、md-t o-pdf(文档转换)、wechat-formatter(微信格式化)、humanizer-zh(中文人性化)、shared-lib(核心API库)
ComfyUI 节点,为 MiniMax H3 而设计,可串联多个视频片段,具备运动上下文、磁盘缓存、动态图像参考、音频参考支持以及最终视频/音频的无缝解码ComfyUI node for MiniMax H3 that chains multiple video clips with motion context, disk caching, dynamic image references, audio reference support, and seamless final video/audio decoding.
面向生成式接口的无障碍 React 加载状态:流式文本、内联活动指示与图像生成。Accessible React loading states for generative interfaces: streamed text, inline activity, and image generation.
语音听写,同时结合你所指屏幕位置的上下文。Voice dictation with the screen context you point at.
用双手操控屏幕——基于摄像头的免穿戴、免手柄手势追踪界面,让你的 AI 直接响应动作Move things on your screen with your bare hands. A webcam-powered, hand-tracked interface for your AI. No headset. No controllers.
综述论文"A Systematic Review of Deep Learning-based Research on Radiology Report Generation"的官方 GitHub 仓库The official GitHub repository of the survey paper "A Systematic Review of Deep Learning-based Research on Radiology Report Generation".
PCB_lightgraph_portable 是一款点击即可运行的 PCB 图片智能分层与图纸导出的html工具。
OpenRouter 的 MCP server——与 300+ LLM(Claude、Gemini、GPT)对话,分析图像/音频/视频,生成图像/语音/音乐/视频(Veo 3.1、Sora、Seedance、Wan),支持 Claude Desktop、Cursor、Kiro、VS CodeMCP server for OpenRouter — chat with 300+ LLMs (Claude, Gemini, GPT), analyze images / audio / video, generate images / speech / music / video (Veo 3.1, Sora, Seedance, Wan) from Claude Desktop, Cursor, Kiro, VS Code.
https://abox.tools/ 的源码——一个面向图像、视频、音频、PDF 与文本的小型 Web 工具集。文件永不离开本机,因为不存在任何能将它们外发的代码路径。Source for https://abox.tools/ — a box of small web tools for images, video, audio, PDFs and text. Your files never leave your machine, because there is no code path that could send them anywhere.