Transformers:面向文本、视觉、音频及多模态 SOTA 机器学习模型的模型定义框架,同时支持推理与训练。🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
仓库/Skill 库
167 个 · 多模态
π RuView 将现成 WiFi 信号转化为实时空间智能、生命体征监测和存在检测,全程无需任何视频画面。π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video.
将任意 PDF 或图片文档转换为结构化数据供 AI 使用。强大而轻量的 OCR 工具集,弥合图像/PDF 与 LLM 之间的鸿沟,支持 100+ 种语言。Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
OpenAI Whisper 模型的 C/C++ 移植版本。Port of OpenAI's Whisper model in C/C++
LocalAI 是开源 AI 引擎,可在任何硬件上运行任意模型 —— LLMs、视觉、语音、图像、视频,无需 GPU。LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.
开源桌面应用,面向本地 LLM。支持文本、视觉、tool-calling,提供 OpenAI/Anthropic 兼容 API。100% 隐私保护。Open-source desktop app for local LLMs. Text, vision, tool-calling, OpenAI/Anthropic-compatible API. 100% private.
🐸💬 —— 一个历经研究与生产环境考验的 Text-to-Speech 深度学习工具包。🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
面向实时和流媒体、跨平台可定制的 ML 解决方案。Cross-platform, customizable ML solutions for live and streaming media.
OCRmyPDF 为扫描的 PDF 文件添加 OCR 文本层,使其可被搜索OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
开箱即用的 OCR,支持 80+ 种语言及所有主流书写系统,包括拉丁文、中文、阿拉伯文、天城文、西里尔文等。Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.
一款免费、开源且可扩展的语音转文字应用程序,完全离线运行。A free, open source, and extensible speech-to-text application that works completely offline.
DeepSpeech 是一个开源嵌入式(离线、端侧)语音转文字引擎,可在从 Raspberry Pi 4 到高性能 GPU 服务器的设备上实时运行。DeepSpeech is an open source embedded (offline, on-device) speech-to-text engine which can run in real time on devices ranging from a Raspberry Pi 4 to high power GPU servers.
开源无限制的 AI 视频平台替代方案 —— 免费 AI 图像与视频生成工作室,内置 200+ 模型(Flux、Midjourney、Kling、Sora、Veo)。无内容过滤,自托管,MIT 许可。Unrestricted Open-source alternative to AI video platforms — Free AI image & video generation studio with 500+ models (Flux, Midjourney, Kling, Sora, Veo). No content filters. Self-hosted, MIT licensed.
LabelImg 现已成为 Label Studio 社区的一部分。由 Tzutalin 创建的热门图像标注工具已不再积极开发,但你可以查看 Label Studio——这款开源数据标注工具支持图像、文本、超文本、音频、视频和时序数据。LabelImg is now part of the Label Studio community. The popular image annotation tool created by Tzutalin is no longer actively being developed, but you can check out Label Studio, the open source data labeling tool for images, text, hypertext, audio, video and time-series data.
围绕 PyTorch 在视觉、文本、强化学习等领域的一组示例。A set of examples around pytorch in Vision, Text, Reinforcement Learning, etc.
Unlimited OCR Works:迈入一键长文档解析的时代。Unlimited OCR Works: Welcome the Era of One-shot Long-horizon Parsing.
OpenUI 可让你凭想象力描述 UI,并实时看到渲染效果。OpenUI let's you describe UI using your imagination, then see it rendered live.
工业级可控、高效的零样本文本转语音系统An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
为研究人员和开发者构建的可扩展生成式 AI 框架,面向 LLM、多模态和语音 AI(自动语音识别与文字转语音)。A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
✨✨多模态大语言模型最新进展。:sparkles::sparkles:Latest Advances on Multimodal Large Language Models
pix2tex:使用 ViT 将公式图像转换为 LaTeX 代码pix2tex: Using a ViT to convert images of equations into LaTeX code.
Toonflow 是开源一站式 AI 短剧创作工具,将小说、剧本快速转化为动画短剧。集成 AI 编剧、智能分镜、角色与视频生成,跨平台桌面端轻量部署,助力创作者低成本批量产出视觉内容。Toonflow is an open-source AI tool that turns stories and scripts into animated short dramas. Features AI scriptwriting, storyboarding, character and video generation. A cross-platform desktop app for efficient content creation.
将参考图像中的物体重建为纯代码、程序化、质量可控、可直接用于动画的 Three.js 模型。Token 高效的图像转 3D。Rebuild the object in a reference image as a code-only, procedural, quality-gated, animation-ready Three.js model. Token-efficient image-to-3D.
在 Krita 中使用 AI 生成图像的精简界面。支持 Inpaint 与 Outpaint,可选文本提示,无需调参。Streamlined interface for generating images with AI in Krita. Inpaint and outpaint with optional text prompt, no tweaking required.
AudioGPT:理解与生成语音、音乐、声音与说话人头像AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
通过修改一行代码即可将 GPT 替换为任意 LLM。Xinference 让你在云端、本地或笔记本上运行开源、语音和多模态模型,全部通过统一的生产就绪推理 API。Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.
轻松微调、评估和部署 Gemma 4、Qwen3.5、Qwen3.6、gpt-oss、DeepSeek-R1 或任意开源 LLM / VLM。Easily fine-tune, evaluate and deploy Qwen, Gemma, or any open weight LLM!
LTX-2 音视频生成模型的官方 Python 推理与 LoRA 训练包。Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.
中文小黑怪诞正文配图生成 Skill | 16:9 白底手绘 | 少量红橙蓝批注 | Codex Skill
将任意品牌转化为可滚动 3D 世界落地页的 skillA skill that turn any brand into a scrollable 3D world landing page
OpenMMLab 多模态高级生成与智能创作工具箱。释放魔力🪄:AIGC、易用 API、丰富模型库、扩散模型,支持文生图、图像/视频修复与增强等任务OpenMMLab Multimodal Advanced, Generative, and Intelligent Creation Toolbox. Unlock the magic 🪄: Generative-AI (AIGC), easy-to-use APIs, awsome model zoo, diffusion models, for text-to-image generation, image/video restoration/enhancement, etc.
通过将文本上下文渲染为图像,将 Fable 5 的 token 使用量降低cut Fable 5 token usage by rendering text context as images