🐸💬 —— 一个历经研究与生产环境考验的 Text-to-Speech 深度学习工具包。🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
仓库/Skill 库
20 个 · 多模态 · 应用 · AI 核心
一款免费、开源且可扩展的语音转文字应用程序,完全离线运行。A free, open source, and extensible speech-to-text application that works completely offline.
DeepSpeech 是一个开源嵌入式(离线、端侧)语音转文字引擎,可在从 Raspberry Pi 4 到高性能 GPU 服务器的设备上实时运行。DeepSpeech is an open source embedded (offline, on-device) speech-to-text engine which can run in real time on devices ranging from a Raspberry Pi 4 to high power GPU servers.
OpenUI 可让你凭想象力描述 UI,并实时看到渲染效果。OpenUI let's you describe UI using your imagination, then see it rendered live.
工业级可控、高效的零样本文本转语音系统An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
在 Krita 中使用 AI 生成图像的精简界面。支持 Inpaint 与 Outpaint,可选文本提示,无需调参。Streamlined interface for generating images with AI in Krita. Inpaint and outpaint with optional text prompt, no tweaking required.
通过修改一行代码即可将 GPT 替换为任意 LLM。Xinference 让你在云端、本地或笔记本上运行开源、语音和多模态模型,全部通过统一的生产就绪推理 API。Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.
轻松微调、评估和部署 Gemma 4、Qwen3.5、Qwen3.6、gpt-oss、DeepSeek-R1 或任意开源 LLM / VLM。Easily fine-tune, evaluate and deploy Qwen, Gemma, or any open weight LLM!
面向 Apple Silicon 的端侧语音 AI。On-device Speech AI for Apple Silicon
面向 LLM、VLM、DiT 和 REC 模型的高性能推理引擎,针对多种 AI 加速器进行了优化。该项目托管于 OpenAtom 基金会。A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
高性能推理引擎,支持 LLM、VLM、DiT 和 REC 模型,针对多种 AI 加速器优化A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators.
[COLMW'26] AdaFlash:通过 On-Policy 蒸馏扩散 Drafters 实现的自适应投机解码[COLMW'26] AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters
在 LIBERO 与 RoboTwin2.0 上,基于状态融合解码与冻结视频骨干网络训练面向机器人操作的轻量级世界动作模型。Train lightweight world action models for robot manipulation using state-fusion decoding and frozen video backbones on LIBERO and RoboTwin2.0.
AI 教材数字化与互动辅导平台:PDF/扫描教材经 MinerU OCR 结构化出题,文本 + 视觉双模型审校与确定性质量门禁;学生端七种题型互动、分层提示与错题多轮陪练闭环;FastAPI + PostgreSQL JSONB + React 19
基于 Llama 3.2 的本地多模态 RAG AI 助手,可与 PDF、图片和视频对话,完全离线,零数据外泄。Local multimodal RAG AI assistant where you can chat with PDFs, images, and video using Llama 3.2, fully offline, zero data exposure.
开源 Raspberry Pi 5 rover 项目,用于建图、视觉、语音、本地 LLM 及液态神经网络实验。Open-source Raspberry Pi 5 rover project for mapping, vision, voice, local LLMs, and liquid-neural-network experiments.
使用 Qwen3-TTS 在本地 GPU 上克隆声音并从文本生成语音,提供端到端训练流水线。Clone voices and generate speech from text locally on your GPU using Qwen3-TTS with an end-to-end training pipeline.
保留图表的转换器 —— DOCX/XLSX 转 Markdown,采用原生 OOXML 图表数据提取(无需光栅化/OCR/VLM),并提供零损耗的复合图表标记。Converters where figures survive — DOCX/XLSX to Markdown with native OOXML chart-data extraction (no rasterize/OCR/VLM) and zero-loss composite-figure markers
基于 Whisper、Gemini、Streamlit、yt-dlp 与 FFmpeg 构建的 AI 应用,可即时转录并总结任意 YouTube 视频。Transcribe and summarize any YouTube video instantly with AI-powered app using Whisper, Gemini, Streamlit, yt-dlp & FFmpeg.
🌐 在数据受限条件下重新思考多模态大语言模型的设计与扩展,以 NaViL 通过 Native Training 提升效率与性能🌐 Rethink Multimodal Large Language Models design and scaling under data constraints with NaViL, enhancing efficiency and performance through Native Training.