Transformers:面向文本、视觉、音频及多模态 SOTA 机器学习模型的模型定义框架,同时支持推理与训练。🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
仓库/Skill 库
7 个 · 多模态 · 模型 · AI 核心
LTX-2 音视频生成模型的官方 Python 推理与 LoRA 训练包。Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.
开源 LLM/VLM 负载均衡器与服务平台,用于规模化自托管 LLM(和 VLM)🏓🦙 作为 llm-d、Docker Model Runner 等项目的替代方案,组件更少、部署更简单,基于 ggml 生态构建。支持 CPU 和 GPU。Open-source LLM/VLM load balancer and serving platform for self-hosting LLMs (and VLMs) at scale 🏓🦙 Alternative to projects like llm-d, Docker Model Runner, etc but with less moving parts and simple deployments built around ggml ecosystem. Runs on CPU and GPU.
自托管、多用户转录平台:录制或上传音频。支持说话人标记、带时间戳的转录、跨录音识别说话人、摘要、提取行动项,并可使用自有 OpenAI 兼容 LLM 与转录内容对话。你的音频、你的服务器、你的模型。已在笔记本 RTX4070、台式机 RTX3090 与 RTX5090 上测试。Self-hosted, multi-user transcription platform: record or upload audio. Speaker-labeled, timestamped transcripts, Recognize speakers across recordings, Summarize, extract action items and chat over your transcripts with your own OpenAI-compatible LLM. Your Audio, your Server, your Model. Tested on Laptop RTX4070, Desktop RTX3090 and RTX5090
使用 vision-language model 处理视觉与文本数据,执行多模态推理与图像理解任务。Process visual and textual data with this vision-language model for multimodal reasoning and image understanding tasks.