Repositories · organized/repo_cards

仓库/Skill 库

48 个 · 多模态 · AI 核心

排序 Stars 周增
mananp-2730/BridgeBuild-AI-PM-Tool
Python · 2026-08-11 多模态 工具 实验 Stars 1 周增 +0

AI 驱动的企业级 Agile OS,使用 Gemini 1.5 将原始客户音频与笔记即时转化为客户 pitch deck、PM epic、UI 规格与后端工程架构。An AI-powered Enterprise Agile OS that instantly translates raw client audio and notes into client pitch decks, PM epics, UI specs, and backend engineering architectures using Gemini 1.5.

rag
Yashwanth-23/Omnisense
Python · 2026-08-16 多模态 应用 实验 Stars 0 周增 +0

基于 Llama 3.2 的本地多模态 RAG AI 助手,可与 PDF、图片和视频对话,完全离线,零数据外泄。Local multimodal RAG AI assistant where you can chat with PDFs, images, and video using Llama 3.2, fully offline, zero data exposure.

ragmultimodalllm-infra
XenoVoyage/xeno-rover
未知语言 · 2026-08-14 多模态 应用 实验 Stars 0 周增 +0

开源 Raspberry Pi 5 rover 项目,用于建图、视觉、语音、本地 LLM 及液态神经网络实验。Open-source Raspberry Pi 5 rover project for mapping, vision, voice, local LLMs, and liquid-neural-network experiments.

multimodalllm-infra
redocto/image-text-structurizer
Python · 2026-08-22 多模态 工具 实验 Stars 0 周增 +0
rag
Macr7523/Video-Summarizer
Python · 2026-08-11 多模态 教程 实验 Stars 0 周增 +0

本地使用 AI 总结视频,从讲座、会议和教程中提取视觉亮点与文本摘要,基于 RTX 40 系列 GPU 的 CUDA 加速Summarize videos locally with AI, extracting visual highlights and text summaries from lectures, meetings, and tutorials using CUDA on RTX 40-series GPUs

multimodalllm-infra
lol-dungeonmaster/official-daily-banana
Jupyter Notebook · 2026-08-24 多模态 数据集 实验 Stars 0 周增 +0

每日提示词,激发 nano banana 生成灵感。Daily prompts to inspire nano banana generation.

multimodalllm-infra
kaulzeejai/AIA-Academic-Illustrator-
JavaScript · 2026-08-13 多模态 工具 实验 Stars 0 周增 +0

🎨 利用 GPT、Gemini 等模型,AI 驱动的学术图表一键生成与定制工具🎨 Generate academic diagrams effortlessly with this AI-driven tool, leveraging models like GPT and Gemini for seamless creation and customization.

ragllm-infra
Homologic-bid91/voice_clone_lab
Python · 2026-08-11 多模态 应用 实验 Stars 0 周增 +0

使用 Qwen3-TTS 在本地 GPU 上克隆声音并从文本生成语音,提供端到端训练流水线。Clone voices and generate speech from text locally on your GPU using Qwen3-TTS with an end-to-end training pipeline.

engineeringllm-infra
HelgDemidov/refigure
Python · 2026-08-20 多模态 应用 实验 Stars 0 周增 +0

保留图表的转换器 —— DOCX/XLSX 转 Markdown,采用原生 OOXML 图表数据提取(无需光栅化/OCR/VLM),并提供零损耗的复合图表标记。Converters where figures survive — DOCX/XLSX to Markdown with native OOXML chart-data extraction (no rasterize/OCR/VLM) and zero-loss composite-figure markers

ragllm-infra
Erikalaylafajri15/MOSS-VL
未知语言 · 2026-08-11 多模态 模型 实验 Stars 0 周增 +0

使用 vision-language model 处理视觉与文本数据,执行多模态推理与图像理解任务。Process visual and textual data with this vision-language model for multimodal reasoning and image understanding tasks.

multimodalllm-infra
Davimoren9040/Youtube-Video-Transcribe-Summarizer-LLM-App
未知语言 · 2026-08-18 多模态 应用 实验 Stars 0 周增 +0

基于 Whisper、Gemini、Streamlit、yt-dlp 与 FFmpeg 构建的 AI 应用,可即时转录并总结任意 YouTube 视频。Transcribe and summarize any YouTube video instantly with AI-powered app using Whisper, Gemini, Streamlit, yt-dlp & FFmpeg.

multimodalllm-infra
ajgarciaj/NaViL
未知语言 · 2026-08-15 多模态 应用 实验 Stars 0 周增 +0

🌐 在数据受限条件下重新思考多模态大语言模型的设计与扩展,以 NaViL 通过 Native Training 提升效率与性能🌐 Rethink Multimodal Large Language Models design and scaling under data constraints with NaViL, enhancing efficiency and performance through Native Training.

agentragmultimodalllm-infra