研究库 开源仓库
Repositories · organized/repo_cards

仓库/Skill 库

55 个 · 多模态 · AI 核心

排序 Stars 周增
fz-zsl/CAPruner
Python · 2026-09-17 多模态 库 研究原型 Stars 6 周增 +0

ACL 2026 论文 CAPruner 的官方实现:通过概念邻接场景图剪枝增强大语言模型的 3D 空间推理能力The official implementation for ACL 2026 paper CAPruner: Conceptual-Adjacent Scene Graph Pruner for Enhancing 3D Spatial Reasoning of Large Language Models

llm-infra
ZinYY/AdaFlash
Python · 2026-08-11 多模态 应用 实验 Stars 3 周增 +0

[COLMW'26] AdaFlash:通过 On-Policy 蒸馏扩散 Drafters 实现的自适应投机解码[COLMW'26] AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters

llm-infra
pbz134/LocalYT
HTML · 2026-08-29 多模态 库 实验 Stars 3 周增 +0

可自托管且功能丰富的视频库Self-hostable and feature-rich video library

multimodalllm-infra
kenhayward/Diariz
C# · 2026-08-11 多模态 模型 实验 Stars 3 周增 +0

自托管、多用户转录平台:录制或上传音频。支持说话人标记、带时间戳的转录、跨录音识别说话人、摘要、提取行动项,并可使用自有 OpenAI 兼容 LLM 与转录内容对话。你的音频、你的服务器、你的模型。已在笔记本 RTX4070、台式机 RTX3090 与 RTX5090 上测试。Self-hosted, multi-user transcription platform: record or upload audio. Speaker-labeled, timestamped transcripts, Recognize speakers across recordings, Summarize, extract action items and chat over your transcripts with your own OpenAI-compatible LLM. Your Audio, your Server, your Model. Tested on Laptop RTX4070, Desktop RTX3090 and RTX5090

ragdatabasellm-infra
Lauraineabsurd944/Light-WAM
Python · 2026-09-22 多模态 应用 实验 Stars 2 周增 +0

在 LIBERO 与 RoboTwin2.0 上,基于状态融合解码与冻结视频骨干网络训练面向机器人操作的轻量级世界动作模型。Train lightweight world action models for robot manipulation using state-fusion decoding and frozen video backbones on LIBERO and RoboTwin2.0.

ragmultimodalllm-infra
ningkaikok/dotty-tutor
Python · 2026-08-20 多模态 应用 实验 Stars 1 周增 +0

AI 教材数字化与互动辅导平台:PDF/扫描教材经 MinerU OCR 结构化出题,文本 + 视觉双模型审校与确定性质量门禁;学生端七种题型互动、分层提示与错题多轮陪练闭环;FastAPI + PostgreSQL JSONB + React 19

ragdatabasellm-infra
mananp-2730/BridgeBuild-AI-PM-Tool
Python · 2026-08-11 多模态 工具 实验 Stars 1 周增 +0

AI 驱动的企业级 Agile OS,使用 Gemini 1.5 将原始客户音频与笔记即时转化为客户 pitch deck、PM epic、UI 规格与后端工程架构。An AI-powered Enterprise Agile OS that instantly translates raw client audio and notes into client pitch decks, PM epics, UI specs, and backend engineering architectures using Gemini 1.5.

rag
Yashwanth-23/Omnisense
Python · 2026-08-16 多模态 应用 实验 Stars 0 周增 +0

基于 Llama 3.2 的本地多模态 RAG AI 助手,可与 PDF、图片和视频对话,完全离线,零数据外泄。Local multimodal RAG AI assistant where you can chat with PDFs, images, and video using Llama 3.2, fully offline, zero data exposure.

ragmultimodalllm-infra
XenoVoyage/xeno-rover
未知语言 · 2026-08-14 多模态 应用 实验 Stars 0 周增 +0

开源 Raspberry Pi 5 rover 项目,用于建图、视觉、语音、本地 LLM 及液态神经网络实验。Open-source Raspberry Pi 5 rover project for mapping, vision, voice, local LLMs, and liquid-neural-network experiments.

multimodalllm-infra
redocto/image-text-structurizer
Python · 2026-08-22 多模态 工具 实验 Stars 0 周增 +0
rag
nhemrajani/ppg-personalisation
Python · 2026-09-29 多模态 应用 研究原型 Stars 0 周增 +0

Neeharika Hemrajani,耶鲁管理学院与耶鲁大学计算机科学系 2026 年秋季独立研究项目。本项目针对光电容积脉搏波 (PPG) 的基础模型与逐人适配技术进行研究,衡量现有方法在成本与准确率之间的权衡。Neeharika Hemrajani, Independent Study, Yale School of Management and Yale University Department of Computer Science, Fall 2026. The following project is an independent study of foundation models for photoplethysmography (PPG) and per-person adaptation techniques to measure the cost versus accuracy gains across existing methods.

llm-infra
Macr7523/Video-Summarizer
Python · 2026-08-28 多模态 教程 实验 Stars 0 周增 +0

本地使用 AI 总结视频,从讲座、会议和教程中提取视觉亮点与文本摘要,基于 RTX 40 系列 GPU 的 CUDA 加速Summarize videos locally with AI, extracting visual highlights and text summaries from lectures, meetings, and tutorials using CUDA on RTX 40-series GPUs

multimodalllm-infra
lol-dungeonmaster/official-daily-banana
Jupyter Notebook · 2026-09-06 多模态 数据集 实验 Stars 0 周增 +0

每日提示词,激发 nano banana 生成灵感。Daily prompts to inspire nano banana generation.

multimodalllm-infra
kaulzeejai/AIA-Academic-Illustrator-
JavaScript · 2026-10-03 多模态 工具 实验 Stars 0 周增 +0

🎨 利用 GPT、Gemini 等模型,AI 驱动的学术图表一键生成与定制工具🎨 Generate academic diagrams effortlessly with this AI-driven tool, leveraging models like GPT and Gemini for seamless creation and customization.

ragllm-infra
Homologic-bid91/voice_clone_lab
Python · 2026-08-11 多模态 应用 实验 Stars 0 周增 +0

使用 Qwen3-TTS 在本地 GPU 上克隆声音并从文本生成语音,提供端到端训练流水线。Clone voices and generate speech from text locally on your GPU using Qwen3-TTS with an end-to-end training pipeline.

engineeringllm-infra
HelgDemidov/refigure
Python · 2026-08-20 多模态 应用 实验 Stars 0 周增 +0

保留图表的转换器 —— DOCX/XLSX 转 Markdown,采用原生 OOXML 图表数据提取(无需光栅化/OCR/VLM),并提供零损耗的复合图表标记。Converters where figures survive — DOCX/XLSX to Markdown with native OOXML chart-data extraction (no rasterize/OCR/VLM) and zero-loss composite-figure markers

ragllm-infra
Erikalaylafajri15/MOSS-VL
未知语言 · 2026-09-22 多模态 模型 实验 Stars 0 周增 +0

使用 vision-language model 处理视觉与文本数据,执行多模态推理与图像理解任务。Process visual and textual data with this vision-language model for multimodal reasoning and image understanding tasks.

multimodalllm-infra
Davimoren9040/Youtube-Video-Transcribe-Summarizer-LLM-App
未知语言 · 2026-10-02 多模态 应用 实验 Stars 0 周增 +0

基于 Whisper、Gemini、Streamlit、yt-dlp 与 FFmpeg 构建的 AI 应用,可即时转录并总结任意 YouTube 视频。Transcribe and summarize any YouTube video instantly with AI-powered app using Whisper, Gemini, Streamlit, yt-dlp & FFmpeg.

multimodalllm-infra
ajgarciaj/NaViL
未知语言 · 2026-09-19 多模态 应用 实验 Stars 0 周增 +0

🌐 在数据受限条件下重新思考多模态大语言模型的设计与扩展,以 NaViL 通过 Native Training 提升效率与性能🌐 Rethink Multimodal Large Language Models design and scaling under data constraints with NaViL, enhancing efficiency and performance through Native Training.

agentragmultimodalllm-infra