研究库 开源仓库
Repositories · organized/repo_cards

仓库/Skill 库

236 个 · 多模态

排序 Stars 周增
fz-zsl/CAPruner
Python · 2026-09-17 多模态 库 研究原型 Stars 6 周增 +0

ACL 2026 论文 CAPruner 的官方实现:通过概念邻接场景图剪枝增强大语言模型的 3D 空间推理能力The official implementation for ACL 2026 paper CAPruner: Conceptual-Adjacent Scene Graph Pruner for Enhancing 3D Spatial Reasoning of Large Language Models

llm-infra
HanQingHub/PaperLens
TypeScript · 2026-10-06 多模态 应用 研究原型 Stars 5 周增 +0

Local-first 学术 PDF 阅读器:离线 CPU LLM 翻译与注释、OCR、ECDICT 词典。Tauri + FastAPI。Local-first academic PDF reader: offline CPU LLM translation & glossing, OCR, ECDICT dictionary. Tauri + FastAPI.

llm-infra
ZinYY/AdaFlash
Python · 2026-08-11 多模态 应用 实验 Stars 3 周增 +0

[COLMW'26] AdaFlash:通过 On-Policy 蒸馏扩散 Drafters 实现的自适应投机解码[COLMW'26] AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters

llm-infra
pbz134/LocalYT
HTML · 2026-08-29 多模态 库 实验 Stars 3 周增 +0

可自托管且功能丰富的视频库Self-hostable and feature-rich video library

multimodalllm-infra
kenhayward/Diariz
C# · 2026-08-11 多模态 模型 实验 Stars 3 周增 +0

自托管、多用户转录平台:录制或上传音频。支持说话人标记、带时间戳的转录、跨录音识别说话人、摘要、提取行动项,并可使用自有 OpenAI 兼容 LLM 与转录内容对话。你的音频、你的服务器、你的模型。已在笔记本 RTX4070、台式机 RTX3090 与 RTX5090 上测试。Self-hosted, multi-user transcription platform: record or upload audio. Speaker-labeled, timestamped transcripts, Recognize speakers across recordings, Summarize, extract action items and chat over your transcripts with your own OpenAI-compatible LLM. Your Audio, your Server, your Model. Tested on Laptop RTX4070, Desktop RTX3090 and RTX5090

ragdatabasellm-infra
Sunrich-HT/authorship-aware-writing
Python · 2026-08-19 多模态 应用 实验 Stars 2 周增 +7

作者感知的科学写作分析:作者空间、目标论文风格、校准的生成证据,以及安全合规的修订Authorship-aware scientific writing analysis: author space, target-paper style, calibrated generation evidence, and integrity-safe revision.

multimodal
zhenhuanZ/DIA-Review
未知语言 · 2026-08-13 多模态 应用 研究原型 Stars 2 周增 +0

论文《Deep Learning in Dental Image Analysis: A Systematic Review of Datasets, Methodologies, and Emerging Challenges》的官方项目主页。The official project homepage for the paper "Deep Learning in Dental Image Analysis: A Systematic Review of Datasets, Methodologies, and Emerging Challenges"

multimodal
yanlingLabs/video-extract-mcp
TypeScript · 2026-08-20 多模态 应用 实验 Stars 2 周增 +0

MCP server,可从 URL 提取任意视频(包括微信视频号)。可选择性地分析并提取场景感知的去重关键帧,并对视频进行转录。完全本地运行MCP server that extracts any video from URL(inc/ WeChat channels). Optionally analyzes and extracts scene-aware deduplicated key frames and transcribes the video. ALL LOCAL

agentmultimodalllm-infra
Lauraineabsurd944/Light-WAM
Python · 2026-09-22 多模态 应用 实验 Stars 2 周增 +0

在 LIBERO 与 RoboTwin2.0 上,基于状态融合解码与冻结视频骨干网络训练面向机器人操作的轻量级世界动作模型。Train lightweight world action models for robot manipulation using state-fusion decoding and frozen video backbones on LIBERO and RoboTwin2.0.

ragmultimodalllm-infra
jongan69/YouTubeResearchAI
JavaScript · 2026-08-14 多模态 应用 实验 Stars 2 周增 +0

将任意视频链接转化为博士级研究报告。自动化学术流水线:一键完成下载、转写、检索同行评审文献、核查论据并生成带引用的报告。免费且开源。Turn any video url into a PhD-grade research report. Automated academic pipeline: download, transcribe, search peer-reviewed literature, verify claims, generate cited reports — all with one command. Free & open source.

multimodalengineeringllm-infra
attavit14203638/transformer-tree-survey
Python · 2026-08-20 多模态 收藏榜 研究原型 Stars 2 周增 +0

论文《Transformer-Based Tree Extraction from Remote Sensing Imagery: A Systematic Review》的补充代码库。Supplementary repository for: Transformer-Based Tree Extraction from Remote Sensing Imagery: A Systematic Review

multimodal
ningkaikok/dotty-tutor
Python · 2026-08-20 多模态 应用 实验 Stars 1 周增 +0

AI 教材数字化与互动辅导平台:PDF/扫描教材经 MinerU OCR 结构化出题,文本 + 视觉双模型审校与确定性质量门禁;学生端七种题型互动、分层提示与错题多轮陪练闭环;FastAPI + PostgreSQL JSONB + React 19

ragdatabasellm-infra
msamribeiro/speech-generation-wiki
未知语言 · 2026-09-12 多模态 应用 实验 Stars 1 周增 +0

TTS、语音转换与口语对话 Agent 的活体系统综述Living systematic review of TTS, voice conversion, and spoken conversational agents

agent
mananp-2730/BridgeBuild-AI-PM-Tool
Python · 2026-08-11 多模态 工具 实验 Stars 1 周增 +0

AI 驱动的企业级 Agile OS,使用 Gemini 1.5 将原始客户音频与笔记即时转化为客户 pitch deck、PM epic、UI 规格与后端工程架构。An AI-powered Enterprise Agile OS that instantly translates raw client audio and notes into client pitch decks, PM epics, UI specs, and backend engineering architectures using Gemini 1.5.

rag
fuliao603/paper-reader
JavaScript · 2026-09-29 多模态 应用 实验 Stars 1 周增 +0

面向学术阅读的本地桌面 PDF 阅读器,支持 AI 翻译、OCR、高亮、笔记、翻译历史以及导入/导出。A local desktop PDF reader for academic reading, with AI translation, OCR, highlights, notes, translation history, and import/export support.

Fatma3598/EEG-Imagined-Speech-Preprocessing-Generalization
Jupyter Notebook · 2026-08-12 多模态 应用 实验 Stars 1 周增 +0

《Comparative Analysis of EEG Signal Preprocessing Impact on Generalization of a Transformer-Based EEG Imagined Speech Recognition》研究的可复现研究仓库。Reproducible research repository for 'Comparative Analysis of EEG Signal Preprocessing Impact on Generalization of a Transformer-Based EEG Imagined Speech Recognition' research work

choxos/NMAViz
JavaScript · 2026-09-07 多模态 模型 实验 Stars 1 周增 +0

阅读网络 meta 分析的证据结构。在浏览器中拟合图论模型并以八种方式绘制:精度几何、试验、证据流、贡献、带符号的研究级重构、不一致性的 Hodge 分解、扩散与弹簧图。Read the evidence structure of a network meta-analysis. Fits the graph-theoretical model in your browser and draws it eight ways: precision geometry, trials, evidence flow, contributions, a signed study-level reconstruction, a Hodge split of inconsistency, diffusion and springs.

aidadbeni/Computational-Pathology
Jupyter Notebook · 2026-08-19 多模态 收藏榜 研究原型 Stars 1 周增 +0

围绕病理 foundation models(尤其是 Prov-GigaPath)的文献综述与可复现性研究Literature review and reproducibility work on pathology foundation models, with a focus on Prov-GigaPath.

Yashwanth-23/Omnisense
Python · 2026-08-16 多模态 应用 实验 Stars 0 周增 +0

基于 Llama 3.2 的本地多模态 RAG AI 助手,可与 PDF、图片和视频对话,完全离线,零数据外泄。Local multimodal RAG AI assistant where you can chat with PDFs, images, and video using Llama 3.2, fully offline, zero data exposure.

ragmultimodalllm-infra
XenoVoyage/xeno-rover
未知语言 · 2026-08-14 多模态 应用 实验 Stars 0 周增 +0

开源 Raspberry Pi 5 rover 项目,用于建图、视觉、语音、本地 LLM 及液态神经网络实验。Open-source Raspberry Pi 5 rover project for mapping, vision, voice, local LLMs, and liquid-neural-network experiments.

multimodalllm-infra
wozengyi/embodied-daily
JavaScript · 2026-10-09 多模态 应用 实验 Stars 0 周增 +0

Embodied Daily:每日具身智能 AI 论文推荐(HF Daily Papers + arXiv)Embodied Daily: daily embodied-AI paper recommendations (HF Daily Papers + arXiv)

tobias-weiss-ai-xr/robotics-research
Python · 2026-08-31 多模态 数据集 实验 Stars 0 周增 +0

数据驱动、自动验证的机器人研究文献综述语料库(manipulation、locomotion、perception、planning、learning、HRI、multi-robot、simulation)。Robotics Research Corpus: Data-driven, auto-validated literature review for robotics research (manipulation, locomotion, perception, planning, learning, HRI, multi-robot, simulation)

shokrydev/master
Python · 2026-09-09 多模态 教程 实验 Stars 0 周增 +0

本仓库最初是我毕设项目的 PyTorch Lightning 模板,后来演进为包含完整项目实现本身。研究问题是地理定位如何影响遥感视觉-语言模型的训练与推理。注:仓库名称与描述可能变化。This repo originally was a PyTorch Lightning template for my thesis project. It has since evolved to contain the full project implementation itself. The research question is how geolocation influences the training and inference of vision-language models for remote sensing. Note: The repository name and description may change

multimodalllm-infra
seasalim/happy-photon
C# · 2026-08-13 多模态 应用 生产可用 Stars 0 周增 +0

更简洁的照片编辑体验;开源 RAW 与 JPEG 流水线。Photo editing, simplified. An open-source RAW and JPEG workflow.

redocto/image-text-structurizer
Python · 2026-08-22 多模态 工具 实验 Stars 0 周增 +0
rag
outrotim/ideal-to-real-publication-figure
未知语言 · 2026-08-24 多模态 工具 实验 Stars 0 周增 +0

Codex skill:将已批准的科研图表蓝图转化为可溯源、使用真实数据的发表级图表。Codex skill for turning approved scientific figure blueprints into traceable real-data publication figures

outrotim/evidence-aligned-scientific-revision
Python · 2026-08-17 多模态 应用 实验 Stars 0 周增 +0

面向中英双语生物医学研究的术语循证、主张强度与科学写作审计 skill。Evidence-aligned terminology, claim-strength, and scientific-writing audit skill for Chinese and English biomedical research.

agent
no-problem-dev/swift-latex-view
Swift · 2026-08-11 多模态 库 实验 Stars 0 周增 +0

SwiftUI 原生 LaTeX 数学渲染,对 LLM 输出具有鲁棒性——支持分隔符归一化、流式渲染与自动主题适配SwiftUI-native LaTeX math rendering robust to LLM output — delimiter normalization, streaming support, and automatic theming

llm-infraengineering
nhemrajani/ppg-personalisation
Python · 2026-09-29 多模态 应用 研究原型 Stars 0 周增 +0

Neeharika Hemrajani,耶鲁管理学院与耶鲁大学计算机科学系 2026 年秋季独立研究项目。本项目针对光电容积脉搏波 (PPG) 的基础模型与逐人适配技术进行研究,衡量现有方法在成本与准确率之间的权衡。Neeharika Hemrajani, Independent Study, Yale School of Management and Yale University Department of Computer Science, Fall 2026. The following project is an independent study of foundation models for photoplethysmography (PPG) and per-person adaptation techniques to measure the cost versus accuracy gains across existing methods.

llm-infra
matguo/Phyadv_pipeline
未知语言 · 2026-08-12 多模态 应用 实验 Stars 0 周增 +0

本仓库包含论文补充材料,旨在提升系统综述的透明度、可复现性与完整性,涵盖筛选文档、检索策略以及计算机视觉中物理对抗攻击的详细分析。This repository contains the supplementary materials accompanying our paper. The materials are designed to support the transparency, reproducibility, and comprehensiveness of the systematic review, including screening documentation, search strategies, and detailed analysis of physical adversarial attacks in computer vision.

multimodalrisk
Macr7523/Video-Summarizer
Python · 2026-08-28 多模态 教程 实验 Stars 0 周增 +0

本地使用 AI 总结视频,从讲座、会议和教程中提取视觉亮点与文本摘要,基于 RTX 40 系列 GPU 的 CUDA 加速Summarize videos locally with AI, extracting visual highlights and text summaries from lectures, meetings, and tutorials using CUDA on RTX 40-series GPUs

multimodalllm-infra
lol-dungeonmaster/official-daily-banana
Jupyter Notebook · 2026-09-06 多模态 数据集 实验 Stars 0 周增 +0

每日提示词,激发 nano banana 生成灵感。Daily prompts to inspire nano banana generation.

multimodalllm-infra
littlecookie0722/wairc-2026
Python · 2026-08-23 多模态 应用 实验 Stars 0 周增 +0

基于 IQ 信号、采用 STFT 频谱图、视觉模型、k-fold 训练与集成推理的多节点 RF 无人机识别可复现研究工具包。A reproducible research toolkit for multi-node RF drone identification from IQ signals using STFT spectrograms, vision models, k-fold training, and ensemble inference.

multimodalllm-infra
Lacenedihia/Google-Search-Ranking-Discoverability-Capstone
HTML · 2026-10-05 多模态 应用 实验 Stars 0 周增 +0

研究问题与暂定方向Research Question and Provisional Lane

multimodal
kaulzeejai/AIA-Academic-Illustrator-
JavaScript · 2026-10-03 多模态 工具 实验 Stars 0 周增 +0

🎨 利用 GPT、Gemini 等模型,AI 驱动的学术图表一键生成与定制工具🎨 Generate academic diagrams effortlessly with this AI-driven tool, leveraging models like GPT and Gemini for seamless creation and customization.

ragllm-infra
karim1988781/agrirobotic_imaging
Jupyter Notebook · 2026-08-22 多模态 数据集 研究原型 Stars 0 周增 +0

面向农业机器人成像系统综述与 Meta 分析的数据与分析代码。Data and analysis code for a systematic review and meta-analysis of agricultural robotic imaging