研究库 开源仓库
Repositories · organized/repo_cards

仓库/Skill 库

112 个 · 多模态 · 应用

排序 Stars 周增
matlowai/ComfyUI-MAINodes
Python · 2026-08-14 多模态 应用 实验 Stars 54 周增 +0

MatlowAI 的 MiniMax-H3 ComfyUI 节点:Contact-Sheet diffusion + Motion Lab(针对快速运动的 test-time 去绳状畸变)MatlowAI's MiniMax-H3 ComfyUI nodes: Contact-Sheet diffusion + Motion Lab (test-time de-roping of fast motion)

zlab-princeton/VisionFoundry
Python · 2026-09-08 多模态 应用 实验 Stars 52 周增 +0

VisionFoundry:使用合成图像教会 VLM 视觉感知。VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images

multimodalllm-infra
itrcz/calab
TypeScript · 2026-10-02 多模态 应用 实验 Stars 52 周增 +0

受 Discord、Telegram 和 Zoom 启发的开源团队聊天和视频会议工具。OpenSource team chats and video conferences inspired with Discord, Telegram & Zoom

multimodal
xmutfyh/dsh-plugin-writing-guard
JavaScript · 2026-10-05 多模态 应用 实验 Stars 45 周增 +3

面向 AI 辅助研究的科学写作与文档完整性守护——论证经济性、确定性完整性、安全的 Word 编辑。本地 · 确定性 · 零 LLM。Scientific writing & document integrity guard for AI-assisted research — argument economy, deterministic integrity, safe Word editing. Local · Deterministic · Zero LLM.

llm-infra
Syh1906/openai-compatible-imagegen
JavaScript · 2026-08-21 多模态 应用 实验 Stars 34 周增 +0

独立的 Skill 与 Codex 插件,支持 OpenAI 兼容的图像生成、编辑、批处理工作流、QA 以及聚焦画布编辑。Standalone Skill and Codex Plugin for OpenAI-compatible image generation, editing, batch workflows, QA, and focused canvas editing.

agentmultimodal
yuanzh0u/embodied-learning
HTML · 2026-09-28 多模态 应用 实验 Stars 7 周增 +0

证据优先的具身 AI 研究枢纽,覆盖 VLA、世界模型、多模态感知、文献综述、声明审计以及可复用的 Agent skillEvidence-first Embodied AI research hub for VLA, world models, multimodal sensing, literature reviews, claim audits, and reusable agent skills.

agentmultimodal
HanQingHub/PaperLens
TypeScript · 2026-10-06 多模态 应用 研究原型 Stars 5 周增 +0

Local-first 学术 PDF 阅读器:离线 CPU LLM 翻译与注释、OCR、ECDICT 词典。Tauri + FastAPI。Local-first academic PDF reader: offline CPU LLM translation & glossing, OCR, ECDICT dictionary. Tauri + FastAPI.

llm-infra
ZinYY/AdaFlash
Python · 2026-08-11 多模态 应用 实验 Stars 3 周增 +0

[COLMW'26] AdaFlash:通过 On-Policy 蒸馏扩散 Drafters 实现的自适应投机解码[COLMW'26] AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters

llm-infra
Sunrich-HT/authorship-aware-writing
Python · 2026-08-19 多模态 应用 实验 Stars 2 周增 +7

作者感知的科学写作分析:作者空间、目标论文风格、校准的生成证据,以及安全合规的修订Authorship-aware scientific writing analysis: author space, target-paper style, calibrated generation evidence, and integrity-safe revision.

multimodal
zhenhuanZ/DIA-Review
未知语言 · 2026-08-13 多模态 应用 研究原型 Stars 2 周增 +0

论文《Deep Learning in Dental Image Analysis: A Systematic Review of Datasets, Methodologies, and Emerging Challenges》的官方项目主页。The official project homepage for the paper "Deep Learning in Dental Image Analysis: A Systematic Review of Datasets, Methodologies, and Emerging Challenges"

multimodal
yanlingLabs/video-extract-mcp
TypeScript · 2026-08-20 多模态 应用 实验 Stars 2 周增 +0

MCP server,可从 URL 提取任意视频(包括微信视频号)。可选择性地分析并提取场景感知的去重关键帧,并对视频进行转录。完全本地运行MCP server that extracts any video from URL(inc/ WeChat channels). Optionally analyzes and extracts scene-aware deduplicated key frames and transcribes the video. ALL LOCAL

agentmultimodalllm-infra
Lauraineabsurd944/Light-WAM
Python · 2026-09-22 多模态 应用 实验 Stars 2 周增 +0

在 LIBERO 与 RoboTwin2.0 上,基于状态融合解码与冻结视频骨干网络训练面向机器人操作的轻量级世界动作模型。Train lightweight world action models for robot manipulation using state-fusion decoding and frozen video backbones on LIBERO and RoboTwin2.0.

ragmultimodalllm-infra
jongan69/YouTubeResearchAI
JavaScript · 2026-08-14 多模态 应用 实验 Stars 2 周增 +0

将任意视频链接转化为博士级研究报告。自动化学术流水线:一键完成下载、转写、检索同行评审文献、核查论据并生成带引用的报告。免费且开源。Turn any video url into a PhD-grade research report. Automated academic pipeline: download, transcribe, search peer-reviewed literature, verify claims, generate cited reports — all with one command. Free & open source.

multimodalengineeringllm-infra
ningkaikok/dotty-tutor
Python · 2026-08-20 多模态 应用 实验 Stars 1 周增 +0

AI 教材数字化与互动辅导平台:PDF/扫描教材经 MinerU OCR 结构化出题,文本 + 视觉双模型审校与确定性质量门禁;学生端七种题型互动、分层提示与错题多轮陪练闭环;FastAPI + PostgreSQL JSONB + React 19

ragdatabasellm-infra
msamribeiro/speech-generation-wiki
未知语言 · 2026-09-12 多模态 应用 实验 Stars 1 周增 +0

TTS、语音转换与口语对话 Agent 的活体系统综述Living systematic review of TTS, voice conversion, and spoken conversational agents

agent
fuliao603/paper-reader
JavaScript · 2026-09-29 多模态 应用 实验 Stars 1 周增 +0

面向学术阅读的本地桌面 PDF 阅读器,支持 AI 翻译、OCR、高亮、笔记、翻译历史以及导入/导出。A local desktop PDF reader for academic reading, with AI translation, OCR, highlights, notes, translation history, and import/export support.

Fatma3598/EEG-Imagined-Speech-Preprocessing-Generalization
Jupyter Notebook · 2026-08-12 多模态 应用 实验 Stars 1 周增 +0

《Comparative Analysis of EEG Signal Preprocessing Impact on Generalization of a Transformer-Based EEG Imagined Speech Recognition》研究的可复现研究仓库。Reproducible research repository for 'Comparative Analysis of EEG Signal Preprocessing Impact on Generalization of a Transformer-Based EEG Imagined Speech Recognition' research work

Yashwanth-23/Omnisense
Python · 2026-08-16 多模态 应用 实验 Stars 0 周增 +0

基于 Llama 3.2 的本地多模态 RAG AI 助手,可与 PDF、图片和视频对话,完全离线,零数据外泄。Local multimodal RAG AI assistant where you can chat with PDFs, images, and video using Llama 3.2, fully offline, zero data exposure.

ragmultimodalllm-infra
XenoVoyage/xeno-rover
未知语言 · 2026-08-14 多模态 应用 实验 Stars 0 周增 +0

开源 Raspberry Pi 5 rover 项目,用于建图、视觉、语音、本地 LLM 及液态神经网络实验。Open-source Raspberry Pi 5 rover project for mapping, vision, voice, local LLMs, and liquid-neural-network experiments.

multimodalllm-infra
wozengyi/embodied-daily
JavaScript · 2026-10-09 多模态 应用 实验 Stars 0 周增 +0

Embodied Daily:每日具身智能 AI 论文推荐(HF Daily Papers + arXiv)Embodied Daily: daily embodied-AI paper recommendations (HF Daily Papers + arXiv)

seasalim/happy-photon
C# · 2026-08-13 多模态 应用 生产可用 Stars 0 周增 +0

更简洁的照片编辑体验;开源 RAW 与 JPEG 流水线。Photo editing, simplified. An open-source RAW and JPEG workflow.

outrotim/evidence-aligned-scientific-revision
Python · 2026-08-17 多模态 应用 实验 Stars 0 周增 +0

面向中英双语生物医学研究的术语循证、主张强度与科学写作审计 skill。Evidence-aligned terminology, claim-strength, and scientific-writing audit skill for Chinese and English biomedical research.

agent
nhemrajani/ppg-personalisation
Python · 2026-09-29 多模态 应用 研究原型 Stars 0 周增 +0

Neeharika Hemrajani,耶鲁管理学院与耶鲁大学计算机科学系 2026 年秋季独立研究项目。本项目针对光电容积脉搏波 (PPG) 的基础模型与逐人适配技术进行研究,衡量现有方法在成本与准确率之间的权衡。Neeharika Hemrajani, Independent Study, Yale School of Management and Yale University Department of Computer Science, Fall 2026. The following project is an independent study of foundation models for photoplethysmography (PPG) and per-person adaptation techniques to measure the cost versus accuracy gains across existing methods.

llm-infra
matguo/Phyadv_pipeline
未知语言 · 2026-08-12 多模态 应用 实验 Stars 0 周增 +0

本仓库包含论文补充材料,旨在提升系统综述的透明度、可复现性与完整性,涵盖筛选文档、检索策略以及计算机视觉中物理对抗攻击的详细分析。This repository contains the supplementary materials accompanying our paper. The materials are designed to support the transparency, reproducibility, and comprehensiveness of the systematic review, including screening documentation, search strategies, and detailed analysis of physical adversarial attacks in computer vision.

multimodalrisk
littlecookie0722/wairc-2026
Python · 2026-08-23 多模态 应用 实验 Stars 0 周增 +0

基于 IQ 信号、采用 STFT 频谱图、视觉模型、k-fold 训练与集成推理的多节点 RF 无人机识别可复现研究工具包。A reproducible research toolkit for multi-node RF drone identification from IQ signals using STFT spectrograms, vision models, k-fold training, and ensemble inference.

multimodalllm-infra
Lacenedihia/Google-Search-Ranking-Discoverability-Capstone
HTML · 2026-10-05 多模态 应用 实验 Stars 0 周增 +0

研究问题与暂定方向Research Question and Provisional Lane

multimodal
jakemorgan-research/codex-sci-research-lifecycle
Python · 2026-08-22 多模态 应用 实验 Stars 0 周增 +0

可审计的 Codex skill 与 Python 项目包,用于证据可追溯的研究、系统综述、投稿与修订。Auditable Codex skill and Python project pack for evidence-traceable research, systematic reviews, submission, and revision.

multimodal
Homologic-bid91/voice_clone_lab
Python · 2026-08-11 多模态 应用 实验 Stars 0 周增 +0

使用 Qwen3-TTS 在本地 GPU 上克隆声音并从文本生成语音,提供端到端训练流水线。Clone voices and generate speech from text locally on your GPU using Qwen3-TTS with an end-to-end training pipeline.

engineeringllm-infra
HelgDemidov/refigure
Python · 2026-08-20 多模态 应用 实验 Stars 0 周增 +0

保留图表的转换器 —— DOCX/XLSX 转 Markdown,采用原生 OOXML 图表数据提取(无需光栅化/OCR/VLM),并提供零损耗的复合图表标记。Converters where figures survive — DOCX/XLSX to Markdown with native OOXML chart-data extraction (no rasterize/OCR/VLM) and zero-loss composite-figure markers

ragllm-infra
Harryphan72007/minh-phan-portfolio
TypeScript · 2026-08-26 多模态 应用 实验 Stars 0 周增 +0

Minh Phan 作品集:机器学习系统、计算机视觉、可复现研究与软件工程Minh Phan portfolio: ML systems, computer vision, reproducible research, and software engineering.

multimodal
hanjiarui1025-a11y/rebuttal-revision
未知语言 · 2026-08-16 多模态 应用 实验 Stars 0 周增 +0

Codex Skill:学术期刊 rebuttal 与修改工作流Codex skill for academic journal rebuttal and revision workflows.

multimodal
gong8/paper-reader
HTML · 2026-08-25 多模态 应用 实验 Stars 0 周增 +0

学术 PDF 阅读器:在浏览器中并排展示印刷版页面与重排阅读视图,由同一模型解析并联动。An academic PDF reader: the printed page and a reflowed reading view, side by side, linked by one model parsed in your browser.

galaxy99881/galaxy99881
未知语言 · 2026-08-16 多模态 应用 实验 Stars 0 周增 +0

Zhixin Li 的学术主页:超声医学、医学影像 AI、多模态学习与可复现研究。Academic homepage of Zhixin Li: ultrasound medicine, medical imaging AI, multimodal learning, and reproducible research.

multimodal
fhlyongko/graduate-academic-writing-ebook-kr
TypeScript · 2026-08-09 多模态 应用 实验 Stars 0 周增 +0

面向研究生学术英语写作、研究设计、伦理、修改与发表的韩文交互式现场指南Interactive Korean field guide for graduate academic English writing, research design, ethics, revision, and publication

multimodal
docxology/DuckRabbit
Python · 2026-09-27 多模态 应用 实验 Stars 0 周增 +0

Python 包,作为可复现的科学刺激生成视觉、听觉和视听错觉 —— 17 个已编目错觉导出为 PNG、WAV、GIF、MP4 和 NPZ,附带 SHA-256 清单与确定性随机种子,含 15 个附带源数据边车的发表级图表,测试覆盖率 90% 以上Python package that generates optical, auditory and audio-visual illusions as reproducible scientific stimuli — 17 catalogued illusions exported to PNG, WAV, GIF, MP4 and NPZ with SHA-256 manifests, deterministic seeds, 15 publication figures with source-data sidecars, and 90%+ test coverage.

ragmultimodal
Davimoren9040/Youtube-Video-Transcribe-Summarizer-LLM-App
未知语言 · 2026-10-02 多模态 应用 实验 Stars 0 周增 +0

基于 Whisper、Gemini、Streamlit、yt-dlp 与 FFmpeg 构建的 AI 应用,可即时转录并总结任意 YouTube 视频。Transcribe and summarize any YouTube video instantly with AI-powered app using Whisper, Gemini, Streamlit, yt-dlp & FFmpeg.

multimodalllm-infra