Repositories · organized/repo_cards

仓库/Skill 库

167 个 · 多模态

排序 Stars 周增
jlucasmcrell/ComfyUI-H3-Multishot
Python · 2026-08-11 多模态 工具 实验 Stars 62 周增 +0
horvitzs/Interactive_Segmentation_Models
未知语言 · 2018-05-18 多模态 收藏榜 研究原型 Stars 58 周增 +0

交互式标注/分割的文献综述Literature Review for Interactive annotation/segmentation

multimodal
youkely/awesome-visual-localization
未知语言 · 2022-07-12 多模态 收藏榜 研究原型 Stars 57 周增 +0

视觉定位的文献综述Literature review of visual localization.

ragmultimodal
matlowai/ComfyUI-MAINodes
Python · 2026-08-14 多模态 应用 实验 Stars 54 周增 +0

MatlowAI 的 MiniMax-H3 ComfyUI 节点:Contact-Sheet diffusion + Motion Lab(针对快速运动的 test-time 去绳状畸变)MatlowAI's MiniMax-H3 ComfyUI nodes: Contact-Sheet diffusion + Motion Lab (test-time de-roping of fast motion)

YZCU/OOTB
C · 2025-07-11 多模态 评测集 实验 Stars 53 周增 +0

[ISPRS 2024] 卫星视频单目标跟踪:系统综述与定向目标跟踪基准[ISPRS 2024] Satellite Video Single Object Tracking: A Systematic Review and An Oriented Object Tracking Benchmark

multimodalevaluation
xulihang/ImageTrans_plugins
B4X · 2026-08-11 多模态 实验 Stars 36 周增 +0

ImageTrans 插件集合。Plugins for ImageTrans

multimodalllm-infra
Syh1906/openai-compatible-imagegen
JavaScript · 2026-08-21 多模态 应用 实验 Stars 34 周增 +0

独立的 Skill 与 Codex 插件,支持 OpenAI 兼容的图像生成、编辑、批处理工作流、QA 以及聚焦画布编辑。Standalone Skill and Codex Plugin for OpenAI-compatible image generation, editing, batch workflows, QA, and focused canvas editing.

agentmultimodal
xmutfyh/dsh-plugin-writing-guard
JavaScript · 2026-08-24 多模态 应用 实验 Stars 22 周增 +13

DeepSeek Harness (DSH) 论文写作护栏——AI 写作风格检测、证据保留、期刊匹配校准、手稿校对、writing_audit 与自动检查。本地运行,零网络,零 LLM。DeepSeek Harness (DSH) academic writing guard for papers — 论文去AI味 / AI-writing style detection, evidence preservation, journal-fit calibration, manuscript proofreading, writing_audit & automatic checks. Local, zero network, zero LLM.

llm-infraengineering
WenyuChiou/academic-writing-skills
Python · 2026-08-11 多模态 应用 实验 Stars 11 周增 +0

面向严谨学术论文写作、修订与投稿的 Claude Code skill。跨领域通用,支持按论文设置期刊覆盖规则。Claude Code skill for rigorous academic paper writing, revision, and submission. Field-agnostic with per-paper journal overrides.

multimodal
HanQingHub/PaperLens
Python · 2026-08-25 多模态 应用 研究原型 Stars 5 周增 +0

Local-first 学术 PDF 阅读器:离线 CPU LLM 翻译与注释、OCR、ECDICT 词典。Tauri + FastAPI。Local-first academic PDF reader: offline CPU LLM translation & glossing, OCR, ECDICT dictionary. Tauri + FastAPI.

llm-infra
yuanzh0u/embodied-learning
Python · 2026-08-18 多模态 应用 实验 Stars 4 周增 +0

证据优先的具身 AI 研究枢纽,覆盖 VLA、世界模型、多模态感知、文献综述、声明审计以及可复用的 Agent skillEvidence-first Embodied AI research hub for VLA, world models, multimodal sensing, literature reviews, claim audits, and reusable agent skills.

agentmultimodal
ZinYY/AdaFlash
Python · 2026-08-11 多模态 应用 实验 Stars 3 周增 +0

[COLMW'26] AdaFlash:通过 On-Policy 蒸馏扩散 Drafters 实现的自适应投机解码[COLMW'26] AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters

llm-infra
pbz134/LocalYT
HTML · 2026-08-12 多模态 实验 Stars 3 周增 +0

可自托管且功能丰富的视频库Self-hostable and feature-rich video library

multimodalllm-infra
Lauraineabsurd944/Light-WAM
Python · 2026-08-11 多模态 应用 实验 Stars 3 周增 +0

在 LIBERO 与 RoboTwin2.0 上,基于状态融合解码与冻结视频骨干网络训练面向机器人操作的轻量级世界动作模型。Train lightweight world action models for robot manipulation using state-fusion decoding and frozen video backbones on LIBERO and RoboTwin2.0.

ragmultimodalllm-infra
kenhayward/Diariz
C# · 2026-08-11 多模态 模型 实验 Stars 3 周增 +0

自托管、多用户转录平台:录制或上传音频。支持说话人标记、带时间戳的转录、跨录音识别说话人、摘要、提取行动项,并可使用自有 OpenAI 兼容 LLM 与转录内容对话。你的音频、你的服务器、你的模型。已在笔记本 RTX4070、台式机 RTX3090 与 RTX5090 上测试。Self-hosted, multi-user transcription platform: record or upload audio. Speaker-labeled, timestamped transcripts, Recognize speakers across recordings, Summarize, extract action items and chat over your transcripts with your own OpenAI-compatible LLM. Your Audio, your Server, your Model. Tested on Laptop RTX4070, Desktop RTX3090 and RTX5090

ragdatabasellm-infra
Sunrich-HT/authorship-aware-writing
Python · 2026-08-19 多模态 应用 实验 Stars 2 周增 +7

作者感知的科学写作分析:作者空间、目标论文风格、校准的生成证据,以及安全合规的修订Authorship-aware scientific writing analysis: author space, target-paper style, calibrated generation evidence, and integrity-safe revision.

multimodal
zhenhuanZ/DIA-Review
未知语言 · 2026-08-13 多模态 应用 研究原型 Stars 2 周增 +0

论文《Deep Learning in Dental Image Analysis: A Systematic Review of Datasets, Methodologies, and Emerging Challenges》的官方项目主页。The official project homepage for the paper "Deep Learning in Dental Image Analysis: A Systematic Review of Datasets, Methodologies, and Emerging Challenges"

multimodal
yanlingLabs/video-extract-mcp
TypeScript · 2026-08-20 多模态 应用 实验 Stars 2 周增 +0

MCP server,可从 URL 提取任意视频(包括微信视频号)。可选择性地分析并提取场景感知的去重关键帧,并对视频进行转录。完全本地运行MCP server that extracts any video from URL(inc/ WeChat channels). Optionally analyzes and extracts scene-aware deduplicated key frames and transcribes the video. ALL LOCAL

agentmultimodalllm-infra
jongan69/YouTubeResearchAI
JavaScript · 2026-08-14 多模态 应用 实验 Stars 2 周增 +0

将任意视频链接转化为博士级研究报告。自动化学术流水线:一键完成下载、转写、检索同行评审文献、核查论据并生成带引用的报告。免费且开源。Turn any video url into a PhD-grade research report. Automated academic pipeline: download, transcribe, search peer-reviewed literature, verify claims, generate cited reports — all with one command. Free & open source.

multimodalengineeringllm-infra
attavit14203638/transformer-tree-survey
Python · 2026-08-20 多模态 收藏榜 研究原型 Stars 2 周增 +0

论文《Transformer-Based Tree Extraction from Remote Sensing Imagery: A Systematic Review》的补充代码库。Supplementary repository for: Transformer-Based Tree Extraction from Remote Sensing Imagery: A Systematic Review

multimodal
ningkaikok/dotty-tutor
Python · 2026-08-20 多模态 应用 实验 Stars 1 周增 +0

AI 教材数字化与互动辅导平台:PDF/扫描教材经 MinerU OCR 结构化出题,文本 + 视觉双模型审校与确定性质量门禁;学生端七种题型互动、分层提示与错题多轮陪练闭环;FastAPI + PostgreSQL JSONB + React 19

ragdatabasellm-infra
msamribeiro/speech-generation-wiki
未知语言 · 2026-08-23 多模态 应用 实验 Stars 1 周增 +0

TTS、语音转换与口语对话 Agent 的活体系统综述Living systematic review of TTS, voice conversion, and spoken conversational agents

agent
mananp-2730/BridgeBuild-AI-PM-Tool
Python · 2026-08-11 多模态 工具 实验 Stars 1 周增 +0

AI 驱动的企业级 Agile OS,使用 Gemini 1.5 将原始客户音频与笔记即时转化为客户 pitch deck、PM epic、UI 规格与后端工程架构。An AI-powered Enterprise Agile OS that instantly translates raw client audio and notes into client pitch decks, PM epics, UI specs, and backend engineering architectures using Gemini 1.5.

rag
fuliao603/paper-reader
JavaScript · 2026-08-10 多模态 应用 实验 Stars 1 周增 +0

面向学术阅读的本地桌面 PDF 阅读器,支持 AI 翻译、OCR、高亮、笔记、翻译历史以及导入/导出。A local desktop PDF reader for academic reading, with AI translation, OCR, highlights, notes, translation history, and import/export support.

Fatma3598/EEG-Imagined-Speech-Preprocessing-Generalization
Jupyter Notebook · 2026-08-12 多模态 应用 实验 Stars 1 周增 +0

《Comparative Analysis of EEG Signal Preprocessing Impact on Generalization of a Transformer-Based EEG Imagined Speech Recognition》研究的可复现研究仓库。Reproducible research repository for 'Comparative Analysis of EEG Signal Preprocessing Impact on Generalization of a Transformer-Based EEG Imagined Speech Recognition' research work

aidadbeni/Computational-Pathology
Jupyter Notebook · 2026-08-19 多模态 收藏榜 研究原型 Stars 1 周增 +0

围绕病理 foundation models(尤其是 Prov-GigaPath)的文献综述与可复现性研究Literature review and reproducibility work on pathology foundation models, with a focus on Prov-GigaPath.

Yashwanth-23/Omnisense
Python · 2026-08-16 多模态 应用 实验 Stars 0 周增 +0

基于 Llama 3.2 的本地多模态 RAG AI 助手,可与 PDF、图片和视频对话,完全离线,零数据外泄。Local multimodal RAG AI assistant where you can chat with PDFs, images, and video using Llama 3.2, fully offline, zero data exposure.

ragmultimodalllm-infra
XenoVoyage/xeno-rover
未知语言 · 2026-08-14 多模态 应用 实验 Stars 0 周增 +0

开源 Raspberry Pi 5 rover 项目,用于建图、视觉、语音、本地 LLM 及液态神经网络实验。Open-source Raspberry Pi 5 rover project for mapping, vision, voice, local LLMs, and liquid-neural-network experiments.

multimodalllm-infra
wozengyi/embodied-daily
JavaScript · 2026-08-25 多模态 应用 实验 Stars 0 周增 +0

Embodied Daily:每日具身智能 AI 论文推荐(HF Daily Papers + arXiv)Embodied Daily: daily embodied-AI paper recommendations (HF Daily Papers + arXiv)

tobias-weiss-ai-xr/robotics-research
Python · 2026-08-23 多模态 数据集 实验 Stars 0 周增 +0

数据驱动、自动验证的机器人研究文献综述语料库(manipulation、locomotion、perception、planning、learning、HRI、multi-robot、simulation)。Robotics Research Corpus: Data-driven, auto-validated literature review for robotics research (manipulation, locomotion, perception, planning, learning, HRI, multi-robot, simulation)

shokrydev/master
Python · 2026-08-25 多模态 教程 实验 Stars 0 周增 +0

本仓库最初是我毕设项目的 PyTorch Lightning 模板,后来演进为包含完整项目实现本身。研究问题是地理定位如何影响遥感视觉-语言模型的训练与推理。注:仓库名称与描述可能变化。This repo originally was a PyTorch Lightning template for my thesis project. It has since evolved to contain the full project implementation itself. The research question is how geolocation influences the training and inference of vision-language models for remote sensing. Note: The repository name and description may change

multimodalllm-infra
seasalim/happy-photon
C# · 2026-08-13 多模态 应用 生产可用 Stars 0 周增 +0

更简洁的照片编辑体验;开源 RAW 与 JPEG 流水线。Photo editing, simplified. An open-source RAW and JPEG workflow.

redocto/image-text-structurizer
Python · 2026-08-22 多模态 工具 实验 Stars 0 周增 +0
rag
outrotim/ideal-to-real-publication-figure
未知语言 · 2026-08-24 多模态 工具 实验 Stars 0 周增 +0

Codex skill:将已批准的科研图表蓝图转化为可溯源、使用真实数据的发表级图表。Codex skill for turning approved scientific figure blueprints into traceable real-data publication figures

outrotim/evidence-aligned-scientific-revision
Python · 2026-08-17 多模态 应用 实验 Stars 0 周增 +0

面向中英双语生物医学研究的术语循证、主张强度与科学写作审计 skill。Evidence-aligned terminology, claim-strength, and scientific-writing audit skill for Chinese and English biomedical research.

agent
no-problem-dev/swift-latex-view
Swift · 2026-08-11 多模态 实验 Stars 0 周增 +0

SwiftUI 原生 LaTeX 数学渲染,对 LLM 输出具有鲁棒性——支持分隔符归一化、流式渲染与自动主题适配SwiftUI-native LaTeX math rendering robust to LLM output — delimiter normalization, streaming support, and automatic theming

llm-infraengineering