Repositories · organized/repo_cards

仓库/Skill 库

167 个 · 多模态

排序 Stars 周增
matguo/Phyadv_pipeline
未知语言 · 2026-08-12 多模态 应用 实验 Stars 0 周增 +0

本仓库包含论文补充材料,旨在提升系统综述的透明度、可复现性与完整性,涵盖筛选文档、检索策略以及计算机视觉中物理对抗攻击的详细分析。This repository contains the supplementary materials accompanying our paper. The materials are designed to support the transparency, reproducibility, and comprehensiveness of the systematic review, including screening documentation, search strategies, and detailed analysis of physical adversarial attacks in computer vision.

multimodalrisk
Macr7523/Video-Summarizer
Python · 2026-08-11 多模态 教程 实验 Stars 0 周增 +0

本地使用 AI 总结视频,从讲座、会议和教程中提取视觉亮点与文本摘要,基于 RTX 40 系列 GPU 的 CUDA 加速Summarize videos locally with AI, extracting visual highlights and text summaries from lectures, meetings, and tutorials using CUDA on RTX 40-series GPUs

multimodalllm-infra
lol-dungeonmaster/official-daily-banana
Jupyter Notebook · 2026-08-24 多模态 数据集 实验 Stars 0 周增 +0

每日提示词,激发 nano banana 生成灵感。Daily prompts to inspire nano banana generation.

multimodalllm-infra
littlecookie0722/wairc-2026
Python · 2026-08-23 多模态 应用 实验 Stars 0 周增 +0

基于 IQ 信号、采用 STFT 频谱图、视觉模型、k-fold 训练与集成推理的多节点 RF 无人机识别可复现研究工具包。A reproducible research toolkit for multi-node RF drone identification from IQ signals using STFT spectrograms, vision models, k-fold training, and ensemble inference.

multimodalllm-infra
Lacenedihia/Google-Search-Ranking-Discoverability-Capstone
HTML · 2026-08-20 多模态 应用 实验 Stars 0 周增 +0

研究问题与暂定方向Research Question and Provisional Lane

multimodal
kaulzeejai/AIA-Academic-Illustrator-
JavaScript · 2026-08-13 多模态 工具 实验 Stars 0 周增 +0

🎨 利用 GPT、Gemini 等模型,AI 驱动的学术图表一键生成与定制工具🎨 Generate academic diagrams effortlessly with this AI-driven tool, leveraging models like GPT and Gemini for seamless creation and customization.

ragllm-infra
karim1988781/agrirobotic_imaging
Jupyter Notebook · 2026-08-22 多模态 数据集 研究原型 Stars 0 周增 +0

面向农业机器人成像系统综述与 Meta 分析的数据与分析代码。Data and analysis code for a systematic review and meta-analysis of agricultural robotic imaging

karim1988781/agri-robotic-imaging-meta-analysis
Jupyter Notebook · 2026-08-21 多模态 数据集 实验 Stars 0 周增 +0

农业机器人成像系统综述与 meta-analysis 的数据与分析代码。Data and analysis code for a systematic review and meta-analysis of agricultural robotic imaging

JessiePBhalerao/miscanthus-yield-maps-viz-demo
未知语言 · 2026-08-14 多模态 教程 实验 Stars 0 周增 +0

基于 GitHub Page 的 2026 年文献综述 Miscanthus 产量制图演示。Demonstration of GitHub Page with Miscanthus Yield mapping for Literature Review 2026

jakemorgan-research/codex-sci-research-lifecycle
Python · 2026-08-22 多模态 应用 实验 Stars 0 周增 +0

可审计的 Codex skill 与 Python 项目包,用于证据可追溯的研究、系统综述、投稿与修订。Auditable Codex skill and Python project pack for evidence-traceable research, systematic reviews, submission, and revision.

multimodal
Homologic-bid91/voice_clone_lab
Python · 2026-08-11 多模态 应用 实验 Stars 0 周增 +0

使用 Qwen3-TTS 在本地 GPU 上克隆声音并从文本生成语音,提供端到端训练流水线。Clone voices and generate speech from text locally on your GPU using Qwen3-TTS with an end-to-end training pipeline.

engineeringllm-infra
HelgDemidov/refigure
Python · 2026-08-20 多模态 应用 实验 Stars 0 周增 +0

保留图表的转换器 —— DOCX/XLSX 转 Markdown,采用原生 OOXML 图表数据提取(无需光栅化/OCR/VLM),并提供零损耗的复合图表标记。Converters where figures survive — DOCX/XLSX to Markdown with native OOXML chart-data extraction (no rasterize/OCR/VLM) and zero-loss composite-figure markers

ragllm-infra
Harryphan72007/minh-phan-portfolio
TypeScript · 2026-08-19 多模态 应用 实验 Stars 0 周增 +0

Minh Phan 作品集:机器学习系统、计算机视觉、可复现研究与软件工程Minh Phan portfolio: ML systems, computer vision, reproducible research, and software engineering.

multimodal
hanjiarui1025-a11y/rebuttal-revision
未知语言 · 2026-08-16 多模态 应用 实验 Stars 0 周增 +0

Codex Skill:学术期刊 rebuttal 与修改工作流Codex skill for academic journal rebuttal and revision workflows.

multimodal
gong8/paper-reader
HTML · 2026-08-25 多模态 应用 实验 Stars 0 周增 +0

学术 PDF 阅读器:在浏览器中并排展示印刷版页面与重排阅读视图,由同一模型解析并联动。An academic PDF reader: the printed page and a reflowed reading view, side by side, linked by one model parsed in your browser.

galaxy99881/galaxy99881
未知语言 · 2026-08-16 多模态 应用 实验 Stars 0 周增 +0

Zhixin Li 的学术主页:超声医学、医学影像 AI、多模态学习与可复现研究。Academic homepage of Zhixin Li: ultrasound medicine, medical imaging AI, multimodal learning, and reproducible research.

multimodal
fhlyongko/graduate-academic-writing-ebook-kr
TypeScript · 2026-08-09 多模态 应用 实验 Stars 0 周增 +0

面向研究生学术英语写作、研究设计、伦理、修改与发表的韩文交互式现场指南Interactive Korean field guide for graduate academic English writing, research design, ethics, revision, and publication

multimodal
Erikalaylafajri15/MOSS-VL
未知语言 · 2026-08-11 多模态 模型 实验 Stars 0 周增 +0

使用 vision-language model 处理视觉与文本数据,执行多模态推理与图像理解任务。Process visual and textual data with this vision-language model for multimodal reasoning and image understanding tasks.

multimodalllm-infra
Davimoren9040/Youtube-Video-Transcribe-Summarizer-LLM-App
未知语言 · 2026-08-18 多模态 应用 实验 Stars 0 周增 +0

基于 Whisper、Gemini、Streamlit、yt-dlp 与 FFmpeg 构建的 AI 应用,可即时转录并总结任意 YouTube 视频。Transcribe and summarize any YouTube video instantly with AI-powered app using Whisper, Gemini, Streamlit, yt-dlp & FFmpeg.

multimodalllm-infra
BOHARRY/swipeta-research
Python · 2026-08-19 多模态 框架 研究原型 Stars 0 周增 +0

SwipeTA 背后关于图表识别、市场模式与噪声反馈的可复现研究Reproducible research on chart-reading, market patterns, and noisy feedback behind SwipeTA.

benisonodigie-dev/MMLA-systematic-Review
Jupyter Notebook · 2026-08-22 多模态 数据集 实验 Stars 0 周增 +0

MMLA 在发展中经济体与 LMIC 中的综合系统综述。A Comprehensive Systematic review of MMLA in developing economics and LMIC.

api-evangelist/kibin
未知语言 · 2026-08-25 多模态 应用 生产可用 Stars 0 周增 +0

Kibin 是一家运营 kibin.com 的消费学习与学术写作公司,自 2011 年起受到学生信赖。其最新产品是一款 AI 驱动的学习伙伴(iOS 应用),可将讲座、笔记、PDF 和 YouTube 视频转化为个性化学习指南、测验、闪卡、转录稿与讲解,帮助学生更快学习……Kibin is a consumer study and academic-writing company operating at kibin.com, trusted by students since 2011. Its newest product is an AI-powered study partner (iOS app) that turns lectures, notes, PDFs, and YouTube videos into personalized study guides, quizzes, flashcards, transcripts, and explanations to help students learn faster and build…

multimodal
ajgarciaj/NaViL
未知语言 · 2026-08-15 多模态 应用 实验 Stars 0 周增 +0

🌐 在数据受限条件下重新思考多模态大语言模型的设计与扩展,以 NaViL 通过 Native Training 提升效率与性能🌐 Rethink Multimodal Large Language Models design and scaling under data constraints with NaViL, enhancing efficiency and performance through Native Training.

agentragmultimodalllm-infra