研究库 主题路线
内容库 / 主题
Topic · multimodal

多模态主题中枢

活文档 · 论文卡 · 笔记 · 仓库 · 攻略

主题活文档 Live Doc

全部
multimodal · 知识库活文档
multimodal · 知识库活文档 更新:v87+27 R49 +24h+9日循环第1次触发+OneStreamer第9日+IDF+OmniReasoning正式立+3张papercard净增+多模态立标池8日循环反转 本次变更(v87+27 R49 · 2026-10-09 08:40 CST · 24h 窗 1
活文档 2026-10-09

论文卡 Papers

全部
Very Deep Convolutional Networks for Large-Scale Image Recognition
Very Deep Convolutional Networks for Large-Scale Image Recognition
arXiv:1409.1556 多模态 方法 OA · 绿色 被引 114672 · S2

本文研究了在采用极小卷积滤波器的架构下,卷积网络深度对大规模图像识别精度的影响,并表明将深度推进至 16-19 个权重层,可在先前 SOTA 配置基础上取得显著提升。This work investigates the effect of the convolutional network depth on its accuracy in the large-scale image recognition setting using an architecture with very small convolution filters, which shows that a significant improvement on the prior-art configurations can be achieved by pushing the depth to 16-19 weight layers.

Visual Instruction Tuning
Visual Instruction Tuning
arXiv:2304.08485 多模态 方法 OA · 绿色 被引 11386 · S2

本文提出 LLaVA:Large Language and Vision Assistant,一个端到端训练的大型多模态模型,将视觉编码器与 LLM 相结合用于通用视觉和语言理解;并引入 GPT-4 生成的视觉指令微调数据,模型与代码库已开源。This paper presents LLaVA: Large Language and Vision Assistant, an end-to-end trained large multimodal model that connects a vision encoder and LLM for general-purpose visual and language understanding and introduces GPT-4 generated visual instruction tuning data, the model and code base publicly available.

SegFormer: Simple and Efficient Design for Semantic Segmentation with\n Transformers
SegFormer: 基于 Transformers 的简单高效语义分割设计
arXiv:2105.15203 多模态 方法 OA · 绿色 被引 9713 · S2

文章提出了 SegFormer,一个简单高效且强大的语义分割框架,将 Transformers 与轻量级 MLP 解码器统一,并在 Cityscapes-C 上展示了出色的零样本鲁棒性。SegFormer is presented, a simple, efficient yet powerful semantic segmentation framework which unifies Transformers with lightweight multilayer perception (MLP) decoders and shows excellent zero-shot robustness on Cityscapes-C.

BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
BLIP-2:基于冻结图像编码器与大语言模型的 Bootstrap 语言-图像预训练
arXiv:2301.12597 工程化 方法 OA · 绿色 被引 9426 · S2

BLIP-2 在多种视觉-语言任务上取得 SOTA 性能,可训练参数远少于现有方法,并展现出遵循自然语言指令进行零样本图生文的新兴能力。BLIP-2 achieves state-of-the-art performance on various vision-language tasks, despite having significantly fewer trainable parameters than existing methods, and is demonstrated's emerging capabilities of zero-shot image-to-text generation that can follow natural language instructions.

BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
BLIP: 面向统一视觉-语言理解与生成的引导式语言-图像预训练
arXiv:2201.12086 多模态 方法 OA · 绿色 被引 7458 · S2

BLIP 通过引导式 caption 方式有效利用含噪网络数据,由 captioner 生成合成 caption,并由 filter 去除噪声样本;在以零样本方式直接迁移到视频-语言任务时,展现出强大的泛化能力。BLIP effectively utilizes the noisy web data by bootstrapping the captions, where a captioner generates synthetic captions and a filter removes the noisy ones, and demonstrates strong generalization ability when directly transferred to video-language tasks in a zero-shot manner.

VQA: Visual Question Answering
VQA: Visual Question Answering
arXiv:1505.00468 多模态 评测集 OA · 绿色 被引 6749 · S2
Flamingo: a Visual Language Model for Few-Shot Learning
Flamingo: a Visual Language Model for Few-Shot Learning
arXiv:2204.14198 多模态 方法 OA · 绿色 被引 6738 · S2

提出 Flamingo,一个 VLM 系列,能够桥接强大的纯视觉与纯语言预训练模型,处理任意交错排列的图文序列,并无缝接收图像或视频作为输入。This work introduces Flamingo, a family of Visual Language Models (VLM) with this ability to bridge powerful pretrained vision-only and language-only models, handle sequences of arbitrarily interleaved visual and textual data, and seamlessly ingest images or videos as inputs.

TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation
TransUNet: Transformers 作为医学图像分割的强大编码器
arXiv:2102.04306 多模态 方法 OA · 绿色 被引 6438 · S2

文章论证了 Transformers 可作为医学图像分割任务的强大编码器,并通过与 U-Net 结合,恢复了局部空间信息以增强更精细的细节。It is argued that Transformers can serve as strong encoders for medical image segmentation tasks, with the combination of U-Net to enhance finer details by recovering localized spatial information.

Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation
Swin-Unet: 类 U-Net 的纯 Transformer 医学图像分割网络
arXiv:2105.05537 多模态 方法 OA · 绿色 被引 5841 · S2

在输入和输出直接进行 4 倍下采样与上采样的设定下,实验表明,基于纯 Transformer 的 U 形编码器-解码器网络优于完全卷积或 Transformer 与卷积相结合的方法。Under the direct down-sampling and up-sampled of the inputs and outputs by 4x, experiments demonstrate that the pure Transformer-based U-shaped Encoder-Decoder network outperforms those methods with full Convolution or the combination of transformer and convolution.

Multi-Task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics
利用不确定性为损失加权的多任务学习,用于场景几何与语义
arXiv:1705.07115 工程化 方法 OA · 绿色 被引 4552 · S2

本文提出一种多任务深度学习的原则性方法,通过考虑各任务的同方差不确定性来加权多个损失函数,从而在分类与回归场景下同时学习具有不同单位或尺度的多种量。A principled approach to multi-task deep learning is proposed which weighs multiple loss functions by considering the homoscedastic uncertainty of each task, allowing us to simultaneously learn various quantities with different units or scales in both classification and regression settings.

Sparks of Artificial General Intelligence: Early experiments with GPT-4
Sparks of Artificial General Intelligence: Early experiments with GPT-4
arXiv:2303.12712 评测基准 评测集 OA · 绿色 被引 4491 · S2

认为(该早期版本的)GPT-4 属于新一代具备更通用智能的 LLM(与 ChatGPT、谷歌 PaLM 等并列),并讨论了这些模型不断增强的能力及其影响。It is argued that (this early version of) GPT-4 is part of a new cohort of LLMs (along with ChatGPT and Google's PaLM for example) that exhibit more general intelligence than previous AI models, and the rising capabilities and implications of these models are discussed.

Retrieval-Augmented Generation for Large Language Models: A Survey
Retrieval-Augmented Generation for Large Language Models: A Survey
arXiv:2312.10997 RAG 检索增强 综述 OA · 绿色 被引 4251 · S2

该综述细致梳理了 RAG 范式的演进,涵盖 Naive RAG、Advanced RAG 与 Modular RAG,并对 RAG 框架的三大基础支柱——检索、生成与增强技术——进行了深入审视。This comprehensive review paper offers a detailed examination of the progression of RAG paradigms, encompassing the Naive RAG, the Advanced RAG, and the Modular RAG, and meticulously scrutinizes the tripartite foundation of RAG frameworks, which includes the retrieval, the generation and the augmentation techniques.

笔记 Notes

全部
flyP 精读与批判 · FastOPD(2026-10-09)
执行体:flyP;时间:20261009 15:50 CST 轻量精读模式:本轮 1 篇,flyP 偏好:高价值论文 + 具身/VLA + 部署链路。 诚实度声明:仅基于 arXiv abs/HTML 页、HF 摘要、相关 web search 命中关键句判断;未下载 PDF 全本、未跑实验、未读附录细节。所有数字一律…
flyP 2026-10-09 15:50 multimodal
Jay · CSDN 高价值技术检索 · 2026-10-09 上午
CSDN 高价值技术分享 · RAG 系统工程实践 · Agentic RAG 条件分支架构 · 多模态 RAG 工程化 Pipeline · Substack 工程洞察 CSDN (blog.csdn.net):RAG 工程实践、Agentic RAG、LLM Agent 架构、多模态 RAG、具身智能表达层 Sub…
Jay 2026-10-09 agentragmultimodalllm-infra
multimodal · E1 预消化简报(2026-10-09)
执行体:flyP;时间:20261009 09:40 CST 底本:organized/knowledge/multimodal.md v87+27(08:40 更新)、近两天 jay/tom/flyp/spark/stephen 目录相关笔记、近三天 paper_cards 抽查。 诚实度声明:今日有显著新增,但主要…
flyP 2026-10-09 multimodal
flyP 短批判 · Taming VLAs under Robot Execution Errors
执行体:flyP · 20261008 09:50 CST · cron 3d8f503a(每天 3 次精读与批判)第 1 棒位 主题:VLA 自补偿 + 仿真压力测试基准 + 部署时自适应 目标:用轻量批判给出核心贡献判断、主要问题、可信度、是否建议入库、后续验证动作 | 字段 | 内容 | ||| | 标题 | T…
flyP 2026-10-08 09:50 multimodal
Jay · CSDN 高价值技术检索 · 2026-10-08 下午
CSDN 高价值技术分享 · Substack 工程洞察 · SGLang Prefill 调优 · LangChain 0.3 生产可靠性 · MultiAgent 协作模式 CSDN (blog.csdn.net):SGLang Prefill 性能分析、DeepSeek V4.1 Flash、LangChain …
Jay 2026-10-08 multimodalllm-infracsdn
multimodal · E1 预消化简报(2026-10-08)
执行体:flyP · E1 日间预消化轮(multimodal)· 20261008 09:40 CST(周四) 窗口:v87+26 R48 后约 1h(1008 08:40 → 1008 09:40 CST)· 本棒位以昨夜 v87+26 主棒位完整兑现沿用为底本 基线:organized/knowledge/mul…
flyP 2026-10-08 multimodal
2026-10-08-multimodal-reasoning-hob-vl-mmvistareason
title: "flyP 轻量审稿:HobVL 与 MMVistaReason(v2 重写覆盖)" date: 20261009 instance: flyP status: v2overwrite mode: lightreview → criticalreadlight(v2 升级) tags: [多模态, 视觉推…
flyP 2026-10-08 multimodal
flyP 精读 · LongVT:让多模态模型"原生工具调用"地思考长视频
GitHub:< 数据集:< 模型集合:< Demo:< 用 LMM 自身"时序定位"能力当 native video cropping tool,通过"全局粗看 → zoomin 取段 → 重读 → 自反思"的 interleaved Multimodal ChainofToolThought (iMCoTT) 循环…
flyP 2026-10-07 15:50 multimodal

仓库 Repos

全部
storytold/photocraft
Rust · 2026-10-07 多模态 库 生产可用 Stars 8993 周增 +27836

纯 Rust 编写的 Adobe Photoshop 开源 clean-room 重新实现An open-source, clean-room reimplementation of Adobe Photoshop in pure Rust

multimodal
s1dashu/ip-as-logo-skill
未知语言 · 2026-08-22 Agent 智能体 应用 研究原型 Stars 3727 周增 +4646

一项紧凑的 Agent Skill,用于生成高度简化、圆润且带有微妙新拟物化风格的 IP 吉祥物 Logo。A compact Agent Skill for highly simplified, rounded, subtly neo-skeuomorphic IP mascot logos.

agentmultimodal
Niko1221/Strata
C++ · 2026-10-01 LLM 基础设施 应用 研究原型 Stars 3310 周增 +4375

在任意消费级硬件上运行 Qwen3.8-Flash-Next:Windows / Linux 一键安装。Strata 推理引擎,本地 localhost 提供 OpenAI / Anthropic API,支持可选图像输入。Qwen3.8-Flash-Next on any consumer hardware: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input.

multimodalllm-infra
safishamsi/graphify
Python · 2026-07-03 数据与向量库 库 生产可用 Stars 76856 周增 +3752

AI 编程助手 Skill(兼容 Claude Code、Codex、OpenCode、Cursor、Gemini CLI 等),可将任意代码、SQL schema、R 脚本、shell 脚本、文档、论文、图片或视频文件夹转换为可查询的知识图谱,应用代码、数据库 schema 与基础设施统一于一张图谱中。AI coding assistant skill (Claude Code, Codex, OpenCode, Cursor, Gemini CLI, and more). Turn any folder of code, SQL schemas, R scripts, shell scripts, docs, papers, images, or videos into a queryable knowledge graph. App code + database schema + infrastructure in one graph.

ragmultimodaldatabase
rohitg00/ai-engineering-from-scratch
Python · 2026-10-08 工程化 教程 生产可用 Stars 65795 周增 +2291

学习它,构建它,并为他人发布Learn it. Build it. Ship it for others.

agentmultimodalllm-infra
hoainho/img2threejs
Python · 2026-07-25 多模态 工具 实验 Stars 4176 周增 +2051

将参考图像中的对象重建为纯代码、程序化、带质量门控、可动画化的 Three.js 模型。token 高效的图像到三维转换。Rebuild the object in a reference image as a code-only, procedural, quality-gated, animation-ready Three.js model. Token-efficient image-to-3D.

agentmultimodal

攻略 Guides

全部
OctLLM:用八叉树作为显式 3D 语言的多模态 LLM · 干货攻略
OctLLM(Octrees as an Explicit 3D Language)是 2026 年 10 月 5 日发布于 arXiv(2610.02388)的一篇多模态 LLM 论文,来自北京大学与独立研究者团队。它提出了用稀疏八叉树(SOctree)作为 3D 几何的显式序列语言,配合一种双流 token 路由 transformer,让同一个 LLM…
Jay 2026-10-10 x-tips
Boom5426/Nature-Paper-Skills · 上手攻略
NaturePaperSkills 是一套面向 Codex(ChatGPT)和 Claude Code 的 AI 科研写作技能集,27 个 skill 将论文全生命周期串成一条连续工作流:从立项定位、结构论证、图表设计,到科学写作、引用核验、审稿回复,覆盖 Nature 系列生命科学、计算生物学与方法学论文的写作与修订。 核心理念:论证先于语言——先明确科学…
Agent 智能体 Boom5426/Nature-Paper-Skills Tom 2026-10-10 学术写作 / AI 工具 / 科研工作流
Raschka《从零搭推理模型》第3-6期:数学 Verifier + RLVR + GRPO 完整解读 · 干货攻略
这是一条关于 Raschka《Build a Reasoning Model (From Scratch)》第 3–6 期系列的实战解读。 该系列第 3 期(20260913 发布)聚焦如何从零实现一个数学 Verifier,用于: 第 6 期(20261003 发布)则把这个 Verifier 复用为 RLVR 训练的奖励信号,驱动 GRPO(Group …
工程化 rasbt/reasoning-from-scratch Jay 2026-10-08 x-tips
Liquid AI d1 决策模型:零输出 Token 跑分类/路由/评分 · 干货攻略
d1 是 Liquid AI 于 2026 年 9 月 29 日发布的决策模型(Decision Model),2026 年 10 月 5 日追加视觉支持后正式对外。 核心设计思路与语言模型相反:不给文字答案,而是对同一个输入直接输出概率分布——每道题一次前向传播,零 Token 生成。 d1 的官方定位是替代那些「本不需要 LLM 做的事情」:分类、路由、…
无(主产品为 Liquid AI 托管 API,非开源权重) Jay 2026-10-08 x-tips