研究库 论文知识库
Papers · organized/paper_cards

论文

165 张论文卡片 · RAG 检索增强 · 方法

开放获取 全部 绿色 · 1640
HistoRAG: Embedding Historical Methodology in Retrieval-Augmented Generation Through Critical Technical Practice
HistoRAG:通过批判性技术实践将历史学方法论嵌入 RAG
arXiv:2606.18103 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 HistoRAG,一个将史学原则转化为具体架构干预的框架,为特定领域认识论承诺如何转化为 RAG 设计决策提供模型,并可迁移至其他处理大规模语料的诠释性学科。HistoRAG is introduced, a framework that translates historiographical principles into concrete architectural interventions and offers a model for how domain-specific epistemological commitments can be translated into RAG design decisions, and may transfer to other interpretive disciplines working with large corpora.

A Unified Framework for Context-Aware and Relation-Aware Graph Retrieval-Augmented Generation
面向上下文感知与关系感知图 RAG 的统一框架
arXiv:2606.18075 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 HyGRAG,一种分层图 RAG 框架,通过解决三个核心挑战超越源文档限制:构建真正融合上下文与关系信息的摘要、利用这些综合表示在检索阶段访问涌现知识、以及为动态语料高效更新分层结构。HyGRAG is proposed, a hierarchical graph RAG framework that transcends source documents by addressing three core challenges: constructing summaries that genuinely integrate contextual and relational information, leveraging these synthesized representations to access emergent knowledge during retrieval, and efficiently updating hierarchical structures for dynamic corpora.

When Rules Learn: A Self-Evolving Agent for Legal Case Retrieval
当规则学会学习:面向法律案例检索的自进化 Agent
arXiv:2606.17220 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一个面向规则驱动查询改写的自进化框架,无需任何参数训练即可增强 BM25,并揭示 LLM 利用先前实验结果的能力以及其对规则消除的内在知识,在通过自进化精炼规则集方面起到关键作用。This work proposes a self-evolving framework for rule-driven query rewriting that enhances BM25 without any parameter training, and reveals that LLM's capabilities to leverage previous experimental results and its intrinsic knowledge of rule elimination play critical roles in refining the rule set via self-evolution.

RL-Index: Reinforcement Learning for Retrieval Index Reasoning
RL-Index:面向检索索引推理的强化学习
arXiv:2606.16316 RAG 检索增强 方法 OA · 绿色 被引 2 · S2

RL-Index 被提出,是一个将检索索引推理建模为强化学习问题的索引框架,能持续提升检索与下游问答性能,同时显著降低在线推理延迟。RL-Index is proposed, an indexing framework that formulates retrieval index reasoning as a reinforcement learning problem that consistently improves both retrieval and downstream question-answering performance, while significantly reducing online inference latency.

MAGE-RAG: Multigranular Adaptive Graph Evidence for Agentic Multimodal RAG in Long-Document QA
MAGE-RAG:面向长文档问答中 Agentic 多模态 RAG 的多粒度自适应图证据
arXiv:2606.15906 RAG 检索增强 方法 OA · 绿色 被引 2 · S2

本文提出 MAGE-RAG,一个面向长文档多模态问答的多粒度自适应图证据框架,并建立了涵盖 Direct MLLM、Text RAG、Page-level Visual RAG 与 Graph/Agentic RAG 的统一比较与分析协议。This paper proposes MAGE-RAG, a multigranular adaptive graph evidence framework for long-document multimodal QA, and establishes a unified comparison and analysis protocol covering Direct MLLM, Text RAG, Page-level Visual RAG, and Graph/Agentic RAG.

Ricci-Filtration: Boosting Retrieval-Augmented Generation Reranker to Query-Answer Tasks by Discrete Ricci Flow
Ricci-Filtration:通过离散 Ricci Flow 将 RAG 重排序器提升至 Query-Answer 任务
arXiv:2606.15482 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文理论上证明,归一化离散 Ricci Flow 可通过识别边权中的不同渐近行为来检测社区结构,并支持移除相对于 query 节点具有大权重与负 Ricci 曲率的"噪声"文档片段。It is theoretically prove that normalized discrete Ricci flow can detect community structures by identifying distinct asymptotic behaviors in edge weights, and supports the removal of ``noisy''document chunks characterized by large weights and negative Ricci curvature relative to the query node.

CQC-RAG: Robust Retrieval-Augmented Generation via Cross-Query Consistency
CQC-RAG:通过跨查询一致性实现鲁棒的检索增强生成
arXiv:2606.13438 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

CQC-RAG 框架被提出,它协同设计查询级多样性注入与跨查询一致性评估,无需外部监督即可实现自我评估,验证了跨查询一致性在过滤噪声引发幻觉方面的有效性。CQC-RAG, a framework that co-designs query-level diversity injection with cross-query consistency evaluation and enables self-evaluation without external supervision, is introduced, validating the effectiveness of cross-query consistency for filtering noise-induced hallucinations.

Large Behavior Model: A Promptable Digital Twin of the Retail Customer
Large Behavior Model:零售客户的可提示数字孪生
arXiv:2607.06993 RAG 检索增强 方法 OA · 绿色 被引 1 · S2

结果表明,交易历史中编码的行为知识可以被语言模型有效学习,为客户数字孪生和行为模拟提供了可扩展的基础。The results demonstrate that behavioral knowledge encoded in transaction histories can be effectively learned by language models, providing a scalable foundation for customer digital twins and behavior simulation.

Conversational Retrieval and On-the-Fly Knowledge Modeling of Historical Penitentiary Repression Records
历史监狱压迫记录的对话式检索与即时知识建模
arXiv:2607.08459 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一种面向历史数字图书馆管理的文档分析系统,支持即时知识建模,并促进生成更丰富、更全面的信息。This article presents a document analysis system designed for the management of historical digital libraries that supports on-the-fly knowledge modeling and facilitates the generation of richer and more comprehensive information.

PolyUQuest: Verifiable Structure-Aware Web RAG over Heterogeneous Graphs
PolyUQuest:异构图上的可验证结构感知 Web RAG
arXiv:2607.08269 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 PolyUQuest,一个可验证、感知结构的 Web RAG 框架,基于异构图构建,统一了页面间超链接拓扑、页面内 DOM 层级以及跨页面实体关系知识。PolyUQuest is presented, a verifiable, structure-aware web RAG framework built on a heterogeneous graph that unifies hyperlink topology between pages, DOM hierarchy within pages, and entity-relation knowledge across pages.

Deceptive Grounding: Entity Attribution Failure in Clinical Retrieval-Augmented Generation
欺骗性 grounding:临床 RAG 中的实体归因失败
arXiv:2607.09349 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

一项受控消融实验揭示了机制:从检索到的文档中移除特定实体的临床证据,可彻底消除实体归因失败,使所有失败转移到虚构生成。A controlled ablation identifies the mechanism: removing entity-specific clinical evidence from retrieved documents eliminates entity-attribution failure entirely, shifting all failures to confabulation.

Augmenting Fundamental Analysis with Large Language Models: A RAG-Based System for Generating Investor Briefs
利用 LLM 增强基本面分析:基于 RAG 的投资者简报生成系统
arXiv:2607.09121 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

论文探讨了 LLM 为公司基本面分析各方面带来的机会,分析依据包括公司报告、描述宏观经济状况(如 GDP 和通胀变化)的数据与文件,以及提交至美国证券交易委员会(SEC)的文件。The opportunities brought by Large Language Models (LLMs) to various aspects of fundamental analysis of companies based on their reports as well as data and documents describing macroeconomic situation like GDP and inflation changes as well as documents filled to the U.S. Securities and Exchange Commission (SEC) are examined.

EvoGraph-R1: Self-Evolving Multimodal Knowledge Hypergraphs for Agentic Retrieval
EvoGraph-R1:面向 Agentic 检索的自演化多模态知识超图
arXiv:2607.12764 RAG 检索增强 方法 OA · 绿色 被引 2 · S2

提出 EvoGraph-R1,一个自演化 GraphRAG 框架,将知识图谱重新概念化为由 Agent 交互塑造的动态环境,将自演化知识图谱确立为跨模态的基础范式。EvoGraph-R1 is introduced, a self-evolving GraphRAG framework that reconceptualizes knowledge graphs as dynamic environments shaped through agent interactions, establishing self-evolving knowledge graphs as a fundamental paradigm across modalities.

Earthquaker-AI: A Retrieval-Augmented Generation Framework with Rubric-Based Assessment for Primary School Earthquake Education
Earthquaker-AI:面向小学地震教育的、采用评分量表评估的 RAG 框架
arXiv:2607.14046 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Earthquaker-AI,一个混合式教育框架,在已有教育机器人项目基础上集成基于 RAG 的对话式 AI 助手,旨在提升小学生的地震应急准备与主动行动意识。该系统将曾获奖的 STEM 项目 Earthquaker 从 Lego WeDo2 的机械模拟拓展至认知与元认知层面:机器人组件利用 Lego WeDo2 自动化模拟地震响应,使学生能够与传感器和执行器进行交互。This paper presents Earthquaker-AI, a hybrid educational framework building upon a previously implemented educational robotics project by integrating a conversational AI assistant based on Retrieval-Augmented Generation. It aims to enhance earthquake preparedness and conscious action among primary-school students. The system extends the award-winning STEM project Earthquaker moving from mechanical simulation with Lego WeDo2 to cognitive and metacognitive processing. The robotics component uses Lego WeDo2 automation to simulate seismic response, letting students interact with sensors and actuat

GRASP: GRanularity-Aware Search Policy for Agentic RAG
GRASP:面向 Agentic RAG 的粒度感知搜索策略
arXiv:2607.10463 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 GRASP,一个用于训练智能体在多步推理过程中自适应协调互补检索工具的强化学习(RL)框架,并指出学会协调检索信号与上下文粒度对智能体的正确推理至关重要。GRASP is introduced, a reinforcement learning (RL) framework for training agents to adaptively coordinate complementary retrieval tools during multi-step reasoning, and it is suggested that learning to coordinate retrieval signals and context granularity is critical for agent's correct reasoning.

Chat2Scenic: An Iterative RAG-Based Framework for Scenario Generation in Autonomous Driving
Chat2Scenic:面向自动驾驶场景生成的迭代式 RAG 框架
arXiv:2607.14387 RAG 检索增强 方法 OA · 绿色 被引 1 · S2

提出 Chat2Scenic,是首个以领域特定语言 (DSL) 生成场景脚本的迭代式检索增强框架,并构建了一个涵盖 NHTSA、联合国车辆法规及其他来源共 123 个场景的开源场景生成基准。Chat2Scenic is presented, the first iterative retrieval-augmented framework to generate scenario scripts in Domain Specific Language (DSL) and proposes an open benchmark for scenario generation comprising 123 scenarios from various regulations, including NHTSA and United Nations Vehicle Regulations, as well as other sources.

RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM
RAGU:基于紧凑领域适配LLM的多步GraphRAG引擎
arXiv:2607.11683 RAG 检索增强 方法 OA · 绿色 被引 1 · S2

RAGU是一个开源模块化GraphRAG引擎,通过将抽取与整合分离来解决抽取-整合问题:实体和关系经过两阶段类型化抽取、基于DBSCAN的去重、LLM摘要和Leiden社区检测。RAGU, an open-source modular GraphRAG engine, addresses extraction from consolidation by separating extraction from consolidation: entities and relations pass through two-stage typed extraction, DBSCAN-backed deduplication, LLM summarization, and Leiden community detection.

NOWJ@COLIEE 2026: Adaptive Pipelines for Legal Retrieval and Reasoning
NOWJ@COLIEE 2026:面向法律检索与推理的自适应流水线
arXiv:2607.16603 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文介绍了 NOWJ 团队参加 COLIEE 2026 全部五项任务的方法与结果,采用基于稠密检索、注意力重排序和小样本提示 LLM 推理的检索增强生成框架。This paper presents the methodologies and results of the NOWJ team's participation across all five tasks of the COLIEE 2026 competition and adopts a retrieval-augmented generation framework with dense retrieval, attention-based reranking, and few-shot-prompted LLM reasoning.

Vector Search As Nearest Neighbor Matching: RAG-based Policy Learning in Causal Inference
向量搜索作为最近邻匹配:基于RAG的因果推断策略学习
arXiv:2607.18225 RAG 检索增强 方法 OA · 绿色 被引 1 · S2

该工作将两步方法的遗憾分解为候选生成遗憾和候选内选择遗憾,并利用最近邻估计器和Transformer的预测误差保证对后者进行了界。This work decomposes the regret of the two-step method into candidate-generation regret and within-candidate choice regret, and bound the latter using prediction-error guarantees for nearest-neighbor estimators and transformers.

AutoIndex: Learning Representation Programs for Retrieval
AutoIndex:为检索学习表征程序
arXiv:2607.18603 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

研究结果表明,文档表示不应被视为检索开始前一次性的固定预处理选择,而应作为一个明确的优化目标。The results suggest that document representation should not be treated as a fixed preprocessing choice made before retrieval begins, but as an explicit optimization target.

IteraSim RAG: A Multi-Stage Retrieval-Augmented Agentic Back-End for OpenFOAM-Based Computational Fluid Dynamics
IteraSim RAG:基于 OpenFOAM 计算流体力学的多阶段检索增强 Agentic 后端
arXiv:2607.20346 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 IteraSim RAG,一个面向自动化 OpenFOAM 算例生成的 RAG 软件后端,围绕三大局限构建:求解器选择、湍流闭合、边界条件与有限体积默认值。IteraSim RAG is presented, a retrieval-augmented software back-end for automated OpenFOAM case generation built around three limitations: solver selection, turbulence closures, boundary conditions and finite-volume defaults.

LAMAR: An Open Language-Aware Multilingual Alignment Reranker
LAMAR:一种开放的语言感知多语言对齐 Reranker
arXiv:2607.22042 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作发布了 LAMAR,一种具备语言感知能力的多语种 cross encoder,在训练中兼顾语义相关性与语言连贯性,在通用多语种 reranking 基准上整体以及各语言单独评估中均达到最佳性能。This work releases LAMAR, a language aware multilingual cross encoder trained to account for both semantic relevance and language coherence, which achieves the best performance overall and across all languages examined individually on general multilingual reranking benchmarks.

DeCoRAG: Cognitive Decoupling and Semantic-Aware Cropping for Complex Document Understanding
DeCoRAG:面向复杂文档理解的认知解耦与语义感知裁剪
arXiv:2607.24554 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

DeCoRAG 是一个多模态 Graph RAG pipeline,将知识处理从耦合的视觉-语义推理转向认知层面的 Decoupling,进而把推理空间从稠密、带噪的背景推向纯净、意图驱动的语义簇。DeCoRAG is a multimodal Graph RAG pipeline that shifts knowledge processing from coupled visual-semantic reasoning to cognitive Decoupling, and subsequently drives the reasoning space from dense, noisy backgrounds to purified, intent-driven semantic clusters.

Leveraging External Knowledge for Historical Document Restoration via Retrieval-Augmented Large Language Models
通过检索增强型大语言模型利用外部知识进行历史文档修复
arXiv:2607.21936 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出一种面向历史文档修复的新框架,利用搭载 RAG 的大语言模型,有效缓解了推断上下文相关专有名词的难题。A novel framework for historical document restoration that leverages large language models with retrieval-augmented generation (RAG) and effectively mitigates the challenge of inferring context-dependent proper nouns is introduced.

Reasoning Denoiser: Denoising Reasoning Traces for Hallucination Detection in Large Reasoning Models
Reasoning Denoiser:用于大推理模型幻觉检测的推理轨迹去噪
arXiv:2607.22098 RAG 检索增强 方法 OA · 绿色 被引 1 · S2

RedE 利用 final-answer attention 作为自动监督信号来塑造 step-level 表征空间,使其中的噪声步骤可被可靠识别与过滤,并在检测性能上超越有竞争力的基线。RedE leverages final-answer attention as an automatic supervision signal to shape the step-level representation space, yielding refined embeddings in which noisy steps can be reliably identified and filtered and improves detection performance over competitive baselines.

Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansion
将 CVE 映射到 MITRE ATT&CK 技术:精选金标准分类器与 LLM 辅助标签扩展的局限
arXiv:2607.25572 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文在由专家级 MITRE Center for Threat-Informed Defense 标注构成的、包含 1,207 条 CVE 的精选 gold 数据集上训练多标签分类器,结果表明该分类器受限于标签质量而非数据规模。A multi-label classifier is trained on a curated gold dataset of 1,207 CVEs from expert MITRE Center for Threat-Informed Defense mappings, indicating that the classifier is limited by label quality rather than dataset size.

Detecting Knowledge Inconsistencies Across Text, Tables, and Knowledge Graphs
跨文本、表格与知识图谱的知识不一致检测
arXiv:2607.25959 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Kontrast,一种利用 Text-to-SPARQL 与 LLM 推理将基于表格的答案与 KG 证据进行对比并对所产生的不一致性进行分类的自动框架,并表明文本、表格与 KG 可通过系统性对比相互补充与纠错。Kontrast is presented, an automatic framework that uses Text-to-SPARQL and LLM reasoning to compare table-based answers with KG evidence and categorize the resulting inconsistencies, and shows that text, tables, and KGs can complement and correct one another through systematic comparison.

Beyond Self-Knowledge: Propagating Uncertainty Across Reasoning and Retrieval in LLMs
超越自我认知:在 LLM 推理与检索间传播不确定性
arXiv:2607.25600 RAG 检索增强 方法 OA · 绿色 被引 1 · S2

BeyondUncertainty 首先引出结构化的临时答案与置信度估计,然后应用在 held-out 验证数据上选定并在测试评估前冻结的模型特定阈值,揭示了更具选择性的证据获取与端到端 token 效率之间的权衡。BeyondUncertainty first elicits a structured provisional answer and confidence estimate, then applies a model-specific threshold selected on held-out validation data and frozen before test evaluation, revealing a trade-off between more selective evidence acquisition and end-to-end token efficiency.

Temporal-Distance JEPA: Plan-Aware Representation Learning for Latent World Model Predictive Control
时序距离 JEPA:面向潜在世界模型预测控制的规划感知表示学习
arXiv:2607.25337 RAG 检索增强 方法 OA · 绿色 被引 8 · S2

Temporal-Distance-JEPA 通过发现离线日志中的时间进展结构,并将代价形式与规划时部署协同设计,缩小了 JEPA world-model 规划器的训练-规划差距。Temporal-Distance-JEPA narrows the train--plan gap for JEPA world-model planners by discovering temporal progress structure in offline logs and co-designing cost form with plan-time deployment.

DualG-MRAG: Decoupling Macro-Reasoning and Micro-Matching for Multimodal Retrieval-Augmented Generation
DualG-MRAG:面向多模态 RAG 的宏观推理与微观匹配解耦
arXiv:2607.28580 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 DualG-MRAG,一种面向多模态 RAG 的双层框架,解耦 Macro-reasoning 与 Micro-matching Graph 两类图结构,通过分离全局结构推理与细粒度证据匹配来抑制检索噪声,并引入动态规划解码机制,从 GNN 前向过程中直接提取显式推理路径。DualG-MRAG is proposed, a Dual-tier framework that introduces a decoupled architecture comprising Macro-reasoning and Micro-matching Graphs for Multimodal RAG to suppress retrieval noise by isolating global structural reasoning from fine-grained evidence matching, and introduces a dynamic programming decoding mechanism that extracts explicit reasoning paths directly from the GNN's forward pass.

GLM-RAG: Graph Language Models for Graph-Based Retrieval-Augmented Generation
GLM-RAG:面向图基 RAG 的图语言模型
arXiv:2607.28397 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

引入一个基于 GLM 的 retriever,并在单跳与多跳 RAG 场景下对比分析 GLM-based、GNN-based 与传统向量检索 retriever 的相对优势,指出微调后的 GLM retriever 具有更好的跨域泛化能力。This work introduces a GLM-based retriever and investigates the comparative strengths of GLM-based, GNN-based, and traditional vector-search-based retrievers in single- and multi-hop RAG settings, and suggests that finetuned GLM retrievers generalize better out of domain.

ConMem: Contribution-Aware Memory for Long-Horizon Manufacturing Inspection Logs
ConMem:面向长周期制造巡检日志的贡献感知记忆
arXiv:2607.28126 RAG 检索增强 方法 OA · 绿色 被引 7 · S2

提出 ConMem,一个面向 LLM 辅助设备巡检的贡献感知 memory 框架,支持人在环的早期风险筛查,并在受限 memory 预算下保留高价值证据。This work proposes ConMem, a contribution-aware memory framework for LLM-assisted equipment inspection, supporting a human-in-the-loop early-risk screening system and retaining high-value evidence under a constrained memory budget.

OptGraph: Large Language Models Enhanced Evolutionary Optimization Via Graph Retrieval-Augmented Generation
OptGraph:通过图 RAG 增强的大语言模型进化优化
arXiv:2607.27918 RAG 检索增强 方法 OA · 绿色 被引 1 · S2

OptGraph 是首个引入 GraphRAG 的优化 agentic workflow,首次将可复用经验构建为类型化 graph,刻画建模模式、问题形式化、实现细节与错误修正之间的关系。OptGraph is the first optimization agentic workflow that introduces graph retrieval-augmented generation (GraphRAG) and first constructs reusable experience as a typed graph, capturing the relationships among modeling patterns, problem formalization, implementation details, and error corrections.

Bridging the Question-Answer Gap in Retrieval-Augmented Generation: Hypothetical Prompt Embeddings
弥合 RAG 中的问答差距:假设提示嵌入
arXiv:2607.29402 RAG 检索增强 方法 被引 9 · S2

提出Hypothetical Prompt Embeddings (HyPE),将假设内容的生成从查询阶段前移到索引阶段,把检索转化为问题-问题匹配任务,无需运行时合成答案生成。This work proposes Hypothetical Prompt Embeddings (HyPE), a framework that shifts the generation of hypothetical content from query time to the indexing phase, and transforms retrieval into a question-question matching task, bypassing the need for runtime synthetic answer generation.

UEmbed: Unified Sparse and Dense Multimodal Embeddings
UEmbed:统一的稀疏与密集多模态 Embedding
arXiv:2608.02583 RAG 检索增强 方法 OA · 绿色 被引 2 · S2

UEmbed (Unified Embedding)是一种decoder-only多模态嵌入模型,在单次因果前向中同时产出稀疏词项与稠密表示,提供新范式:在单一模型中统一稠密与稀疏嵌入,并将稀疏检索扩展以统一文本与多模态输入。UEmbed (Unified Embedding), a decoder-only multimodal embedding model that produces both sparse lexical and dense representations in one causal forward pass, offers a new paradigm: it unifies dense and sparse embeddings in one model, while further extending sparse retrieval to unify text and multimodal inputs.

Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation
更优、更强、更快、更广:面向基于 MLLM 分割的结构化全 Mask 预测
arXiv:2608.02791 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

STAMPlus 解决了超出单目标预测的三难问题,解耦自回归对话与非自回归 mask 预测,取得 SOTA 分割性能,同时保持通用多模态指令遵循能力,并降低 12 类别延迟。STAMPlus resolves the trilemma beyond single-target prediction, decoupling autoregressive dialogue from non-autoregressive mask prediction and achieves state-of-the-art segmentation performance, preserves general multimodal instruction following, and reduces 12-category latency.