本文提出 NE-R1,一种面向自适应检索增强 NER 的新框架,在多个基准上达到 SOTA 性能,域内评估平均 F1 提升 2.52%,零样本跨域评估平均 F1 提升 1.18%。This paper proposes NE-R1, a novel framework for adaptive retrieval-augmented NER, which achieves state-of-the-art performance on various benchmarks, with an average F1 score gain of 2.52% in in-domain evaluation and 1.18% in zero-shot cross-domain evaluation.
论文
165 张论文卡片 · RAG 检索增强 · 方法
本文提出 KBMR,首个面向 KB-VQA 的基于 MLLM 的 embedding retriever,并引入一个基于 MLLM 的语义判别器以生成连续的实体一致性权重,应对维基百科规模检索中的噪声监督挑战。KBMR is proposed, the first MLLM-based embedding retriever tailored for KB-VQA, and an MLLM-based semantic discriminator that generates continuous entity-consistency weights is introduced to tackle the challenge of noisy supervision in Wikipedia-scale retrieval.
Sparse Readout Prism (SRP) 仅使用 readout 的权重对其进行分解,将任意 token logit 或 logit 差表示为来自稀疏 readout 特征贡献之和,揭示了 readout 特征作为 lens 解读新单元的价值,暴露出 token 身份可能掩盖的结构,并支持跨 token、上下文、层与 lens 的比较。Sarse Readout Prism (SRP), which decomposes the readout using only its weights and expresses any token logit or logit difference as a sum of contributions from sparse readout features, reveals readout features as a new unit of analysis for lens readings, exposing structure that token identities can obscure and enabling comparisons across tokens, contexts, layers, and lenses.
空间资产定价模型将企业间交互结构视为已知,并利用语言模型表征从企业的信息环境中推断该结构;语言模型表征充当资本市场中潜在企业间信息结构的测量工具。Spatial asset-pricing models take the structure of inter-firm interaction as given and infer that structure from firms' information environments using language-model representations, which serve as a measurement instrument for latent inter-firm information structure in capital markets.
该工作提出 LatentStream,一种渐进式 latent working memory 框架,将流式记忆从"存储-检索"转变为"检索-内化",在现有在线和离线视频 benchmark 上取得新的 SOTA 结果。This work introduces LatentStream, a progressive latent working memory framework that shifts streaming memory from store-and-retrieve to retrieve-and-internalize, and achieves new state-of-the-art results on existing online and offline video benchmarks.
提出 PACE 数据集,用于评估模型能否识别以自我中心知识或事件形式表达的潜在约束(这些约束使看似合理的用户请求变得不当),以及 PaceMaker 多 Agent 框架,其中专门 Agent 通过查询重构、多跳图遍历与冲突感知过滤进行协调,以检索上下文决定性证据。PACE is introduced, a dataset for evaluating whether models can identify latent constraints, expressed as egocentric knowledge or events, that render seemingly reasonable user requests inappropriate, and PaceMaker, a multi-agent framework in which specialized agents coordinate across query reformulation, multi-hop graph traversal, and conflict-aware filtering to retrieve contextually decisive evidence.
本工作提出 RoboTok,一个可扩展的数据引擎:利用人类操作视频作为查询,从互联网检索与操作相关的演示以训练灵巧机器人策略,并从以演员为中心的参考系下估计的 3D 手部轨迹中学习一个潜在运动空间This work introduces RoboTok, a scalable data engine that uses a query human manipulation video to retrieve manipulation-relevant internet demonstrations for training dexterous robot policies and learns a latent motion space from 3D hand trajectories expressed in estimated actor-centered reference frames.
研究结果强调了方向特定的迁移测试、严格的 embedding space 隔离,以及在 memory migrations 中为 memory repair 保留源历史的必要性。Findings highlight the necessity of direction-specific migration testing, strict embedding space isolation, and the retention of source histories for memory repair in memory migrations.
我们提出 ENEAS,一种用于实例追踪与语义发现的统一且文本可提示的方法。包括 SAM 3 在内的文本可提示分割模型仍存在时间幻觉、空间碎片化与语义误分类问题:目标离开视野时无法报告目标缺失;极端特写下只分割局部纹理而非完整目标;将视觉特征置于本体事实之上,从而把雕像、绘画或反射等视觉相似的物体误分割为目标。We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-promptable segmentation models, including the latest foundation models such as SAM 3, still suffer from temporal hallucinations, spatial fragmentation, and semantic misclassification: they fail to report target absence when an object leaves the field of view, segment local textures instead of the complete object during extreme close-ups, and prioritize visual features over ontological reality, so that visually similar artifacts such as statues, paintings, or reflections are segmented as targe
本文提出 Conformal Relevance 框架,利用上下文学习的示例筛选与集成构造打分函数,在保持覆盖的同时以极低人工成本提升简洁性。The Conformal Relevance framework is introduced which uses in-context learning example curation and ensembling to create a score function which maintains coverage while improving conciseness with minimal manual input.
Generative Late-Interaction Embeddings(GLIE):从归一化质心中学习每个页面 k<<N 个向量,既作为轻量级索引,也作为重建页面完整嵌入集的基础,解码器是其主要设计面。Generative Late-Interaction Embeddings (GLIE): k<<N vectors per page learned from the normalized centroids to serve as both a lightweight index and a basis for regenerating the page's full embedding set, with the decoder as its main design surface.
一个简单、无训练的框架,其中具备推理能力的视觉语言模型在 Wikipedia 上迭代搜索与推理,动态收集证据,表明推理与检索在稀有实体上具有互补性。A simple, training-free framework in which a reasoning-capable vision-language model iteratively searches and reasons over Wikipedia, gathering evidence dynamically, shows that reasoning and retrieval are complementary on rare entities.
提出面向检索增强生成(RAG)的通用微调方案——RAG 模型融合预训练参数化记忆与非参数化记忆进行语言生成;研究发现,相较 SOTA 的纯参数化 seq2seq 基线,RAG 模型生成的文本更具针对性、更多样且更符合事实。A general-purpose fine-tuning recipe for retrieval-augmented generation (RAG) -- models which combine pre-trained parametric and non-parametric memory for language generation, and finds that RAG models generate more specific, diverse and factual language than a state-of-the-art parametric-only seq2seq baseline.
提出结构感知的 RAG 框架 ReMoMask-2:耦合 Hierarchical Bidirectional Momentum 对比学习以对齐全局与部件级特征与文本;采用 Semantic Spatial-Temporal Attention (SSTA) 实现拓扑感知的融合;通过 Topology Structured Masking (TSM) 借助自适应掩码强化鲁棒的部件级 grounding。ReMoMask-2, a structure-aware RAG framework coupling Hierarchical Bidirectional Momentum contrastive learning to align global and part-level features with text; Semantic Spatial-Temporal Attention (SSTA) for topology-aware fusion; and Topology Structured Masking (TSM) to force robust part-level grounding via adaptive masking are introduced.
外部评估显示,尽管引用有效性保持稳健,但在领域偏移下证据利用、片段对齐与拒答校准变得更加困难,表明可信的 RAG 系统需要在检索与最终答案交付之间进行显式验证。External evaluation shows that while citation validity remains strong, evidence utilization, span alignment, and refusal calibration become harder under domain shift, indicating that trustworthy RAG systems require explicit validation between retrieval and final answer delivery.
该工作提出了 SCoRE(Selection and Consolidation for Robust Evidence),一个用于显式证据选择与整合的统一 agent 循环,将最终推理与探索式试错解耦,并通过索引化的声明-图像关联确保严格的视觉锚定。This work proposes SCoRE (Selection and Consolidation for Robust Evidence), a unified agent loop for explicit evidence selection and consolidation, which decouples final reasoning from exploratory trial-and-error while ensuring strict visual grounding via indexed claim-to-image linkages.
检索增强生成(RAG)管线通常依赖在预处理阶段确定的固定索引与检索配置。这种一刀切的设计难以适配领域专家场景,因为异构查询需要不同的分块粒度、元数据约束与来源选择策略。因此,针对某一类查询有效的配置,往往在其他类查询上表现欠佳。本文提出 ORDER(Optimal Routing for Dynamic Evidence Retrieval),一种查询条件化的 RAG 框架,可联合自适应地调整索引与……Retrieval-Augmented Generation (RAG) pipelines typically rely on a fixed indexing and retrieval configuration determined at preprocessing time. This one-size-fits-all design is ill-suited to domain-expert settings, where heterogeneous queries require different chunking granularities, metadata constraints, and source-selection strategies. As a result, configurations that are effective for one family of queries often perform poorly for others. In this paper, we introduce ORDER (Optimal Routing for Dynamic Evidence Retrieval), a query-conditioned RAG framework that jointly adapts indexing and ret
论文提出了 InceptionRAG,一种颠覆 RAG 标准攻击范式的隐蔽攻击机制,展现出更优的规避能力,能够有效绕过针对传统单文档注入的既有防御。This paper introduces InceptionRAG, a stealthy attack mechanism that subverts the standard attack paradigm of RAG, and shows superior evasion capabilities, effectively bypassing established defenses that mitigate traditional single-document injections.
本文重新审视联合多模态表示学习与生成,旨在产生可直接被生成式解码器使用的线性可插值嵌入,并确保其表示同时充当判别性语义描述符和生成条件。This work revisits joint multimodal representation learning and generation to produce linearly interpolatable embeddings that are directly consumable by generative decoders and ensures its representations function as both discriminative semantic descriptors and generative conditions.
本文提出 SpectralShift,一种用于 GDN 长上下文持续预训练的谱重参数化方法,通过重参数化 alpha 投影的初始化以重塑衰减谱,增强慢传播能力,并进一步引入 alpha 投影的学习率缩放以促进长上下文训练。SpectralShift is proposed, a spectral reparameterization approach for long-context continual pretraining of GDNs that reparameterizes the alpha projections initialization to reshape the decay spectrum by enhancing slow propagation capacity, and further introduces a learning-rate scaling for alpha projections to facilitate long-context training.
本文提出 SELF-INDEX,一个无需人工干预即可让索引自我演进的框架;其 Optimizer 可自主诊断检索短板,选择性修改对应的索引键,并在每次更新前对每项修订进行验证。SELF-INDEX is proposed, a framework that enables an index to self-evolve without human intervention, and its Optimizer autonomously diagnoses retrieval shortfalls, selectively revises the responsible index keys, and validates each revision before updating the index.
Mixture of Memory Embeddings (MoME):一种上下文感知的记忆机制,将每个 token 的单一记忆行替换为 M 个 slot 的混合,并通过隐藏状态上的可学习门控选择在每个位置读取哪些 slot;在 sub-billion 规模下展现出更优的记忆容量 scaling 趋势,且训练与推理均保持高效。Mixture of Memory Embeddings (MoME), a context-aware memory mechanism that replaces each token's single memory row with a mixture of M slots and uses a learned gate over the hidden state to choose which slots to read at each position, shows more promising memory-size scaling trend at sub-billion scale and remains efficient in training and inference.
提出 Document Retrieval-Aware Chunking (D-RAC),将 Web Retrieval-Aware Chunking (W-RAC) 框架扩展至任意文档格式,保留 W-RAC 的成本、确定性与可观测性优势,同时将每种可渲染格式解锁为一类输入。Document Retrieval-Aware Chunking (D-RAC), an extension of the Web Retrieval-Aware Chunking (W-RAC) framework to arbitrary document formats, is presented, preserving W-RAC's cost, determinism, and observability benefits while unlocking every renderable format as a first-class input.
一种近期且有效的方法:通过一个小型 low-rank adapter,将大型预训练图像编辑模型适配到图像恢复任务,并以源自退化图像本身的 instruction 替代 text prompt;在同等条件下该方法优于文本条件。A recent and effective recipe adapts a large pretrained image-editing model to restoration using a small low-rank adapter with a text prompt to an instruction derived from the degraded image itself that outperforms text conditioning under a matched comparison.
结果支持将生成式记忆视为对直接检索的选择性修正,并强调何时、以何种方式、以何种强度进行路由干预是核心挑战。The results support generated memory as a selective correction to direct retrieval and highlight routing when, which, and how strongly to intervene as the central challenge.
本文提出 ZooWork-ShopRanker,一族对齐到裁判标注购物偏好的电商 reranker,以及 ShopRank-Bench,一个包含约 10,000 条私有流量偏好对的低污染 benchmark,按承诺该标注的裁判家族数量分档呈现,覆盖多种文本格式。ZooWork-ShopRanker, a family of e-commerce rerankers aligned to judge-labeled shopping preference, and ShopRank-Bench, a contamination-limited benchmark of ~10,000 private-traffic preference pairs in both text formats, tiered by how many judge families committed to each label.
QReason 是一种解耦框架,将面向 query 的推理与针对窗口的 passage 相关性评估分离,显著减少冗余推理,在取得与 reasoning-based reranker 相当乃至更优的排序性能的同时,超越了现有的 query rewriting 模型。QReason is a decoupled framework that separates query-focused reasoning from window-specific passage relevance assessment, and significantly reduces redundant reasoning, achieves ranking performance comparable to or better than strong reasoning-based rerankers, and outperforms existing query rewriting models.
SarseX 模型无关、无需训练,并与 Prefix Cache 兼容,可为多轮对话、检索增强生成 (RAG) 和 Agent 工作流等常见在线服务场景提供统一支持。SarseX is model-agnostic, training-free, and compatible with Prefix Cache, and it provides unified support for common online serving scenarios including multi-round chat, retrieval-augmented generation (RAG), and agent workflows.
BoundaryMORPH 是一种新算法,专门为 LLM 的上下文容量 k 分配 CE 预算,在多个模型与数据集上针对开放性查询取得了 SOTA 集合检索质量。BoundaryMORPH is introduced, a novel algorithm that allocates CE budget specifically for the LLM's context capacity $k$ and achieves state-of-the-art set retrieval quality across multiple models and datasets with open-ended queries.
本文提出一个新颖的初步框架,可定量评估输出格式不同(如列表与消息)的检索系统的 IR 准确度,为客观评估基于关键词与基于语义的对话式检索方法奠定坚实基础。This work introduces a novel, preliminary framework to quantitatively assess IR accuracy of search systems that produce different output formats, such as lists and messages, and provides a strong foundation for objectively assessing keyword-based and semantic chat-based search methods.
本文提出一个 Agentic RAG 框架,使 LLM 能够使用逻辑表达式构建检索意图,同时将检索后端简化为基于倒排索引的系统,并表明将检索过程锚定在逻辑查询上可显著降低生成响应中的幻觉。This paper proposes an agentic RAG framework that enables LLMs to formulate retrieval intents using logical expressions while simplifying the retrieval backend to an inverted-index-based system, and shows that anchoring the retrieval process in logical queries substantially reduces hallucinations in generated responses.
提出一种自适应、自动化的数据提取攻击流程,在黑盒设置下针对 MRAG(其中检索到的视觉产物本身就是答案)发起攻击,表明亟需专门面向多模态数据设计的安全防护。An adaptive and automatic data extraction attack procedure operating in a black box setting against MRAG, a configuration in which the retrieved visual artifact is itself the response, and shows the urgent need for safeguards specifically designed for multimodal data.
提出 MMAgent-R$^2$,一种将视觉重排序与主动拒答作为内部验证机制的 Agentic mRAG 框架,并通过 GRPO 训练实现外部检索、内部验证与答案生成的联合优化。MMAgent-R$^2$, an agentic mRAG framework that integrates visual reranking and active rejection as its internal verification mechanism, is proposed and achieves joint optimization of external retrieval, internal verification, and answer generation via GRPO training.
MatRAG 在检索质量上优于其最强的竞争者;同时,通过避免 KG 构建和基于 LLM 的摘要降低了索引成本,并通过维度感知的相似度降低了查询成本。MatRAG outperforms its strongest competitors in terms of retrieval quality; furthermore, it reduces indexing costs by avoiding KG construction and LLM-based summarization, and lowers query-time costs through dimension-aware similarity.
提出一种基于不确定性感知框架的自适应问答方法,通过 LLM 内部表征中区分知识不足与知识歧义/冲突的显式信号,在单次前向传播中即可由隐状态高效估计。This work proposes an uncertainty-aware framework for adaptive QA based on explicit signals derived from LLM internal representations that distinguish between knowledge insufficiency and knowledge ambiguity or conflict, and efficiently estimate these from hidden states in a single forward pass.
在 Encyclopedic-VQA 和 InfoSeek 上的实验表明,CLIMB 持续优于基于检索增强的多模态基线;消融实验显示互补池化、基于评论家的打分以及迭代置信度控制的精炼各自对最终性能均有贡献。Experiments on Encyclopedic-VQA and InfoSeek show that CLIMB consistently improves over retrieval-augmented multimodal baselines, andlations indicate that complementary pooling, critic-based scoring, and iterative confidence-controlled refinement each contribute to the final performance.