提出 CaSKG,一种反事实因果 Skill 图谱框架,在检索前校准程序关系,将边置信度校准定位为大规模紧凑且可执行的 Skill 检索的有效路径。CaSKG, a counterfactual-causal skill graph framework that calibrates procedural relations before retrieval, is proposed, position edge-confidence calibration as an effective route to compact and executable skill retrieval at scale.
论文
192 张论文卡片 · RAG 检索增强 · OA 绿色
本工作将 embedding 缩放作为正交于稀疏度缩放的强有力维度加以探索,并推出 LongCat-Flash-Lite,一个从零训练的 68.5B 参数、约 30 亿激活参数的模型,不仅超越参数等量级的 MoE 基线,还对同规模现有模型展现出卓越竞争力。This work explores embedding scaling as a potent, orthogonal dimension for scaling sparsity and introduces LongCat-Flash-Lite, a 68.5B parameter model with ~3B activated trained from scratch that not only surpasses parameter-equivalent MoE baselines but also exhibits exceptional competitiveness against existing models of comparable scale.
该研究将 358,896 条消息切分为 22,329 个时间连贯的块,并构建三种搜索表示:raw_text、生成的 summary,以及将 summary 与 raw_text 片段及其他固定文本相结合的 embedding_text。This study segmented 358,896 messages into 22,329 temporally coherent chunks and constructed three search representations: raw_text, a generated summary, and embedding_text, which combines a summary with a raw-text excerpt and other fixed text.
提出 CamoDocs,一种通过将对抗文档伪装在良性内容中来避免直接包含查询的投毒攻击,并表明 TrustRAG 等以擦除为主的聚类防御可降低 ASR,但会在 NeoQA 等依赖检索的基准上造成显著的效用下降。CamoDocs is proposed, a poisoning attack that avoids direct query inclusion by camouflaging adversarial documents among benign content, and shows that erasure-heavy clustering defenses such as TrustRAG can reduce ASR, but only with substantial utility drops on retrieval-dependent benchmarks such as NeoQA.
一种神经符号架构,将逻辑知识图谱(LKG)与动态求解器路由相结合,并引入基于本体的 LKG,将逻辑规则和约束视为一等拓扑节点,从而支持对从文本中抽取的依赖关系进行显式建模。A Neuro-Symbolic architecture that integrates a Logical Knowledge Graph (LKG) with dynamic solver routing, and introduces an ontology-based LKG that treats logical rules and constraints as first-class topological nodes, enabling explicit modeling of dependencies extracted from text.
该候选名单可在不牺牲检索质量的前提下短路全语料库加密搜索:200–500 个候选即可在 5 个零样本语料库(规模从 25K 到 5.4M 文档)中与全语料库检索效果接近匹配。This shortlist short-circuits full-corpus cryptographic search without sacrificing retrieval quality: with 200-500 candidates, it closely matches full-corpus retrieval across five zero-shot corpora spanning 25K to 5.4M documents.
跨数据集分析表明,语义切分在具有显式关系线索的抽取数据集(如 GM-CIHT 和 DDI)上表现更优,而固定切分在密集生化抽取和二分类场景(如 ChemProt 和 ADE)下仍具竞争力甚至更强。Cross-dataset analysis shows that semantic chunking improves extraction datasets with explicit relation cues, such as GM-CIHT and DDI, while fixed chunking remains competitive or stronger for dense biochemical extraction and binary classification settings such as ChemProt and ADE.
结果表明,即便在算力预算匹配的前提下,循环(looping)仍可提升 Transformer,提供了一种将深度复用转化为可衡量增益的实用方案。Results show that looping can improve Transformers even under budget matching, offering a practical recipe that turns depth reuse into measurable gains.
ACToR 在生成过程中识别关键 token,按需触发有针对性的检索,在这些决定性位置提供仓库上下文;并为稠密检索器设计了一种位置感知加权方法,以优先考虑对生成更具信息量的上下文。ACToR identifies critical tokens during generation and triggers targeted retrieval on demand to provide repository context at these decisive positions, and designs a position-aware weighting method for dense retrievers to prioritize context that is more informative for generation.
本文提出 SimLoss,一种面向单轮细粒度图像描述的无参考 embedding 空间目标;结果表明 embedding 空间监督能在单轮描述器延迟下恢复多阶段验证的质量。SimLoss is proposed, a reference-free embedding-space objective for single-pass fine-grained image captioning, and results show that embedding-space supervision can recover the quality of multi-stage verification at the latency of a single-pass captioner.
基于实证对 "Enhanced" 与 "Agentic" RAG 范式进行评估,为真实场景中选取最有效的 RAG 设计(兼顾性能与成本)提供指导。An empirically driven evaluation of the "Enhanced" and "Agentic" RAG paradigms is conducted, offering guidance on selecting the most effective RAG design for real-world applications, considering both performance and costs.
本文提出 SnapBench——首个面向鲁棒"拍照即问"多模态检索的配对基准,以及一种简单的自适应融合方法 MOOR(Modality-anchored, Outlier-aware, Optimal Reweighting),并指出在"拍照即问"检索中需要具备可靠性感知的模态校准。SnapBench is introduced, the first paired benchmark for robust snap-and-ask multimodal retrieval, and MOOR (Modality-anchored, Outlier-aware, Optimal Reweighting), a simple adaptive fusion approach, highlighting the need for reliability-aware modality calibration in snap-and-ask retrieval.
本文提出 ViSAR(Visual Semantic Activation Retrieval),一种面向 late-interaction 视觉文档检索的无训练自适应 k 检索方法,并表明相似度矩阵结构与答案准确率相关,为面向检索质量感知的文档理解指明了未来方向。ViSAR (Visual Semantic Activation Retrieval), a training-free adaptive-$k$ retrieval method for late-interaction visual document retrieval, is introduced and it is shown that the similarity matrix structure correlates with answer accuracy, suggesting future directions for retrieval quality-aware document understanding.
本研究探讨了相较于标准 LLM 生成方式,RAG 与命名实体识别在多大程度上能提升自动生成通俗摘要的质量、事实一致性及可读性。This study investigates the extent to which Retrieval-Augmented Generation and Named Entity Recognition improve the quality, factual consistency, and readability of automatically generated lay summaries compared with standard LLM-based generation.
本文提出 NE-R1,一种面向自适应检索增强 NER 的新框架,在多个基准上达到 SOTA 性能,域内评估平均 F1 提升 2.52%,零样本跨域评估平均 F1 提升 1.18%。This paper proposes NE-R1, a novel framework for adaptive retrieval-augmented NER, which achieves state-of-the-art performance on various benchmarks, with an average F1 score gain of 2.52% in in-domain evaluation and 1.18% in zero-shot cross-domain evaluation.
本文提出 KBMR,首个面向 KB-VQA 的基于 MLLM 的 embedding retriever,并引入一个基于 MLLM 的语义判别器以生成连续的实体一致性权重,应对维基百科规模检索中的噪声监督挑战。KBMR is proposed, the first MLLM-based embedding retriever tailored for KB-VQA, and an MLLM-based semantic discriminator that generates continuous entity-consistency weights is introduced to tackle the challenge of noisy supervision in Wikipedia-scale retrieval.
Sparse Readout Prism (SRP) 仅使用 readout 的权重对其进行分解,将任意 token logit 或 logit 差表示为来自稀疏 readout 特征贡献之和,揭示了 readout 特征作为 lens 解读新单元的价值,暴露出 token 身份可能掩盖的结构,并支持跨 token、上下文、层与 lens 的比较。Sarse Readout Prism (SRP), which decomposes the readout using only its weights and expresses any token logit or logit difference as a sum of contributions from sparse readout features, reveals readout features as a new unit of analysis for lens readings, exposing structure that token identities can obscure and enabling comparisons across tokens, contexts, layers, and lenses.
该工作提出 LatentStream,一种渐进式 latent working memory 框架,将流式记忆从"存储-检索"转变为"检索-内化",在现有在线和离线视频 benchmark 上取得新的 SOTA 结果。This work introduces LatentStream, a progressive latent working memory framework that shifts streaming memory from store-and-retrieve to retrieve-and-internalize, and achieves new state-of-the-art results on existing online and offline video benchmarks.
提出 PACE 数据集,用于评估模型能否识别以自我中心知识或事件形式表达的潜在约束(这些约束使看似合理的用户请求变得不当),以及 PaceMaker 多 Agent 框架,其中专门 Agent 通过查询重构、多跳图遍历与冲突感知过滤进行协调,以检索上下文决定性证据。PACE is introduced, a dataset for evaluating whether models can identify latent constraints, expressed as egocentric knowledge or events, that render seemingly reasonable user requests inappropriate, and PaceMaker, a multi-agent framework in which specialized agents coordinate across query reformulation, multi-hop graph traversal, and conflict-aware filtering to retrieve contextually decisive evidence.
本工作提出 RoboTok,一个可扩展的数据引擎:利用人类操作视频作为查询,从互联网检索与操作相关的演示以训练灵巧机器人策略,并从以演员为中心的参考系下估计的 3D 手部轨迹中学习一个潜在运动空间This work introduces RoboTok, a scalable data engine that uses a query human manipulation video to retrieve manipulation-relevant internet demonstrations for training dexterous robot policies and learns a latent motion space from 3D hand trajectories expressed in estimated actor-centered reference frames.
我们提出 ENEAS,一种用于实例追踪与语义发现的统一且文本可提示的方法。包括 SAM 3 在内的文本可提示分割模型仍存在时间幻觉、空间碎片化与语义误分类问题:目标离开视野时无法报告目标缺失;极端特写下只分割局部纹理而非完整目标;将视觉特征置于本体事实之上,从而把雕像、绘画或反射等视觉相似的物体误分割为目标。We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-promptable segmentation models, including the latest foundation models such as SAM 3, still suffer from temporal hallucinations, spatial fragmentation, and semantic misclassification: they fail to report target absence when an object leaves the field of view, segment local textures instead of the complete object during extreme close-ups, and prioritize visual features over ontological reality, so that visually similar artifacts such as statues, paintings, or reflections are segmented as targe
本文提出 Conformal Relevance 框架,利用上下文学习的示例筛选与集成构造打分函数,在保持覆盖的同时以极低人工成本提升简洁性。The Conformal Relevance framework is introduced which uses in-context learning example curation and ensembling to create a score function which maintains coverage while improving conciseness with minimal manual input.
Generative Late-Interaction Embeddings(GLIE):从归一化质心中学习每个页面 k<<N 个向量,既作为轻量级索引,也作为重建页面完整嵌入集的基础,解码器是其主要设计面。Generative Late-Interaction Embeddings (GLIE): k<<N vectors per page learned from the normalized centroids to serve as both a lightweight index and a basis for regenerating the page's full embedding set, with the decoder as its main design surface.
一个简单、无训练的框架,其中具备推理能力的视觉语言模型在 Wikipedia 上迭代搜索与推理,动态收集证据,表明推理与检索在稀有实体上具有互补性。A simple, training-free framework in which a reasoning-capable vision-language model iteratively searches and reasons over Wikipedia, gathering evidence dynamically, shows that reasoning and retrieval are complementary on rare entities.
提出面向检索增强生成(RAG)的通用微调方案——RAG 模型融合预训练参数化记忆与非参数化记忆进行语言生成;研究发现,相较 SOTA 的纯参数化 seq2seq 基线,RAG 模型生成的文本更具针对性、更多样且更符合事实。A general-purpose fine-tuning recipe for retrieval-augmented generation (RAG) -- models which combine pre-trained parametric and non-parametric memory for language generation, and finds that RAG models generate more specific, diverse and factual language than a state-of-the-art parametric-only seq2seq baseline.
提出结构感知的 RAG 框架 ReMoMask-2:耦合 Hierarchical Bidirectional Momentum 对比学习以对齐全局与部件级特征与文本;采用 Semantic Spatial-Temporal Attention (SSTA) 实现拓扑感知的融合;通过 Topology Structured Masking (TSM) 借助自适应掩码强化鲁棒的部件级 grounding。ReMoMask-2, a structure-aware RAG framework coupling Hierarchical Bidirectional Momentum contrastive learning to align global and part-level features with text; Semantic Spatial-Temporal Attention (SSTA) for topology-aware fusion; and Topology Structured Masking (TSM) to force robust part-level grounding via adaptive masking are introduced.
外部评估显示,尽管引用有效性保持稳健,但在领域偏移下证据利用、片段对齐与拒答校准变得更加困难,表明可信的 RAG 系统需要在检索与最终答案交付之间进行显式验证。External evaluation shows that while citation validity remains strong, evidence utilization, span alignment, and refusal calibration become harder under domain shift, indicating that trustworthy RAG systems require explicit validation between retrieval and final answer delivery.
该工作提出了 SCoRE(Selection and Consolidation for Robust Evidence),一个用于显式证据选择与整合的统一 agent 循环,将最终推理与探索式试错解耦,并通过索引化的声明-图像关联确保严格的视觉锚定。This work proposes SCoRE (Selection and Consolidation for Robust Evidence), a unified agent loop for explicit evidence selection and consolidation, which decouples final reasoning from exploratory trial-and-error while ensuring strict visual grounding via indexed claim-to-image linkages.
检索增强生成(RAG)管线通常依赖在预处理阶段确定的固定索引与检索配置。这种一刀切的设计难以适配领域专家场景,因为异构查询需要不同的分块粒度、元数据约束与来源选择策略。因此,针对某一类查询有效的配置,往往在其他类查询上表现欠佳。本文提出 ORDER(Optimal Routing for Dynamic Evidence Retrieval),一种查询条件化的 RAG 框架,可联合自适应地调整索引与……Retrieval-Augmented Generation (RAG) pipelines typically rely on a fixed indexing and retrieval configuration determined at preprocessing time. This one-size-fits-all design is ill-suited to domain-expert settings, where heterogeneous queries require different chunking granularities, metadata constraints, and source-selection strategies. As a result, configurations that are effective for one family of queries often perform poorly for others. In this paper, we introduce ORDER (Optimal Routing for Dynamic Evidence Retrieval), a query-conditioned RAG framework that jointly adapts indexing and ret
本文重新审视联合多模态表示学习与生成,旨在产生可直接被生成式解码器使用的线性可插值嵌入,并确保其表示同时充当判别性语义描述符和生成条件。This work revisits joint multimodal representation learning and generation to produce linearly interpolatable embeddings that are directly consumable by generative decoders and ensures its representations function as both discriminative semantic descriptors and generative conditions.
本文提出 SpectralShift,一种用于 GDN 长上下文持续预训练的谱重参数化方法,通过重参数化 alpha 投影的初始化以重塑衰减谱,增强慢传播能力,并进一步引入 alpha 投影的学习率缩放以促进长上下文训练。SpectralShift is proposed, a spectral reparameterization approach for long-context continual pretraining of GDNs that reparameterizes the alpha projections initialization to reshape the decay spectrum by enhancing slow propagation capacity, and further introduces a learning-rate scaling for alpha projections to facilitate long-context training.
Mixture of Memory Embeddings (MoME):一种上下文感知的记忆机制,将每个 token 的单一记忆行替换为 M 个 slot 的混合,并通过隐藏状态上的可学习门控选择在每个位置读取哪些 slot;在 sub-billion 规模下展现出更优的记忆容量 scaling 趋势,且训练与推理均保持高效。Mixture of Memory Embeddings (MoME), a context-aware memory mechanism that replaces each token's single memory row with a mixture of M slots and uses a learned gate over the hidden state to choose which slots to read at each position, shows more promising memory-size scaling trend at sub-billion scale and remains efficient in training and inference.
本文将 Agentic 检索-生成循环形式化为有限时域部分可观测马尔可夫决策过程,显式建模其控制策略与状态转移,并构建了全面的分类体系与模块化架构分解,按规划机制、检索编排、记忆范式与工具调用行为对系统进行分类。This paper formalizes agentic retrieval-generation loops as finite-horizon partially observable Markov decision processes, explicitly modeling their control policies and state transitions, and develops a comprehensive taxonomy and modular architectural decomposition that categorizes systems by their planning mechanisms, retrieval orchestration, memory paradigms, and tool-invocation behaviors.
提出 Document Retrieval-Aware Chunking (D-RAC),将 Web Retrieval-Aware Chunking (W-RAC) 框架扩展至任意文档格式,保留 W-RAC 的成本、确定性与可观测性优势,同时将每种可渲染格式解锁为一类输入。Document Retrieval-Aware Chunking (D-RAC), an extension of the Web Retrieval-Aware Chunking (W-RAC) framework to arbitrary document formats, is presented, preserving W-RAC's cost, determinism, and observability benefits while unlocking every renderable format as a first-class input.
提出 Functionalizer,一种无损预分词框架,将正字与结构变体分解为分词前的可组合操作码/操作数前缀流:以 Unicode 私有使用区编码的参数化变换操作符(操作码)为前缀,连接规范基础 token(操作数)。The Functionalizer is presented, a lossless pre-tokenizer framework that factors orthographic and structural variations into a compositional opcode/operand prefix stream before tokenization: a canonical base token (operand) prefixed by parametric transformation operators (opcodes) encoded in the Unicode Private Use Area.
一种近期且有效的方法:通过一个小型 low-rank adapter,将大型预训练图像编辑模型适配到图像恢复任务,并以源自退化图像本身的 instruction 替代 text prompt;在同等条件下该方法优于文本条件。A recent and effective recipe adapts a large pretrained image-editing model to restoration using a small low-rank adapter with a text prompt to an instruction derived from the degraded image itself that outperforms text conditioning under a matched comparison.