总体而言,分阶段检索的影响往往超过"理想证据存在"本身;本文为专业科学场景下 RAG 系统的部署与诊断提供了实践指导,并为构建更可靠、可控的迭代式检索-推理框架奠定了基础。This is the first controlled, mechanism-level diagnostic evaluation of whether synchronized iterative retrieval and reasoning can surpass even an idealized static upper bound (Gold Context) RAG, and practical guidance for deploying and diagnosing RAG in specialized scientific settings.
论文
19 张论文卡片 · RAG 检索增强 · 应用落地
本文介绍 TianoForge,这是面向 TianoCore 开源 UEFI 固件开发生态中 bug 分诊的集成方案,部署 AI(具体为机器学习)领域的 SOTA 方法以实现自动化 bug 分诊。This integrated approach to bug triage in the TianoCore open-source UEFI firmware development ecosystem, called TianoForge, deploys the state of the art in artificial intelligence, specifically machine learning, to enable automated bug triage.
本文通过全面超越传统开环基线,证明了当前主流的单体上下文扩展策略是一种因相关性衰减而受到惩罚的架构陷阱,并确立了以顺序、反馈驱动的编排作为生成式搜索的确定性范式。By dominating classical open-loop baselines, this work proves that the prevailing strategy of monolithic context widening is an architectural trap penalized by relevance decay and establishes sequential, feedback-driven orchestration as the definitive paradigm for generative search.
介绍 WeMM-Embedding,一族通用多模态嵌入模型,支持文本、图像、视频、视觉文档及任意交错的多模态输入,输出维度灵活,在多个公开基准上取得 SOTA 表现。WeMM-Embedding is presented, a family of universal multimodal embedding models supporting text, images, videos, visual documents, and arbitrarily interleaved multimodal inputs with flexible output dimensions and achieves leading performance on multiple public benchmarks.
RAGScope 是一个泄漏受控的评估协议,用于评估仅使用任务输入、检索上下文和答案文本的本地证据门,结合了上下文分组划分、折范围预处理、组自举区间、部署工作点、端到端运行时以及显式的源偏移压力测试。RAGScope, a leakage-controlled protocol for evaluating local evidence gates that use only the task input, retrieved context, and answer text is presented, which combines context-grouped splits, fold-scoped preprocessing, group bootstrap intervals, deployment operating points, end-to-end runtime, and explicit source-shift stress tests.
恶意软件防御通常会尽快移除或隔离可疑程序。该策略虽便于遏制,却也浪费了观察攻击者行为与部署针对性反制的机会。ORCAGen 另辟蹊径:离线利用 GenAI 构建针对特定恶意软件的欺骗剧本,部署前进行校验,运行时仅强制执行已验证的逻辑。ORCAGen 将 RAG 与结构化提示工程相结合,以同时生成 PoC 恶意软件及对应的欺骗编排代码。Malware defenses often remove or isolate suspicious programs as quickly as possible. While effective for containment, this approach can also waste an opportunity to observe attacker behavior and deploy targeted countermeasures. ORCAGen takes a different approach: it uses GenAI to build malware-specific deception playbooks offline, validates them before deployment, and enforces only the verified logic at runtime. ORCAGen combines Retrieval-Augmented Generation (RAG) with structured prompt engineering to generate both proof-of-concept (PoC) malware and corresponding deception orchestration code.
LLM 可在前缀式提取下泄露记忆化的训练序列:给定训练样本的前缀,模型可能为原始续写赋予高概率。但在部署系统中,前缀很少被单独评估,常与指令、检索文档或其他任务相关上下文一同出现,正如 RAG 所做的那样。这促使我们去考察:上下文条件化究竟是缓解了记忆化,还是仅仅改变了可被提取的记忆化样本集合。本文对这一问题展开研究……Large language models (LLMs) can expose memorized training sequences under prefix-based extraction: given a prefix from a training example, the model may assign high probability to the original continuation. In deployed systems, however, prefixes are rarely evaluated in isolation. They often appear together with instructions, retrieved documents, or other task-specific context, as in retrieval-augmented generation (RAG). This motivates examining whether contextual conditioning mitigates memorization or merely changes the set of memorized samples that become extractable. We investigate this iss
本文命名了"检索状态锁定"这一失败模式,通过分离单一置信度分数所混淆的三个对象——答案表面、检索到的证据以及检索状态本身——来诊断该问题,并直接衡量"一致性盲区"。This work names the failure retrieval-state lock-in and diagnose it by separating the three objects a single confidence score conflates: the answer surface, the retrieved evidence, and the retrieval state itself, and measures the agreement blind spot directly.
FRAMe 展示了先进 LLM 如何被部署用于以人为本的任务规划,将自然语言指令转化为安全、高效且灵活的飞行路线。FRAMe signifies how advanced LLMs can be deployed for human-centric mission planning, translating natural language instructions into safe, efficient, and flexible flight routes.
提出 AgentKGV,一种用于知识图谱事实核查的智能体 LLM-RAG 框架,集成动态路由与迭代查询改写,以应对文档级检索中的表层形式不匹配问题。AgentKGV, the Agentic LLM-RAG framework for KG fact Verification, is proposed, that integrates dynamic routing and iterative query rewriting, which handles surface-form mismatch in document-level retrieval.
Cross-encoder 在 RAG 流水线中具有较高的重排序准确率,但推理成本随序列长度呈二次增长,难以实时部署。本文通过两阶段流水线解决该问题:使用 Unsloth 框架与 LoRA 适配器,在自定义的查询-文档相关性数据集上对 LLaMA 3 (8B) 进行监督微调,随后进行 4-bit 量化以提升推理效率。该模型可替换双路检索 RAG 流水线中结合 BM25 与稠密向量检索的 cross-encoder,并在特定领域问答……Cross-encoders achieve high reranking accuracy in Retrieval-Augmented Generation (RAG) pipelines but impose quadratic inference costs that limit real-time deployment. We address this by fine-tuning LLaMA 3 (8B) as a drop-in reranker using a two-stage pipeline: supervised fine-tuning on a custom query-document relevance dataset via the Unsloth framework with LoRA adapters, followed by 4-bit quantization for efficient inference. The resulting model replaces the cross-encoder in a dual-retriever RAG pipeline combining BM25 and dense vector search. Evaluated on a domain-specific question-answering
本文提出 Chunk Coverage (CC),一种独立于 oracle 的 RAG 系统检索组件测试充分性准则,结果表明 CC 在无需测试 oracle 的情况下捕获了与有效测试相关的检索多样性。Chunk Coverage (CC), an oracle-independent test adequacy criterion for testing the retrieval component of RAG systems, is introduced and results show that CC captures retrieval diversity relevant to effective testing without requiring test oracles.
已部署的平台与其面向运维的评估共同构成了一条可信赖、统计上可靠的 AI 辅助工作流,适用于设施运维,并可推广到其他大型科学仪器。Together, the deployed platform and its operations-grounded evaluation present a promising workflow for trustworthy, statistically grounded AI assistance in facility operations, transferable to other large scientific instruments.
结果表明,RAG 增强的 LLM 能显著提升生成文本的事实一致性、领域专属性与规范精度,同时降低产生无支持内容的风险;本地部署的 RAG 增强 LLM 不应仅被视为文本生成工具,而应作为认知计算基础设施中的语义处理模块,在法律和信息高度动态的环境中支撑合规与组织决策。The results demonstrate that augmenting LLMs with RAG significantly improves the factual consistency, domain specificity and normative precision of generated texts while reducing the risk of unsupported content generation and indicate that locally deployed LLMs enhanced with RAG should be regarded not merely as text generation tools but as semantic processing modules within cognitive computing infrastructures supporting regulatory compliance and organizational decision-making in environments characterized by high legal and informational volatility.
提出统一框架,融合时频图结构学习与协变量感知的表示融合,证实其在建模选择性变量交互、利用协变量提升预测精度方面的有效性。This work proposes a unified framework integrating time–frequency graph structure learning with covariate-aware representation fusion, confirming its effectiveness in modeling selective variable interactions and leveraging covariates for improved forecasting accuracy.
结果表明,DEFRAG 缩小了 SLM 与 LLM 之间的精度差距,同时相比集中式服务,成本降低高达 98.4%,峰值吞吐量提升高达 97.8%,证明了 DEFRAG 在边缘实现 LLM 服务普惠化的潜力。Results show that DEFRAG narrows the SLM–LLM accuracy gap, while reducing cost by up to 98.4% and increasing peak throughput by up to 97.8% over centralized services, demonstrating the potential of DEFRAG for democratized LLM services at the edge.
TEngineDB-V通过将全局段解耦索引物化为关系表,使向量搜索成为Tencent OLAP引擎的一等分析原语,消除scatter-gather执行、降低放大效应,并支持原生存储优化。TEngineDB-V makes vector search a first-class analytical primitive in Tencent's OLAP engine through a global segment-decoupled index materialized as relational tables, eliminating scatter-gather execution, reducing amplification, and enabling native storage optimizations.
本文将 S2G-RAG 的结构化充分性-缺口判断适配到冻结的 Search-R1 流程中,并在来自 900 个不相交 HotpotQA 问题的 3009 个状态上训练了一个 Qwen3.5-2B judge,以减少检索次数同时广泛保持答案准确性。This work adapts S2G-RAG's structured sufficiency-and-gap judgment to a frozen Search-R1 pipeline and trains a Qwen3.5-2B judge on 3,009 states from 900 disjoint HotpotQA questions to reduce retrieval while broadly preserving answer accuracy.
在多个 benchmark 数据集上的实证评估表明,RAEF 在准确率和推理开销方面均优于 RAF,并且与零样本及微调基础模型的全面对比显示,RAEF 在避免微计算负担的同时取得了与微调相当或更优的性能。Empirical evaluation across multiple benchmark datasets demonstrates that RAEF outperforms RAF in both accuracy and inference overhead, and comprehensive comparisons with zero-shot and fine-tuned foundation models show that RAEF achieves competitive or superior performance to fine-tuning while avoiding its computational burden.