研究库 论文知识库
Papers · organized/paper_cards

论文

1753 张论文卡片

开放获取 全部 绿色 · 1640
Uncertainty-Aware End-to-End AI Weather Forecasting: Disentangling Observation and Model Contributions
感知不确定性的端到端 AI 天气预报:分离观测与模型的贡献
arXiv:2608.30795 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

端到端天气预报系统直接从原始地球观测数据生成高质量的全球网格与站点预报,仅以数值天气预报流水线(含数据同化)一小部分的成本取而代之。这类系统是确定性的,不输出不确定性结果。本文通过在每个组件上附加一种随机机制,将 Aardvark Weather 模型概率化:在观测编码器中加入学习到的、依赖输入的噪声,以捕捉源自观测系统的偶然不确定性;在处理器中使用 Monte Carlo dropout,以捕捉认知不确定性。End-to-end weather forecasting systems produce skillful global gridded and station forecasts directly from raw Earth observations, replacing the numerical weather prediction pipeline, including data assimilation, at a fraction of its cost. These systems are deterministic and issue no uncertainty. Here we render the Aardvark Weather model probabilistic by attaching one stochastic mechanism to each component: learned, input-dependent noise at the observation encoder, capturing aleatoric uncertainty inherited from the observing system, and Monte Carlo dropout in the processor, capturing epistemic

ContextBias: Controlled Evaluation of Bias Persistence Under Context Shift in Text-to-Image Models
ContextBias:受控评估文本到图像模型在上下文偏移下的偏见持续性
arXiv:2608.29847 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

评估四个 SOTA 模型发现,将角色置于语义无关的上下文中并不会抑制该角色的关联属性;相反,跨角色的属性集中度会上升(合并 BI $+0.047$)。Evaluating four state-of-the-art models finds that placing a role in a semantically unrelated context does not suppress role-linked attributes; instead, cross-role attribute concentration increases (pooled BI $+0.047$).

SpanCalib-VLM: Calibrated Hallucination Span Detection in Vision-Language Models
SpanCalib-VLM:视觉语言模型中校准的幻觉片段检测
arXiv:2608.29974 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 SpanCalib-VLM,这是一种用于 SHROOM-Visions 共享任务的混合双系统,结合了多模态序列标注器(由 XLM-RoBERTa-Large 与 SigLIP 视觉编码器通过交叉注意力融合而成)与微调后的生成式 VLM(Qwen3.5-4B-SHROOM-SFT)。SpanCalib-VLM is presented, a hybrid dual-system for the SHROOM-Visions Shared Task that combines a multimodal sequence tagger, consisting of XLM-RoBERTa-Large fused with a SigLIP vision encoder via cross-attention, with the fine-tuned generative VLM (Qwen3.5-4B-SHROOM-SFT).

MMMMM: A Unified Taxonomy for Investigating the Mechanisms of Multilingual MultiModal Misinformation
MMMMM:用于研究多语言多模态虚假信息机制的统一分类法
arXiv:2608.29681 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

基于对数据的深入定性分析和既有理论研究,构建了一套新颖且全面的多模态错误信息分类法,并由此获得了关于社交媒体用户如何在实际场景中将图像与文本结合以传播错误信息的此前未被记录的洞见。A novel, comprehensive taxonomy of multimodal misinformation grounded in an in-depth qualitative analysis of the data and prior theoretical work is developed, which leads to previously undocumented insights about how social media users combine images with text to spread misinformation in the wild.

EvoGenUI-Bench: Evaluating LLMs as Multi-Turn Generative UI Assistants
EvoGenUI-Bench:评估 LLM 作为多轮生成式 UI 助手
arXiv:2608.29387 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 EvoGenUI-Bench,一个面向多轮界面维护的基准,包含 150 个五轮任务,共计 750 轮,覆盖三种场景:信息呈现、可执行交互和工具驱动的外部状态。EvoGenUI-Bench is introduced, a benchmark for multi-turn interface maintenance comprising 150 five-turn tasks and 750 turns across three scenarios: information presentation, executable interaction, and tool-grounded external state.

RECAP-Forcing: Retaining Content Appearances for Long Video Generation
RECAP-Forcing:面向长视频生成的内容外观保持方法
arXiv:2608.26671 多模态 方法 OA · 绿色 被引 2 · S2

本文提出 RECAP-Forcing,一种无需训练的推理方法,不增加任何可学习参数,在多个强基线上稳定提升视觉质量与语义保真度,并优于现有记忆方法。This work proposes RECAP-Forcing, a training-free inference method with no additional learnable parameters that consistently improves visual quality and semantic fidelity across multiple strong baselines and outperforms existing memory methods.

SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers
SMELT:面向计算匹配的 MoE 循环 Transformer 的扩展定律
arXiv:2609.01343 RAG 检索增强 方法 OA · 绿色 被引 13 · S2

结果表明,即便在算力预算匹配的前提下,循环(looping)仍可提升 Transformer,提供了一种将深度复用转化为可衡量增益的实用方案。Results show that looping can improve Transformers even under budget matching, offering a practical recipe that turns depth reuse into measurable gains.

Safin-1: Safety from Within through Memory-Native State Evolution
Safin-1:通过记忆原生状态演化实现内在安全
arXiv:2609.00092 安全与风险 方法 OA · 绿色 被引 0 · S2 + OpenAlex

路由状态接口在模型的原生计算中统一了上下文记忆与持久的能力适配,将记忆从对历史上下文的被动记录重塑为维持与演化模型行为的主动基质。The routed-state interface unifies contextual memory and persistent capability adaptation within the model's native computation, reframing memory from a passive record of prior context into an active substrate for maintaining and evolving model behavior.

DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory
DiagEvo:通过分层错误记忆实现的诊断引导自我演化
arXiv:2609.00768 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

本文提出 DagEvo,通过利用 self-play 中求解器的失败历史来引导问题生成,无需外部任务资源,并表明混合生成、跨状态拼接的记忆状态更新以及双置信度过滤共同贡献了这些性能提升。DagEvo is introduced, which guides question generation using the solver's failure history from self-play, without external task resources, and shows that mixed generation, memory-state updates with cross-state stitching, and double-confidence filtering contribute to these gains.

6. LLM Research Papers: The 2026 List (Jan–May) — Sebastian Raschka
6. LLM Research Papers:2026 清单(1月–5月) — Sebastian Raschka
arXiv:2603.15569 LLM 基础设施 方法 OA · 绿色 被引 99 · S2

本工作借鉴线性模型的 state space model(SSM)视角,提出三项核心方法改进并组合形成更具表达力的递推结构:源自 SSM 离散化的递推式、用于更丰富状态追踪的复数值状态更新规则,以及在不增加 decode 延迟前提下提升模型性能的多输入多输出(MIMO)建模。This work introduces three core methodological improvements inspired by the state space model (SSM) viewpoint of linear models that combine to form a more expressive recurrence derived from SSM discretization, a complex-valued state update rule that enables richer state tracking, and a multi-input, multi-output (MIMO) formulation for better model performance without increasing decode latency.

Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement
Harness-of-Harness:具备持续改进能力的多日自主软件开发
arXiv:2609.01481 评测基准 方法 OA · 绿色 被引 2 · S2

在持续多天、超过 70 轮迭代的部署中,HoH 自主开发出一款第一人称射击游戏,具备完整的主线剧情、完整实现的核心机制、可供人类游玩的体验、精美的画面与集成的音效。In a multi-day deployment with more than 70 iterations, HoH autonomously develops a first-person-shooter game, featuring a coherent storyline, fully implemented core mechanics, human-playable experience, polished visuals and integrated audio.

ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training
ZimaBlue:通过可扩展视频预训练演化出可泛化的通用世界动作模型
arXiv:2609.00188 多模态 方法 OA · 绿色 被引 2 · S2

本文提出 ZimaBlue,一个可扩展的框架,用于从大规模视频中学习可泛化的 World Action Models (WAMs),并采用异步 Slow-Fast 双系统架构,使生成式 WAM 具备面向实时控制的实用性。This work introduces ZimaBlue, a scalable framework for learning generalizable World Action Models (WAMs) from large-scale video, and adopts an asynchronous Slow-Fast dual-system architecture to make generative WAMs practical for real-time control.

Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving
Qwen-Drive-1.0:迈向自动驾驶视觉-语言基础模型的初步探索
arXiv:2609.00111 多模态 方法 OA · 绿色 被引 7 · S2

实验表明,该方法在大幅保留通用视觉-语言能力的同时,具备出色的 3D 感知与驾驶场景理解能力;在开环、伪闭环与闭环设定下的综合评估进一步显示其运动规划性能具有很强的竞争力。Experiments demonstrate strong 3D perception and driving scene understanding while largely preserving general vision-language capability and comprehensive evaluations across open-loop, pseudo-closed-loop, and closed-loop settings further show highly competitive motion-planning performance.

UI-Venus-2 Technical Report
UI-Venus-2 技术报告
arXiv:2609.00028 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

本文提出 UI-Venus-2,一个面向移动、Web 和桌面环境的通用 GUI 基础 Agent,采用统一的闭环推理-行动框架,并通过集成安全感知机制来确保关键操作的可控执行。UI-Venus-2 is presented, a general-purpose foundation GUI agent designed to operate across mobile, web, and desktop environments through a unified closed-loop reasoning-action framework that integrates safety-aware mechanisms to ensure controlled execution of consequential actions.

Adaptive Critical Token-Aware Retrieval for Repository-Level Code Generation
自适应关键 token 感知的仓库级代码生成检索
arXiv:2609.01601 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

ACToR 在生成过程中识别关键 token,按需触发有针对性的检索,在这些决定性位置提供仓库上下文;并为稠密检索器设计了一种位置感知加权方法,以优先考虑对生成更具信息量的上下文。ACToR identifies critical tokens during generation and triggers targeted retrieval on demand to provide repository context at these decisive positions, and designs a position-aware weighting method for dense retrievers to prioritize context that is more informative for generation.

From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix
从生产流量到后训练:构建覆盖企业请求组合的自托管 LLM
arXiv:2609.01572 工程化 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

数据驻留约束迫使企业自托管 LLM,但不断引入新模型而不下线旧模型会扩张服务集群,分散有限的 GPU 池。我们通过沿指令遵循、函数调用和内部任务分布三个维度,针对生产错误分析所发现的质量差距,将 200 多个内部应用的流量整合到单一模型上。质量通过按生产流量分层的离线基准进行跟踪,并由确定性验证器或经过校准的 LLM 评判器打分。不同于针对Data-residency constraints force enterprises to self-host LLMs, but continuous adoption of newer models without decommissioning their predecessors expands the serving fleet, fragmenting a finite GPU pool. We consolidate traffic from over 200 internal applications onto a single model by closing quality gaps identified through production error analysis along three axes: instruction following, function-calling, and internal task distribution. Quality is tracked by offline benchmarks stratified to production traffic and scored by deterministic verifiers or calibrated LLM judges. Rather than optimi

Recursive Criticality of AI Self-Improvement
AI 自我改进的递归临界性
arXiv:2609.00137 安全与风险 观点 OA · 绿色 被引 0 · S2 + OpenAlex

该框架识别了 AI R&D 系统的若干可测量属性,可用于区分递归放大与其他来源驱动的快速进展,包括递归反馈的强度、改进向后续系统传播的有效性、周期时长,以及进一步取得进展的难度递增。The framework identifies measurable properties of AI R\&D systems that can help distinguish recursive amplification from rapid progress driven by other sources, including the strength of recursive feedback, how effectively improvements propagate into successor systems, cycle duration, and the increasing difficulty of further progress.

Agent Memory Is a Surface for Endogenous Authorization Laundering
Agent 内存是内生授权洗钱的表层载体
arXiv:2609.01836 Agent 智能体 方法 OA · 绿色 被引 3 · S2

本工作在采购、网络安全与金融场景下评估了 5 个 LLM 作为记忆写入者、2 个 LLM 作为执行者,并提出 EAL-Bench,用于衡量持久记忆在多大程度上准确保留不断演化的授权状态,以及错误是否会向下游传播为未授权操作。This work evaluates five LLMs as memory writers and two as executors across procurement, cybersecurity, and finance and introduces EAL-Bench, which measures how accurately persistent memory preserves evolving authorization state and whether errors propagate to downstream unauthorized actions.

DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation
DramaChain Bench:面向短剧生成的端到端基准
arXiv:2609.00646 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 DramaChain Bench,首个覆盖完整生产链各阶段的短剧基准,并验证最终剧集质量并非仅由视频生成决定。DramaChain Bench is presented, the first short-drama benchmark that evaluates every stage of the complete production chain, and confirms that final episode quality is not governed by video generation alone.

Does Imitation Learning Preserve Temporal Robustness in Dexterous Manipulation? An Expert-Learner Comparison Across Task Execution Speeds
模仿学习是否保留灵巧操作的时间鲁棒性?跨任务执行速度的专家-学习者对比
arXiv:2609.01453 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

任务名义成功率相同并不意味着跨执行速度保留专家性能;本文在相同任务条件、初始条件采样与加速倍率下对比了专家与学习者。Equal nominal task success does not imply preservation of expert performance across execution speeds, and an expert and learner under the same task conditions, initial-condition draws, and speedup factors is compared.

6. LLM Research Papers: The 2026 List (Jan–May) — Sebastian Raschka
6. LLM Research Papers:2026 清单(1月–5月) — Sebastian Raschka
arXiv:2604.12374 LLM 基础设施 方法 OA · 绿色 被引 37 · S2

Nemotron 3 Super 是 Nemotron 3 系列中首个采用 NVFP4 进行预训练的模型,借助 LatentMoE(一种同时优化精度 per FLOP 与精度 per parameter 的新型 Mixture-of-Experts 架构),并集成 MTP 层以通过原生 speculative decoding 加速推理。Nemotron 3 Super is the first model in the Nemotron 3 family to be pre-trained in NVFP4, leverage LatentMoE, a new Mixture-of-Experts architecture that optimizes for both accuracy per FLOP and accuracy per parameter, and include MTP layers for inference acceleration through native speculative decoding.

Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
中期训练阶段的知识蒸馏更偏向推理而非事实记忆
arXiv:2609.01532 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Switch Distillation,一种简单的 mid-training 目标:以教师预测熵作为轻量路由信号,仅在教师置信的 token 上蒸馏,其余回退到交叉熵;在不同教师规模下均稳定优于现有蒸馏目标。Switch Distillation is proposed, a simple mid-training objective that distills on tokens where the teacher is confident, using teacher predictive entropy as a lightweight routing signal, and otherwise falls back to cross-entropy, which consistently outperforms existing distillation objectives across teacher sizes.

Adapting Without Gradients: Affine Statistics Transport and What Its Certificate Can Tell You
无梯度自适应:仿射统计迁移及其证书能告诉你什么
arXiv:2609.00374 LLM 基础设施 应用落地 被引 0 · S2

本文提出 CASTER,一种无梯度方法:将源类统计量存储在判别子空间中,从目标批次矩估计一个类共享的仿射变换,并在分类前解析地将源类分布迁移到目标域,使其成为面向冻结模型部署的轻量适配机制。CASTER is introduced, a gradient-free method that stores source class statistics in a discriminative subspace, estimates a class-shared affine transformation from target-batch moments, and analytically transports the source class distributions before classification, which positions CASTER as a lightweight adaptation mechanism for frozen-model deployment.

Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data
通过对非结构化数据的自适应结构化实现 token 高效的数据推理 Agent
arXiv:2608.31082 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 agentic data cracking,将非结构化数据自适应、推测性地结构化为推理自身的副产品,是面向非结构化数据上 Agentic reasoning 的下一代数据基础设施的第一步。This work proposes agentic data cracking, a method that structures unstructured data adaptively and speculatively as a byproduct of reasoning itself, a first step toward next-generation data infrastructure for agentic reasoning over unstructured data.

Learning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal Geometry
学习结果变化的位置:面向多模态几何的可寻址信用推理
arXiv:2608.30457 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 credit-addressable reasoning:推理时暴露的语义单元同时定义学习阶段比较候选与分配 credit 的位置;并实例化为 Code-CoT,保留图示、将视觉关系表示为行可寻址的可执行代码,并将推理组织为类型化事件。This work introduces credit-addressable reasoning, in which the semantic units exposed during inference also define where learning compares alternatives and assigns credit, and instantiates Code-CoT, which retains the diagram, represents visual relations as line-addressable executable code, and organizes reasoning into typed events.

AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling
AgentJudgeBench:一个用于评估 LLM 法官在智能体工具调用方面表现的多难度基准
arXiv:2608.26623 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

首个系统性研究 LLM-as-a-judge 在工作流 DAG 上对 Agentic 工具调用评判可靠性的基准,区别于面向开放式文本或偏好的通用 LLM-as-a-judge 任务;揭示了当前 LLM judge 的根本局限,并给出面向 Agentic 系统可靠评估的实践指南。The first benchmark to systematically study LLM-as-a-judge reliability for agentic tool-calling over workflow DAGs, as distinct from the broader LLM-as-a-judge task of open-ended text or preference evaluation, exposes fundamental limitations of current LLM judges and yields practical guidelines for reliable evaluation in agentic systems.

VibeVoice-ASR-Streaming Technical Report
VibeVoice-ASR-Streaming 技术报告
arXiv:2609.02812 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

VibeVoice-ASR-Streaming 是首批基于 LLM 的端到端流式说话人归属 ASR 方法之一,无需独立 diarization 阶段即可在语音到达时输出"谁说了什么"。VibeVoice-ASR-Streaming is one of the first LLM-based end-to-end approaches to streaming speaker-attributed ASR, allowing the model to produce''who said what''as speech arrives, without a separate diarization stage.

HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
HarnessDev:LLM 能创建并演进自身的 Agent Harness 吗?
arXiv:2609.01437 评测基准 评测集 OA · 绿色 被引 5 · S2

研究发现,生成的 harness 在代码与搜索/研究任务上仍显著落后于成熟的人工参考方案,而在写作与机器学习实验任务上达到或超过所选参考方案,且执行成本差异巨大。It is found that generated harnesses remain substantially behind mature human-engineered references on code and on search and research, while matching or exceeding the selected references on writing and machine-learning experimentation, with large variation in execution cost.

CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing
CRISP:面向结构-质量驱动的悬崖感知输入自适应稀疏预填充
arXiv:2609.01925 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文用 C_struct 替代 Jensen-Shannon Divergence 路由:C_struct 是一种结构化代理,通过度量 Vertical-Slash 兼容位置上的 mass 来复现 JSD 的路由决策,同时消除池化 matmul 与后续 KL 散度开销。This work replaces the Jensen-Shannon Divergence routing with C_struct, a structural proxy that measures mass at Vertical-Slash compatible positions and reproduces JSD's routing decisions while eliminating both the pooled matmul and subsequent KL divergence overhead.

A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss
一瞥即可:基于 SimLoss 的单遍细粒度图像描述生成
arXiv:2609.00591 RAG 检索增强 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 SimLoss,一种面向单轮细粒度图像描述的无参考 embedding 空间目标;结果表明 embedding 空间监督能在单轮描述器延迟下恢复多阶段验证的质量。SimLoss is proposed, a reference-free embedding-space objective for single-pass fine-grained image captioning, and results show that embedding-space supervision can recover the quality of multi-stage verification at the latency of a single-pass captioner.

Aspire: Can Models Self-Evolve from Vague Goals?
Aspire:模型能否从模糊目标中自我演进?
arXiv:2608.31111 评测基准 评测集 OA · 绿色 被引 1 · S2

本文提出 ASPIRE,面向模糊目标驱动自我演化的基准,表明模糊目标会将搜索资源导向目标解释阶段;并在涵盖 6 类目标的 520 题隐藏专家评测集上评估所得系统。This work introduces ASPIRE, a benchmark for vague-goal-driven self-evolution and shows that vague goals redirect search effort toward goal interpretation, and evaluates the resulting systems on a hidden, expert-authored set of 520 items spanning six goals.

5️⃣ arXiv · Is Agentic RAG Worth It? An Experimental Comparison of RAG Approaches(⭐⭐⭐⭐ 高优先级)
5️⃣ arXiv · Agentic RAG 是否值得?RAG 方法的实验对比(⭐⭐⭐⭐ 高优先级)
arXiv:2601.07711 RAG 检索增强 评测集 OA · 绿色 被引 6 · S2

基于实证对 "Enhanced" 与 "Agentic" RAG 范式进行评估,为真实场景中选取最有效的 RAG 设计(兼顾性能与成本)提供指导。An empirically driven evaluation of the "Enhanced" and "Agentic" RAG paradigms is conducted, offering guidance on selecting the most effective RAG design for real-world applications, considering both performance and costs.

S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?
S3Gym:LLM 能将自测试与自评判转化为自我提升吗?
arXiv:2608.31100 评测基准 评测集 OA · 绿色 被引 1 · S2

这些发现表明,仅识别成功动作并不足够;Agent 还需将反馈转化为可执行且可迁移的策略。本文给出统一框架以诊断该过程,并定位阻碍 Agent 将交互经验转化为可靠自我提升的瓶颈。These findings show that recognizing successful actions is insufficient; agents must also transform feedback into executable and transferable policies, and provide a unified framework for diagnosing this process and identifying the bottlenecks that prevent agents from translating interaction experience into reliable self-improvement.

SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions
SnapBench:面向移动端交互的即拍即问多模态检索 benchmark
arXiv:2608.29607 RAG 检索增强 评测集 OA · 绿色 被引 1 · S2

本文提出 SnapBench——首个面向鲁棒"拍照即问"多模态检索的配对基准,以及一种简单的自适应融合方法 MOOR(Modality-anchored, Outlier-aware, Optimal Reweighting),并指出在"拍照即问"检索中需要具备可靠性感知的模态校准。SnapBench is introduced, the first paired benchmark for robust snap-and-ask multimodal retrieval, and MOOR (Modality-anchored, Outlier-aware, Optimal Reweighting), a simple adaptive fusion approach, highlighting the need for reliability-aware modality calibration in snap-and-ask retrieval.

Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers
机构报纸流水线:从历史报纸中提炼数十亿高质量 token
arXiv:2608.18972 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Institutional Newspapers Pipeline,一个模块化系统,旨在从历史报纸扫描件中提取高质量、结构化的数据集;其架构设计使每个步骤都保持可解释和可定制,并使整个 pipeline 在计算上足够精简,可在工作站级硬件上运行。The Institutional Newspapers Pipeline is presented, a modular system designed to extract high-quality, structured datasets from historical newspaper scans that was architected so that each step remains interpretable and customizable, and so that the pipeline as a whole remains computationally frugal enough to run on workstation-level hardware.

ViSAR: Training-Free Adaptive-$k$ Retrieval for Visual Document Question Answering
ViSAR:面向视觉文档问答的无训练自适应 $k$ 检索
arXiv:2609.02486 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 ViSAR(Visual Semantic Activation Retrieval),一种面向 late-interaction 视觉文档检索的无训练自适应 k 检索方法,并表明相似度矩阵结构与答案准确率相关,为面向检索质量感知的文档理解指明了未来方向。ViSAR (Visual Semantic Activation Retrieval), a training-free adaptive-$k$ retrieval method for late-interaction visual document retrieval, is introduced and it is shown that the similarity matrix structure correlates with answer accuracy, suggesting future directions for retrieval quality-aware document understanding.