本文提出一种面向历史数字图书馆管理的文档分析系统,支持即时知识建模,并促进生成更丰富、更全面的信息。This article presents a document analysis system designed for the management of historical digital libraries that supports on-the-fly knowledge modeling and facilitates the generation of richer and more comprehensive information.
论文
471 张论文卡片 · 方法 · OA 绿色
PolyUQuest 是一个基于异构图构建的可验证、结构感知 Web RAG 框架,统一了页面间超链接拓扑、页面内 DOM 层级以及跨页面实体-关系知识,在答案正确性、覆盖度和忠实度上优于现有 RAG 系统,且每次查询消耗的 LLM tokens 显著更少。PolyUQuest, a verifiable, structure-aware web RAG framework built on a heterogeneous graph that unifies hyperlink topology between pages, DOM hierarchy within pages, and entity-relation knowledge across pages, outperforms existing RAG systems in answer correctness, coverage, and faithfulness, while consuming significantly fewer LLM tokens per query.
CineMobile 采用三重优化策略,通过蒸馏引导的剪枝方法得到一个紧凑而高效的模型,保留实现电影级效果所需的核心视频生成能力,证明了其在移动端图像到视频创作中的实用可行性。CineMobile adopts a three-fold optimization strategy, leveraging a distillation-guided pruning approach to derive a compact yet efficient model that retains the essential video generation capabilities required for cinematic effects, demonstrating its practical applicability for mobile-based image-to-video creation.
本文提出 LongE2V,一种利用预训练视频扩散先验来联合处理基于事件视频重建、预测与帧间插值的新方法,并引入自回归展开与自适应上下文切换机制,以缓解超长序列中的时序漂移问题。This work proposes LongE2V, a novel approach that leverages pre-trained video diffusion priors to jointly handle event-based video reconstruction, prediction, and frame interpolation, and introduces Autoregressive Unrolling and Adaptive Context Switching to mitigate temporal drift in extremely long sequences.
借助强大的全景先验,Canvas360 构建了一个统一的上下文全景生成框架,通过 token 级拼接支持多样化下游任务,在任务覆盖范围与建模灵活性上均超越已有方法。Empowered by strong panoramic priors, Canvas360 enables a unified in-context panoramic generation framework that supports diverse downstream tasks via token-level concatenation, surpassing prior methods in both task coverage and modeling flexibility.
本文对 softmax 注意力与四种近期的循环线性注意力架构(DeltaNet、Gated DeltaNet、Kimi Delta Attention 与 Gated DeltaNet-2)进行对比研究,明确阐述它们在表达能力、记忆衰减、擦写控制、训练吞吐量与实现复杂度上的差异。A comparative study of softmax attention and four recent recurrent linear-attention architectures: DeltaNet, Gated DeltaNet, Kimi Delta Attention, and Gated DeltaNet-2 is presented, making explicit how they differ in expressivity, memory decay, erase and write control, training throughput, and implementation complexity.
消融实验表明,选择性干预优于被动记忆库暴露、常驻注入、仅顾问引导和通用检索。Ablations show that selective intervention outperforms passive bank exposure, always-on injection, advisor-only guidance, advisor-only guidance, and general retrieval, and general retrieval and that selective intervention outperforms passive bank exposure, always-on injection, advisor-only guidance, and general retrieval.
SAM-MT 成功将延迟与目标数量解耦,在保持 SAM2 鲁棒视频分割性能的同时,实现了与单目标基线相当的实时速度。SAM-MT successfully decouples latency from the number of targets, achieving real-time speed on par with single-target baselines while maintaining SAM2's robust video segmentation performance.
提出 PAST-TIDE,在官方排行榜上 Subtask A 取得 0.75 的 macro-F1,Subtask B 取得 0.74,表明对预训练模型仅做极少的架构改动即可在低资源场景下保持竞争力。PAST-TIDE is introduced, PAST-TIDE achieves macro-F1 scores of 0.75 for Subtask A and 0.74 for Subtask B on the official leaderboard, indicating that minimal architectural additions to a pre-trained model can remain competitive in low-resource settings.
通过将疾病特异性上下文整合到分子生成中,DrugGen-2 推动了 AI 辅助药物发现,为 de novo 设计和药物再利用提供了强大工具,可同时考虑疾病与分子靶点之间的复杂相互作用。By integrating disease-specific context into molecular generation, DrugGen-2 advances AI-assisted drug discovery, offering a powerful tool for de novo design and drug repurposing that accounts for the complex interplay between diseases and molecular targets.
本工作描述了如何在开源实现中满足使用有限项数的截断态向量来模拟 peaked circuits 的要求,并讨论了其性能与局限性。This work describes how the requirements to simulate peaked circuits using a truncated state vector with a limited number of terms were met in an open-source implementation, and discusses its performance and limitations.
一项受控消融实验揭示了机制:从检索到的文档中移除特定实体的临床证据,可彻底消除实体归因失败,使所有失败转移到虚构生成。A controlled ablation identifies the mechanism: removing entity-specific clinical evidence from retrieved documents eliminates entity-attribution failure entirely, shifting all failures to confabulation.
论文探讨了 LLM 为公司基本面分析各方面带来的机会,分析依据包括公司报告、描述宏观经济状况(如 GDP 和通胀变化)的数据与文件,以及提交至美国证券交易委员会(SEC)的文件。The opportunities brought by Large Language Models (LLMs) to various aspects of fundamental analysis of companies based on their reports as well as data and documents describing macroeconomic situation like GDP and inflation changes as well as documents filled to the U.S. Securities and Exchange Commission (SEC) are examined.
论文提出 PanoWorld,通过固定朝向将相机轨迹简化为平移,并借助 Dense Panoramic Ray-Conditioning 与 Geometry-aware Memory Augmentation 同时支持当前动作建模与长程记忆。PanoWorld is proposed, which simplifies camera trajectories into translations via fixed headings for both current-action modeling and long-range memory through Dense Panoramic Ray-Conditioning and Geometry-aware Memory Augmentation through Dense Panoramic Ray-Conditioning and Geometry-aware Memory Augmentation.
一种简单方法 Self-Guided TTT (S-TTT),可同时提升 Qwen3-4B-Thinking-2507 与 Llama-3.1-8B-Instruct 的准确率,相对改进最高达 15%。A simple method, Self-Guided TTT (S-TTT), which improves accuracy for both Qwen3-4B-Thinking-2507 and Llama-3.1-8B-Instruct, achieving up to a 15% relative improvement.
论文主张语音结构已隐含在自监督语音模型(S3M)的表征中,只需对其进行引导即可同时完成切分与识别任务。It is argued that phonetic structure is already latent in the representations of self-supervised speech models (S3Ms), and one only needs to steer them to solve both segmentation and recognition tasks.
提出 MedPMC——一种自动化、可持续更新的框架,可将宽松许可的文献转化为面向医学多模态模型的高保真基础设施,并公开发布该框架、语料库、基准与预训练模型。MedPMC, an automated, continuously updatable framework that transforms permissively licensed literature into high-fidelity infrastructure for medical multimodal models, is introduced and publicly release the framework, corpus, benchmarks, and pretrained models.
提出 VaseMuseum——一种面向古希腊陶器智能数字博物馆的轻量化、模块化多模态智能体框架,相比启用搜索的 VLM 基线,它提升了引用有效性,减少了知识密集型查询中的幻觉,并在含歧义场景下给出更中立的回答。VaseMuseum is proposed, a lightweight and modular multimodal agent framework for intelligent digital museums of ancient Greek pottery that improves citation validity, reduces hallucinations on knowledge-intensive queries, and produces more neutral answers under ambiguity compared with search-enabled VLM baselines.
提出 ABot-AgentOS,一个通用机器人 Agent Operating System,位于底层控制器之上,提供 deliberation agent 层,支持场景条件规划、上下文隔离的 Skill 执行、多阶段验证、多模态记忆以及边云协同。ABot-AgentOS is presented, a general robotic Agent Operating System that sits above low-level controllers and provides a deliberative agent layer for scene-conditioned planning, context-isolated skill execution, multi-stage verification, multi-modal memory, and edge-cloud collaboration.
提出 StudioRecon,一种通过解耦背景与人体、并利用视频扩散模型合成数百个相机可控新视角,从稀疏低重叠相机重建 4D 人体场景的流水线,在四个真实数据集上达到了 SOTA 的新视角合成效果StudioRecon is proposed, a pipeline that reconstructs 4D human scenes from sparse, low-overlap cameras by decoupling background and humans by synthesizing hundreds of camera-controlled novel views with a video diffusion model and achieves state-of-the-art novel view synthesis across four real-world datasets.
推出 CtrlVTON,一个将试穿重构为图像编辑问题并引入分割掩码作为对服装布局(包括风格、尺寸与身体空间位置)像素级控制的可控 VTO 框架CtrlVTON is introduced, a controllable VTO framework that recasts try-on as an image editing problem and adds segmentation masks as pixel-level control over garment layout, including style, size, and spatial placement on the body.
提出 Direct On-Policy Distillation(Direct-OPD),该方法迁移教师模型由 RL 引起的策略偏移,而非在目标模型上运行稀疏奖励 RL,并一致地利用更弱的教师模型来提升更强的目标模型Direct On-Policy Distillation (Direct-OPD) is proposed, which transfers the teacher's RL-induced policy shift instead of running sparse-reward RL on the target model and consistently leverages weaker teachers to improve stronger target models.
提出 Proxy OPD——一种异步后训练框架,迁移奖励驱动的策略改进而非绝对策略分布,将相对策略更新确立为可大规模、按奖励进行后训练的高复用、可调节资产。Proxy OPD is introduced, an asynchronous post-training framework that transfers reward-induced policy improvements rather than absolute policy distributions and establishes relative policy updates as highly reusable, adjustable assets for scalable, reward-based post-training.
提出 LATO.2,一个因子化 flow matching 框架,将网格生成分解为 vertex flow 和随后以已实现顶点为条件的 connectivity flow,在几何保真度和连通性质量上超越 SOTA 的拓扑感知网格生成方法。LATO.2, a factorized flow matching framework that decomposes mesh generation into a vertex flow followed by a connectivity flow conditioned on the realized vertices, is presented, which surpasses state-of-the-art topology-aware mesh generators in geometric fidelity and connectivity quality.
本文探索了一个预训练、冻结 encoder 的潜空间用于 text-to-image 个性化,并表明可以在该空间及由选定 token 定义的子空间中识别出有意义的编辑方向,从而实现局部化、细粒度且语义一致的编辑。This work explores the latent space of a pre-trained, frozen encoder for text-to-image personalization, and shows that meaningful directions can be identified within this space and within subspaces defined by selected tokens, enabling localized, fine-grained, and semantically coherent edits.
一个全栈系统,可从第一人称人类视频扩展灵巧 VLA 的预训练,并支持数据高效的真实机器人后训练,在 40+ 多种任务上稳健执行自由形式指令,展现出失败恢复、灵巧性与泛化能力。A full-stack system that scales dexterous VLA pre-training from egocentric human videos and enables data-efficient real-robot post-training that robustly executes free-form instructions across 40+ diverse tasks, demonstrating failure recovery, dexterity, and generalization.
本文提出 Multi-Agent Contextual Exploration (MACE),一个通过结构化的对等体选择显式促进探索的轻量级框架,显著改善了探索行为和下游任务表现,并在理论上证明探索价值随 Agent 多样性增加而提升。This work introduces Multi- Agent Contextual Exploration (MACE), a lightweight framework that explicitly promotes exploration through structured peer selection that substantially improves exploration behavior and downstream task performance and shows theoretically that the value of exploration increases with agent diversity.
提出 ST-Evidence,首个同时面向判别式和生成式像素级 grounding 的人工验证 benchmark,并开发可扩展的自动生成流程,构建了 16 万规模、衔接高层推理与细粒度 grounding 的数据集 ST-Evidence-Instruct。ST-Evidence is introduced, the first human-verified benchmark for both discriminative and generative pixel-level grounding, and scalable, automated generation pipelines are developed to create ST-Evidence-Instruct, a 160k-scale dataset bridging high-level reasoning with fine-grained grounding.
针对一系列基本增广与任意具有平稳统计量的图像数据集,以解析方式根据对比损失计算最优表示,结果表明对于某些增广,最优解可由第一层滤波器为正弦函数的 CNN 实现。Analytically computing the optimal representation in terms of a contrastive loss for a range of basic augmentations and any image dataset with stationary statistics shows that for certain augmentations the optimum can be attained by a CNN whose first layer filters are sinusoids.
提出 SpectraReward,一种无需训练的将预训练 MLLM 转化为即用型奖励模型的奖励函数,用于图像生成强化学习;并引入 Self-SpectraReward,这是统一多模态模型的一种特例,其中策略自身的理解分支充当其生成分支的奖励模型。SpectraReward is proposed, a training-free reward function that turns pretrained MLLMs into off-the-shelf reward models for image-generation reinforcement learning, and Self-SpectraReward is introduced, a special case for unified multimodal models where the policy's own understanding branch serves as the reward model for its generation branch.
基于 LLM 的编程 Agent 显著推动了自动化软件问题解决,但由于对仓库理解不足,仍易出现事实性错误。近期方法尝试通过修复前仓库探索来缓解此问题;然而,其修复驱动策略在未识别 Agent 知识缺口的情况下探索仓库,往往产生不精确的上下文,无法弥补潜在的理解不足。本文提出 ACQUIRE,一种面向软件问题解决的 QA 驱动框架,模拟经验丰富的开发者LLM-based coding agents have significantly advanced automated software issue resolution, yet they remain highly prone to factual errors caused by insufficient repository understanding. Recent methods attempt to mitigate this limitation through pre-repair repository exploration; however, their fix-driven strategies explore repositories without identifying the agent's knowledge gaps, often yielding imprecise context that fails to bridge the underlying understanding deficit. In this paper, we propose ACQUIRE, a QA-driven framework for software issue resolution. Mirroring how experienced developer
介绍 AMID,一种面向医学影像模型开发的自主多 Agent 框架,其性能优于所评估的通用 MLE 系统,并在异构任务上接近或匹配强大的人工设计挑战赛方案。AMID is introduced, an autonomous multi-agent framework for medical imaging model development that outperformed evaluated general-purpose MLE systems and approached or matched strong human-designed challenge solutions across heterogeneous tasks.
介绍其潜在原因的理论基础,阐明强化学习算法的渐近性能在性能排名与数据规模之间不存在单调关系。The theoretical foundations of the underlying causes outlining that the asymptotic performance of reinforcement learning algorithms does not have a monotone relationship between performance rankings and data-regimes are introduced.
除领域内增益外,mid-training 还能缓解 agentic post-training 对非 Agent 编程及非编程工具调用基准(tau-bench、BFCL)造成的能力侵蚀:尽管 mid-training 语料仅含 Python 代码,函数调用的归纳偏置在 post-training 后依然保留,带来稳定的增益。Beyond in-domain gains, mid-training mitigates the capability erosion that agentic post-training otherwise inflicts on non-agent coding and non-coding tool-use benchmarks (tau-bench, BFCL): although the mid-training corpus contains Python code only, the function-call inductive bias survives post-training and yields consistent gains.
提出 EvoGraph-R1,一个自演化 GraphRAG 框架,将知识图谱重新概念化为由 Agent 交互塑造的动态环境,将自演化知识图谱确立为跨模态的基础范式。EvoGraph-R1 is introduced, a self-evolving GraphRAG framework that reconceptualizes knowledge graphs as dynamic environments shaped through agent interactions, establishing self-evolving knowledge graphs as a fundamental paradigm across modalities.
尽管视觉-语言模型(VLMs)已取得成功,但误导性图表因欺骗性视觉结构与失真数据表示仍构成重大挑战。我们提出 ChartCynics,一个通过"怀疑式"推理范式揭露视觉欺骗的 Agentic 双路径框架。与整体化模型不同,ChartCynics 将感知与验证解耦:诊断式视觉路径通过策略性 ROI 裁剪捕获结构异常(如倒置坐标轴),OCR 驱动数据路径确保数值根植性。为解决跨模态冲突,我们提出Despite the success of Vision-Language Models (VLMs), misleading charts remain a significant challenge due to their deceptive visual structures and distorted data representations. We present ChartCynics, an agentic dual-path framework designed to unmask visual deception via a "skeptical" reasoning paradigm. Unlike holistic models, ChartCynics decouples perception from verification: a Diagnostic Vision Path captures structural anomalies (e.g., inverted axes) through strategic ROI cropping, while an OCR-Driven Data Path ensures numerical grounding. To resolve cross-modal conflicts, we introduce