提出 Contextuality(上下文性)——即 AI 系统自主访问用户累积知识资本的程度——作为 AI 介导不平等的一个维度,补充但不可化约为 Sharp 等人的框架。Contextuality -- the degree to which an AI system autonomously accesses a user's accumulated knowledge capital -- is proposed as a dimension of AI-mediated inequality that complements, but is not reducible to, the Sharp et al. framework.
论文
724 张论文卡片 · OA 绿色
本文提出一种面向历史数字图书馆管理的文档分析系统,支持即时知识建模,并促进生成更丰富、更全面的信息。This article presents a document analysis system designed for the management of historical digital libraries that supports on-the-fly knowledge modeling and facilitates the generation of richer and more comprehensive information.
PolyUQuest 是一个基于异构图构建的可验证、结构感知 Web RAG 框架,统一了页面间超链接拓扑、页面内 DOM 层级以及跨页面实体-关系知识,在答案正确性、覆盖度和忠实度上优于现有 RAG 系统,且每次查询消耗的 LLM tokens 显著更少。PolyUQuest, a verifiable, structure-aware web RAG framework built on a heterogeneous graph that unifies hyperlink topology between pages, DOM hierarchy within pages, and entity-relation knowledge across pages, outperforms existing RAG systems in answer correctness, coverage, and faithfulness, while consuming significantly fewer LLM tokens per query.
CineMobile 采用三重优化策略,通过蒸馏引导的剪枝方法得到一个紧凑而高效的模型,保留实现电影级效果所需的核心视频生成能力,证明了其在移动端图像到视频创作中的实用可行性。CineMobile adopts a three-fold optimization strategy, leveraging a distillation-guided pruning approach to derive a compact yet efficient model that retains the essential video generation capabilities required for cinematic effects, demonstrating its practical applicability for mobile-based image-to-video creation.
该审计揭示现有基准样本中 55% 可在无视觉输入或时序上下文的情况下被解决,并提出 Video-Oasis,一个用于系统性审计现有视频理解基准的可持续诊断套件。This audit reveals that 55\% of existing benchmark samples are solvable without visual input or temporal context, and introduces Video-Oasis, a sustainable diagnostic suite for systematically auditing existing video understanding benchmarks.
本文论证了稀疏组合监督与动宾学习的不对称性会助长物体驱动的捷径学习,并指出减少捷径诊断可提升组合泛化能力。This work argues that sparse compositional supervision and verb-object learning asymmetry can promote object-driven shortcut learning and reduces shortcut diagnostics and consequently improves compositional generalization.
本文提出 LongE2V,一种利用预训练视频扩散先验来联合处理基于事件视频重建、预测与帧间插值的新方法,并引入自回归展开与自适应上下文切换机制,以缓解超长序列中的时序漂移问题。This work proposes LongE2V, a novel approach that leverages pre-trained video diffusion priors to jointly handle event-based video reconstruction, prediction, and frame interpolation, and introduces Autoregressive Unrolling and Adaptive Context Switching to mitigate temporal drift in extremely long sequences.
借助强大的全景先验,Canvas360 构建了一个统一的上下文全景生成框架,通过 token 级拼接支持多样化下游任务,在任务覆盖范围与建模灵活性上均超越已有方法。Empowered by strong panoramic priors, Canvas360 enables a unified in-context panoramic generation framework that supports diverse downstream tasks via token-level concatenation, surpassing prior methods in both task coverage and modeling flexibility.
本文对 softmax 注意力与四种近期的循环线性注意力架构(DeltaNet、Gated DeltaNet、Kimi Delta Attention 与 Gated DeltaNet-2)进行对比研究,明确阐述它们在表达能力、记忆衰减、擦写控制、训练吞吐量与实现复杂度上的差异。A comparative study of softmax attention and four recent recurrent linear-attention architectures: DeltaNet, Gated DeltaNet, Kimi Delta Attention, and Gated DeltaNet-2 is presented, making explicit how they differ in expressivity, memory decay, erase and write control, training throughput, and implementation complexity.
该紧凑且量化原生、带内置控制的运行时为物联网音频场景下的端侧语义音频提供了实用基础;通过对转向接口的案例分析,可生成在部分属性上具有真实但有界控制的、承载口味联想的音乐。A compact, quantized runtime with built-in control a practical basis for on-device semantic audio in Internet-of-Sounds settings and a case study of the steering interface generates music carrying taste associations with genuine but bounded control for a subset of attributes.
消融实验表明,选择性干预优于被动记忆库暴露、常驻注入、仅顾问引导和通用检索。Ablations show that selective intervention outperforms passive bank exposure, always-on injection, advisor-only guidance, advisor-only guidance, and general retrieval, and general retrieval and that selective intervention outperforms passive bank exposure, always-on injection, advisor-only guidance, and general retrieval.
SAM-MT 成功将延迟与目标数量解耦,在保持 SAM2 鲁棒视频分割性能的同时,实现了与单目标基线相当的实时速度。SAM-MT successfully decouples latency from the number of targets, achieving real-time speed on par with single-target baselines while maintaining SAM2's robust video segmentation performance.
提出 PAST-TIDE,在官方排行榜上 Subtask A 取得 0.75 的 macro-F1,Subtask B 取得 0.74,表明对预训练模型仅做极少的架构改动即可在低资源场景下保持竞争力。PAST-TIDE is introduced, PAST-TIDE achieves macro-F1 scores of 0.75 for Subtask A and 0.74 for Subtask B on the official leaderboard, indicating that minimal architectural additions to a pre-trained model can remain competitive in low-resource settings.
通过将疾病特异性上下文整合到分子生成中,DrugGen-2 推动了 AI 辅助药物发现,为 de novo 设计和药物再利用提供了强大工具,可同时考虑疾病与分子靶点之间的复杂相互作用。By integrating disease-specific context into molecular generation, DrugGen-2 advances AI-assisted drug discovery, offering a powerful tool for de novo design and drug repurposing that accounts for the complex interplay between diseases and molecular targets.
本工作描述了如何在开源实现中满足使用有限项数的截断态向量来模拟 peaked circuits 的要求,并讨论了其性能与局限性。This work describes how the requirements to simulate peaked circuits using a truncated state vector with a limited number of terms were met in an open-source implementation, and discusses its performance and limitations.
一项受控消融实验揭示了机制:从检索到的文档中移除特定实体的临床证据,可彻底消除实体归因失败,使所有失败转移到虚构生成。A controlled ablation identifies the mechanism: removing entity-specific clinical evidence from retrieved documents eliminates entity-attribution failure entirely, shifting all failures to confabulation.
论文探讨了 LLM 为公司基本面分析各方面带来的机会,分析依据包括公司报告、描述宏观经济状况(如 GDP 和通胀变化)的数据与文件,以及提交至美国证券交易委员会(SEC)的文件。The opportunities brought by Large Language Models (LLMs) to various aspects of fundamental analysis of companies based on their reports as well as data and documents describing macroeconomic situation like GDP and inflation changes as well as documents filled to the U.S. Securities and Exchange Commission (SEC) are examined.
Soofi S 30B-A3B 是一个主权开源的 MoE 混合 Mamba Transformer 德英基础模型,在作者的对比中超越了所有欧洲主权基线,包括活跃参数量远超自身的模型。Soofi S 30B-A3B is a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English that outperforms every European sovereign baseline in the authors' comparison, including ones far larger in active parameters.
该工作提出 LongMedBench,一个基于真实 EHR 的长程临床决策基准,并设计了一套包含三个评估维度的分类体系:事实型问答、时序推理、长程决策。This work introduces LongMedBench, a real-world EHR-based benchmark for long-horizon clinical decision-making, and proposes an evaluation taxonomy with three suites: fact-based QA, temporal reasoning, and long-horizon decision-making.
该工作提出 Long-Horizon-Terminal-Bench,一个涵盖九大类共 46 个长程任务的终端基准,包括实验复现、软件工程、多模态分析、交互游戏与科学计算,并分析失败模式与错误规律,以推动长程终端 Agent 的后续研究。This work introduces Long-Horizon-Terminal-Bench, a terminal benchmark of 46 long-horizon tasks spanning nine categories, including experiment reproduction, software engineering, multimodal analysis, interactive games, and scientific computing, and analyzes failure modes and error patterns to support future progress on long-horizon terminal agents.
论文提出 PanoWorld,通过固定朝向将相机轨迹简化为平移,并借助 Dense Panoramic Ray-Conditioning 与 Geometry-aware Memory Augmentation 同时支持当前动作建模与长程记忆。PanoWorld is proposed, which simplifies camera trajectories into translations via fixed headings for both current-action modeling and long-range memory through Dense Panoramic Ray-Conditioning and Geometry-aware Memory Augmentation through Dense Panoramic Ray-Conditioning and Geometry-aware Memory Augmentation.
一种简单方法 Self-Guided TTT (S-TTT),可同时提升 Qwen3-4B-Thinking-2507 与 Llama-3.1-8B-Instruct 的准确率,相对改进最高达 15%。A simple method, Self-Guided TTT (S-TTT), which improves accuracy for both Qwen3-4B-Thinking-2507 and Llama-3.1-8B-Instruct, achieving up to a 15% relative improvement.
论文主张语音结构已隐含在自监督语音模型(S3M)的表征中,只需对其进行引导即可同时完成切分与识别任务。It is argued that phonetic structure is already latent in the representations of self-supervised speech models (S3Ms), and one only needs to steer them to solve both segmentation and recognition tasks.
论文提出 Flow-ERD,一个同时追求真实性与多样性的多 Agent 仿真器,在 WOSAC 测试基准上排名第一,并在可复现基线中主导真实性-多样性 Pareto 前沿。Flow-ERD is introduced, a multi-agent simulator that pursues realism and diversity jointly and ranks first on the WOSAC test benchmark and dominates the realism--diversity Pareto front among reproducible baselines.
提出 AgentKGV,一种用于知识图谱事实核查的智能体 LLM-RAG 框架,集成动态路由与迭代查询改写,以应对文档级检索中的表层形式不匹配问题。AgentKGV, the Agentic LLM-RAG framework for KG fact Verification, is proposed, that integrates dynamic routing and iterative query rewriting, which handles surface-form mismatch in document-level retrieval.
提出 MedPMC——一种自动化、可持续更新的框架,可将宽松许可的文献转化为面向医学多模态模型的高保真基础设施,并公开发布该框架、语料库、基准与预训练模型。MedPMC, an automated, continuously updatable framework that transforms permissively licensed literature into high-fidelity infrastructure for medical multimodal models, is introduced and publicly release the framework, corpus, benchmarks, and pretrained models.
提出 VaseMuseum——一种面向古希腊陶器智能数字博物馆的轻量化、模块化多模态智能体框架,相比启用搜索的 VLM 基线,它提升了引用有效性,减少了知识密集型查询中的幻觉,并在含歧义场景下给出更中立的回答。VaseMuseum is proposed, a lightweight and modular multimodal agent framework for intelligent digital museums of ancient Greek pottery that improves citation validity, reduces hallucinations on knowledge-intensive queries, and produces more neutral answers under ambiguity compared with search-enabled VLM baselines.
本文提出 AdvancedMathBench,一个用于评估高级数学推理能力的 benchmark 套件,并推出 VerifierBench,包含 888 条模型生成的证明轨迹及专家 ground truth,用于评估模型能否正确判断证明有效性并给出合理的验证理由。This work introduces AdvancedMathBench, a benchmark suite designed to evaluate advanced mathematical reasoning capabilities, and introduces VerifierBench, consisting of 888 model-generated proof trajectories paired with expert ground truth, to evaluate whether models can correctly judge proof validity and provide sound verification rationales.
本文首次全面综述了 LLM 元认知的研究现状,涵盖用于测量和评估 LLM 元认知能力的方法与 benchmark、激发、改进与应用 LLM 元认知的技术,以及当前研究的发现与启示。The first comprehensive overview of the current state of knowledge on metacognition for LLMs is presented, including methods and benchmarks to measure and evaluate LLMs'metacognitive abilities, techniques to elicit, improve, and apply metacognition in LLMs, and findings and implications of ongoing research.
Motion4Motion 对视频中角色的 motion flow(而非骨骼)进行建模,使跨物种运动迁移更加容易。Motion4Motionmodels the motion flow of the character in a video instead of skeletons, which makes motion transfer across species easier, which makes motion transfer across species easier.
通过考察包含意识形态话语的 RAG 框架对 LLM 生成答案的影响,发现 RAG 框架倾向于将意识形态话语传递到 LLM 响应中,且采样温度对这种传递的强度有可测量的影响。Examining the influence of the RAG framework, comprising ideological discourses, in LLM-generated answers shows that the RAG framework is prone to transferring ideological discourses into LLM responses, with sampling temperature having a measurable impact on the strength of this transfer.
提出 ABot-AgentOS,一个通用机器人 Agent Operating System,位于底层控制器之上,提供 deliberation agent 层,支持场景条件规划、上下文隔离的 Skill 执行、多阶段验证、多模态记忆以及边云协同。ABot-AgentOS is presented, a general robotic Agent Operating System that sits above low-level controllers and provides a deliberative agent layer for scene-conditioned planning, context-isolated skill execution, multi-stage verification, multi-modal memory, and edge-cloud collaboration.
提出 StudioRecon,一种通过解耦背景与人体、并利用视频扩散模型合成数百个相机可控新视角,从稀疏低重叠相机重建 4D 人体场景的流水线,在四个真实数据集上达到了 SOTA 的新视角合成效果StudioRecon is proposed, a pipeline that reconstructs 4D human scenes from sparse, low-overlap cameras by decoupling background and humans by synthesizing hundreds of camera-controlled novel views with a video diffusion model and achieves state-of-the-art novel view synthesis across four real-world datasets.
推出 CtrlVTON,一个将试穿重构为图像编辑问题并引入分割掩码作为对服装布局(包括风格、尺寸与身体空间位置)像素级控制的可控 VTO 框架CtrlVTON is introduced, a controllable VTO framework that recasts try-on as an image editing problem and adds segmentation masks as pixel-level control over garment layout, including style, size, and spatial placement on the body.
提出 Direct On-Policy Distillation(Direct-OPD),该方法迁移教师模型由 RL 引起的策略偏移,而非在目标模型上运行稀疏奖励 RL,并一致地利用更弱的教师模型来提升更强的目标模型Direct On-Policy Distillation (Direct-OPD) is proposed, which transfers the teacher's RL-induced policy shift instead of running sparse-reward RL on the target model and consistently leverages weaker teachers to improve stronger target models.
提出 Proxy OPD——一种异步后训练框架,迁移奖励驱动的策略改进而非绝对策略分布,将相对策略更新确立为可大规模、按奖励进行后训练的高复用、可调节资产。Proxy OPD is introduced, an asynchronous post-training framework that transfers reward-induced policy improvements rather than absolute policy distributions and establishes relative policy updates as highly reusable, adjustable assets for scalable, reward-based post-training.