研究库 论文知识库
Papers · organized/paper_cards

论文

1640 张论文卡片 · OA 绿色

开放获取 全部 绿色 · 1640
Sliding-window beats linear attention
滑动窗口优于线性注意力
arXiv:2608.28444 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作表明带 sink 的 Sliding Window Attention(SWA)表现不逊于甚至优于后训练的 Linear Attention 模型,并建议改用 SWA 而非后训练线性模型。This work shows that Sliding Window Attention (SWA) with sinks performs as well or better than post-trained Linear Attention models, and recommends switching to SWA instead of post-training linear models.

EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses
EvoUndo:面向 LLM Agent 运行环境的可恢复性约束自演化
arXiv:2608.28363 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,可靠的 Agent 自我进化需要协同设计验证、状态接地、见证语义和恢复语言表达能力,而非仅依赖迭代提示。The results indicate that reliable agent self-evolution requires co-designing verification, state grounding, witness semantics, and recovery-language expressivity rather than relying on iterative prompting alone.

Acquire, Repair, Preserve: A Diagnosis-Guided Post-Training Recipe for Small-Model Dialogue Game Agents
获取、修复、保留:面向小模型对话游戏 Agent 的诊断驱动后训练方案
arXiv:2608.28458 工程化 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,广泛的 SFT 带来模型大部分能力提升;当失败检测精准时,turn-local 监督可发挥作用,且观察到的迁移主要集中在同族模型之间。The results suggest that broad SFT brings most of the model's capability improvement; turn-local supervision can be effective when failure detection is precise, with observed transfer concentrated primarily within-family.

SymbolLKG: Towards Verifiable Logical Reasoning via Logical Knowledge Graph and Symbolic Solvers
SymbolLKG:通过逻辑知识图谱与符号求解器实现可验证的逻辑推理
arXiv:2608.26836 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

一种神经符号架构,将逻辑知识图谱(LKG)与动态求解器路由相结合,并引入基于本体的 LKG,将逻辑规则和约束视为一等拓扑节点,从而支持对从文本中抽取的依赖关系进行显式建模。A Neuro-Symbolic architecture that integrates a Logical Knowledge Graph (LKG) with dynamic solver routing, and introduces an ontology-based LKG that treats logical rules and constraints as first-class topological nodes, enabling explicit modeling of dependencies extracted from text.

Pointing the Way, Hiding the Destination: Practical Private Dense Retrieval at Scale
指路而隐目的地:面向大规模场景的实用化私有稠密检索
arXiv:2608.25735 RAG 检索增强 方法 OA · 绿色 被引 1 · S2

该候选名单可在不牺牲检索质量的前提下短路全语料库加密搜索:200–500 个候选即可在 5 个零样本语料库(规模从 25K 到 5.4M 文档)中与全语料库检索效果接近匹配。This shortlist short-circuits full-corpus cryptographic search without sacrificing retrieval quality: with 200-500 candidates, it closely matches full-corpus retrieval across five zero-shot corpora spanning 25K to 5.4M documents.

MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents
MNIST-PRO:MNIST 作为部分可观察世界回归,用于 AI Agents
arXiv:2608.31022 Agent 智能体 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 MnIST-PRO 基准,将 Agent 感知能力隔离测试:把 MNIST 数字识别转换为带回看约束的序列式 glimpse 搜索任务,并评估了十个多模态模型,结果表明仅获取视觉证据是不够的,Agent 还必须能够构建并更新可靠的感知状态。MnIST-PRO is addressed, a benchmark that isolates agentic perception by converting MNIST digit recognition into a sequential, glimpse-based search task with lookback constraints, and ten multimodal models are evaluated, showing that simply acquiring visual evidence is not enough and agents must also be able to build and update a reliable perceptual state.

Agents in the Large: Perception-Centered Architecture for Persistent Agents
Large Agents:以感知为中心的持久化 Agent 架构
arXiv:2608.30478 Agent 智能体 方法 OA · 绿色 被引 1 · S2

Pera 描述了一种持久化 Agent,围绕感知和控制组件组织,持续从情景任务执行、上下文及周围环境变化中感知服务相关信号,并利用这些信号构建生命周期任务。Pera describes a persistent agent organized around perception and control components that continually perceive service-relevant signals from episodic task executions, internal context, and changes in the surrounding environment, and use these signals to construct lifecycle tasks.

Configurable Semantic Chunking for Biomedical Information Extraction in Retrieval-Augmented Generation
RAG 中面向生物医学信息抽取的可配置语义分块
arXiv:2608.31139 RAG 检索增强 观点 OA · 绿色 被引 0 · S2 + OpenAlex

跨数据集分析表明,语义切分在具有显式关系线索的抽取数据集(如 GM-CIHT 和 DDI)上表现更优,而固定切分在密集生化抽取和二分类场景(如 ChemProt 和 ADE)下仍具竞争力甚至更强。Cross-dataset analysis shows that semantic chunking improves extraction datasets with explicit relation cues, such as GM-CIHT and DDI, while fixed chunking remains competitive or stronger for dense biochemical extraction and binary classification settings such as ChemProt and ADE.

6. LLM Research Papers: The 2026 List (Jan–May) — Sebastian Raschka
LLM 研究论文:2026 年清单(1—5 月)— Sebastian Raschka
arXiv:2602.15763 Agent 智能体 方法 OA · 绿色 被引 514 · S2

GLM-5 在真实编码任务中展现出前所未有的能力,在端到端软件工程挑战的处理上超越既有基线,并提出了新颖的异步 Agent RL 算法,进一步提升了 RL 质量。GLM-5 demonstrates unprecedented capability in real-world coding tasks, surpassing previous baselines in handling end-to-end software engineering challenges and proposing novel asynchronous agent RL algorithms that further improve RL quality.

OntoAligner-Ensemble: Voting-Based Fusion across Heterogeneous Ontology Alignment Techniques
OntoAligner-Ensemble:基于投票的异构本体对齐技术融合
arXiv:2608.31137 安全与风险 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

结果揭示了集成组合对精度–召回权衡的直接影响:异构跨范式集成通常提升精度,而同构 LLM 集成更常取得更高的整体 F1。It is revealed that ensemble composition directly affects the precision-recall trade-off: heterogeneous cross-paradigm ensembles generally improve precision, whereas homogeneous LLM ensembles more often achieve higher overall F1-scores.

DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution
DreamX-Creator:在 2K 分辨率下实现原生音视频生成的民主化
arXiv:2608.31106 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

通过发布紧凑的 7B 生成器和 2K Refiner,该工作致力于使原生音视频生成平民化,并为统一音视频生成建模的未来研究提供可及的基础。By releasing the compact 7B generator and 2K Refiner, this work seeks to democratize native audio-video generation and provide an accessible foundation for future research in unified audio-video generative modeling.

PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback
PaperBanana-Interact:基于多轮人类反馈的科学图表精修
arXiv:2608.30241 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 MTPaperBananaBench,一个面向多轮图表生成的基准,包含 292 张图像和 3,518 条用户需求标注,并引入 PaperBanana-Interact,一个通过内部 critique-and-refine 循环来优化图表的多智能体系统。MTPaperBananaBench, a benchmark for multi-turn diagram generation containing 292 images annotated with 3,518 user requirements, is presented and PaperBanana-Interact, a multi-agent system that refines diagrams via an internal critique-and-refine loop is introduced.

Verification-Aware Training for Speculative Decoding
面向 Speculative Decoding 的验证感知训练
arXiv:2608.30135 工程化 观点 OA · 绿色 被引 2 · S2

提出 Verification-Aware Training,一种插件式框架,在每一步训练中模拟验证并将产生的 accept 与 reject 模式转化为监督信号,在数学、代码和聊天基准上提升了平均接受长度和实际加速比。Verification-Aware Training is introduced, a plug-in framework that simulates verification at every training step and turns the resulting accept and reject patterns into supervision and improves average acceptance length and wall-clock speedup across math, code, and chat benchmarks.

Weaving Visual Narratives: Agentic Image Bundle Composition Beyond Atomic Visual Matching
编织视觉叙事:超越原子视觉匹配的 Agentic Image Bundle Composition
arXiv:2608.28695 Agent 智能体 观点 OA · 绿色 被引 1 · S2

大量实验表明,SOTA embedding 模型和静态 decompose-and-rerank 范式存在关系盲区,而 BundleWeaver 取得了显著性能提升,凸显了从原子打分转向动态关系组合的必要性。Extensive experiments demonstrate that while state-of-the-art embedding models and static decompose-and-rerank paradigms suffer from relational blindness, BundleWeaver achieves substantial performance gains, highlighting the necessity of shifting from atomic scoring to dynamic relational composition.

Super Library Agent: Joint Generation and Maintenance of Multiple Applications Beyond the Single Codebase
Super Library Agent:跨单一代码库的多个应用联合生成与维护
arXiv:2608.29310 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出 Super Library Agent 问题:Agent 顺序生成 N 个相关应用组成的组合,同时维护一个共享的 Super Library 用于跨应用可复用组件,并解决了基于候选引导的代码块摘要抽取、抽取前的代码库整合,以及利用抽取 trace 和调用图信息的上下文感知迁移。This work introduces the Super Library Agent problem, where an agent sequentially generates a portfolio of N related applications while maintaining a shared Super Library of reusable cross-application components, and addresses candidate-guided extraction over code chunk summaries, pre-extraction codebase consolidation, and context-aware migration using extraction traces and call-graph information.

Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered
思维链忠实性随偏好线索的传递位置与方式而变化
arXiv:2608.29464 评测基准 方法 OA · 绿色 被引 1 · S2

结果表明,在所测试的单调用、预填工具场景下,当偏好信息通过工具传入或需从原始产物中推断时,CoT 监控的可靠性可能下降。The results suggest that CoT monitoring may be less reliable when preference information arrives through tools or must be inferred from raw artifacts, within the single-call, prefilled-tool setting tested here.

SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models
SafeAtlas-VL:超越二元判断的大规模多模态安全数据与防护模型
arXiv:2608.29098 安全与风险 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该论文提出 SafeAtlas-VL,一个包含 1.5M 训练实例的数据集,将图像、请求和响应级判断置于五级有序量表上,并通过 target-conditioned tuning 训练 SafeAtlas Guard 系列模型,用于多模态安全检测。This paper introduces SafeAtlas-VL, a dataset of 1.5M training instances that places image-, request-, and response-level judgments on a five-level ordered scale, and trains the SafeAtlas Guard series of models via target-conditioned tuning for multimodal safety detection.

Chat-Edit-3D++: Interactive 3D and 4D Scene Editing via Large Language Models
Chat-Edit-3D++:基于大语言模型的交互式 3D 与 4D 场景编辑
arXiv:2608.29137 多模态 方法 OA · 绿色 被引 1 · S2

提出 Hash-Atlas 网络,将 3D 场景编辑重新表述为对 2D atlas 图像的操作,从而实现 2D 编辑与 3D 重建流程的解耦。The Hash-Atlas network is proposed, which reformulates 3D scene editing as operations on 2D atlas images, thereby achieving a workflow decoupling of the 2D editing and 3D reconstruction processes.

Evaluating the Hidden Costs of Personalization in Large Language Models
评估大语言模型个性化中的隐性代价
arXiv:2608.28833 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 PRISK,一个具备自动化数据生成和定制化指标的动态评估框架,用于揭示当前 LLM 个性化中的系统性局限以及个性化信息如何塑造其响应。PRISK is proposed, a dynamic evaluation framework with automated data generation and tailored metrics that uncovers systematic limitations in current LLM personalization and how personalized information shapes its responses.

6. LLM Research Papers: The 2026 List (Jan–May) — Sebastian Raschka
LLM 研究论文:2026 年清单(1—5 月)— Sebastian Raschka
arXiv:2603.15031 LLM 基础设施 综述 OA · 绿色 被引 67 · S2
Unsupervised Learning of Video Representations using LSTMs
基于 LSTM 的视频表示无监督学习
arXiv:1502.04681 多模态 方法 OA · 绿色 被引 2770 · S2

本文利用 LSTM 网络学习视频序列表示,并通过在 UCF-101 与 HMDB-51 数据集上对人体动作识别这一监督学习任务进行微调来评估所学表示。This work uses Long Short Term Memory networks to learn representations of video sequences and evaluates the representations by finetuning them for a supervised learning problem - human action recognition on the UCF-101 and HMDB-51 datasets.

Uncertainty-Aware End-to-End AI Weather Forecasting: Disentangling Observation and Model Contributions
感知不确定性的端到端 AI 天气预报:分离观测与模型的贡献
arXiv:2608.30795 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

端到端天气预报系统直接从原始地球观测数据生成高质量的全球网格与站点预报,仅以数值天气预报流水线(含数据同化)一小部分的成本取而代之。这类系统是确定性的,不输出不确定性结果。本文通过在每个组件上附加一种随机机制,将 Aardvark Weather 模型概率化:在观测编码器中加入学习到的、依赖输入的噪声,以捕捉源自观测系统的偶然不确定性;在处理器中使用 Monte Carlo dropout,以捕捉认知不确定性。End-to-end weather forecasting systems produce skillful global gridded and station forecasts directly from raw Earth observations, replacing the numerical weather prediction pipeline, including data assimilation, at a fraction of its cost. These systems are deterministic and issue no uncertainty. Here we render the Aardvark Weather model probabilistic by attaching one stochastic mechanism to each component: learned, input-dependent noise at the observation encoder, capturing aleatoric uncertainty inherited from the observing system, and Monte Carlo dropout in the processor, capturing epistemic

ContextBias: Controlled Evaluation of Bias Persistence Under Context Shift in Text-to-Image Models
ContextBias:受控评估文本到图像模型在上下文偏移下的偏见持续性
arXiv:2608.29847 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

评估四个 SOTA 模型发现,将角色置于语义无关的上下文中并不会抑制该角色的关联属性;相反,跨角色的属性集中度会上升(合并 BI $+0.047$)。Evaluating four state-of-the-art models finds that placing a role in a semantically unrelated context does not suppress role-linked attributes; instead, cross-role attribute concentration increases (pooled BI $+0.047$).

SpanCalib-VLM: Calibrated Hallucination Span Detection in Vision-Language Models
SpanCalib-VLM:视觉语言模型中校准的幻觉片段检测
arXiv:2608.29974 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 SpanCalib-VLM,这是一种用于 SHROOM-Visions 共享任务的混合双系统,结合了多模态序列标注器(由 XLM-RoBERTa-Large 与 SigLIP 视觉编码器通过交叉注意力融合而成)与微调后的生成式 VLM(Qwen3.5-4B-SHROOM-SFT)。SpanCalib-VLM is presented, a hybrid dual-system for the SHROOM-Visions Shared Task that combines a multimodal sequence tagger, consisting of XLM-RoBERTa-Large fused with a SigLIP vision encoder via cross-attention, with the fine-tuned generative VLM (Qwen3.5-4B-SHROOM-SFT).

MMMMM: A Unified Taxonomy for Investigating the Mechanisms of Multilingual MultiModal Misinformation
MMMMM:用于研究多语言多模态虚假信息机制的统一分类法
arXiv:2608.29681 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

基于对数据的深入定性分析和既有理论研究,构建了一套新颖且全面的多模态错误信息分类法,并由此获得了关于社交媒体用户如何在实际场景中将图像与文本结合以传播错误信息的此前未被记录的洞见。A novel, comprehensive taxonomy of multimodal misinformation grounded in an in-depth qualitative analysis of the data and prior theoretical work is developed, which leads to previously undocumented insights about how social media users combine images with text to spread misinformation in the wild.

EvoGenUI-Bench: Evaluating LLMs as Multi-Turn Generative UI Assistants
EvoGenUI-Bench:评估 LLM 作为多轮生成式 UI 助手
arXiv:2608.29387 评测基准 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 EvoGenUI-Bench,一个面向多轮界面维护的基准,包含 150 个五轮任务,共计 750 轮,覆盖三种场景:信息呈现、可执行交互和工具驱动的外部状态。EvoGenUI-Bench is introduced, a benchmark for multi-turn interface maintenance comprising 150 five-turn tasks and 750 turns across three scenarios: information presentation, executable interaction, and tool-grounded external state.

RECAP-Forcing: Retaining Content Appearances for Long Video Generation
RECAP-Forcing:面向长视频生成的内容外观保持方法
arXiv:2608.26671 多模态 方法 OA · 绿色 被引 2 · S2

本文提出 RECAP-Forcing,一种无需训练的推理方法,不增加任何可学习参数,在多个强基线上稳定提升视觉质量与语义保真度,并优于现有记忆方法。This work proposes RECAP-Forcing, a training-free inference method with no additional learnable parameters that consistently improves visual quality and semantic fidelity across multiple strong baselines and outperforms existing memory methods.

SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers
SMELT:面向计算匹配的 MoE 循环 Transformer 的扩展定律
arXiv:2609.01343 RAG 检索增强 方法 OA · 绿色 被引 13 · S2

结果表明,即便在算力预算匹配的前提下,循环(looping)仍可提升 Transformer,提供了一种将深度复用转化为可衡量增益的实用方案。Results show that looping can improve Transformers even under budget matching, offering a practical recipe that turns depth reuse into measurable gains.

Safin-1: Safety from Within through Memory-Native State Evolution
Safin-1:通过记忆原生状态演化实现内在安全
arXiv:2609.00092 安全与风险 方法 OA · 绿色 被引 0 · S2 + OpenAlex

路由状态接口在模型的原生计算中统一了上下文记忆与持久的能力适配,将记忆从对历史上下文的被动记录重塑为维持与演化模型行为的主动基质。The routed-state interface unifies contextual memory and persistent capability adaptation within the model's native computation, reframing memory from a passive record of prior context into an active substrate for maintaining and evolving model behavior.

DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory
DiagEvo:通过分层错误记忆实现的诊断引导自我演化
arXiv:2609.00768 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

本文提出 DagEvo,通过利用 self-play 中求解器的失败历史来引导问题生成,无需外部任务资源,并表明混合生成、跨状态拼接的记忆状态更新以及双置信度过滤共同贡献了这些性能提升。DagEvo is introduced, which guides question generation using the solver's failure history from self-play, without external task resources, and shows that mixed generation, memory-state updates with cross-state stitching, and double-confidence filtering contribute to these gains.

6. LLM Research Papers: The 2026 List (Jan–May) — Sebastian Raschka
6. LLM Research Papers:2026 清单(1月–5月) — Sebastian Raschka
arXiv:2603.15569 LLM 基础设施 方法 OA · 绿色 被引 99 · S2

本工作借鉴线性模型的 state space model(SSM)视角,提出三项核心方法改进并组合形成更具表达力的递推结构:源自 SSM 离散化的递推式、用于更丰富状态追踪的复数值状态更新规则,以及在不增加 decode 延迟前提下提升模型性能的多输入多输出(MIMO)建模。This work introduces three core methodological improvements inspired by the state space model (SSM) viewpoint of linear models that combine to form a more expressive recurrence derived from SSM discretization, a complex-valued state update rule that enables richer state tracking, and a multi-input, multi-output (MIMO) formulation for better model performance without increasing decode latency.

Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement
Harness-of-Harness:具备持续改进能力的多日自主软件开发
arXiv:2609.01481 评测基准 方法 OA · 绿色 被引 2 · S2

在持续多天、超过 70 轮迭代的部署中,HoH 自主开发出一款第一人称射击游戏,具备完整的主线剧情、完整实现的核心机制、可供人类游玩的体验、精美的画面与集成的音效。In a multi-day deployment with more than 70 iterations, HoH autonomously develops a first-person-shooter game, featuring a coherent storyline, fully implemented core mechanics, human-playable experience, polished visuals and integrated audio.

ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training
ZimaBlue:通过可扩展视频预训练演化出可泛化的通用世界动作模型
arXiv:2609.00188 多模态 方法 OA · 绿色 被引 2 · S2

本文提出 ZimaBlue,一个可扩展的框架,用于从大规模视频中学习可泛化的 World Action Models (WAMs),并采用异步 Slow-Fast 双系统架构,使生成式 WAM 具备面向实时控制的实用性。This work introduces ZimaBlue, a scalable framework for learning generalizable World Action Models (WAMs) from large-scale video, and adopts an asynchronous Slow-Fast dual-system architecture to make generative WAMs practical for real-time control.

Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving
Qwen-Drive-1.0:迈向自动驾驶视觉-语言基础模型的初步探索
arXiv:2609.00111 多模态 方法 OA · 绿色 被引 7 · S2

实验表明,该方法在大幅保留通用视觉-语言能力的同时,具备出色的 3D 感知与驾驶场景理解能力;在开环、伪闭环与闭环设定下的综合评估进一步显示其运动规划性能具有很强的竞争力。Experiments demonstrate strong 3D perception and driving scene understanding while largely preserving general vision-language capability and comprehensive evaluations across open-loop, pseudo-closed-loop, and closed-loop settings further show highly competitive motion-planning performance.

UI-Venus-2 Technical Report
UI-Venus-2 技术报告
arXiv:2609.00028 Agent 智能体 评测集 OA · 绿色 被引 1 · S2

本文提出 UI-Venus-2,一个面向移动、Web 和桌面环境的通用 GUI 基础 Agent,采用统一的闭环推理-行动框架,并通过集成安全感知机制来确保关键操作的可控执行。UI-Venus-2 is presented, a general-purpose foundation GUI agent designed to operate across mobile, web, and desktop environments through a unified closed-loop reasoning-action framework that integrates safety-aware mechanisms to ensure controlled execution of consequential actions.

Adaptive Critical Token-Aware Retrieval for Repository-Level Code Generation
自适应关键 token 感知的仓库级代码生成检索
arXiv:2609.01601 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

ACToR 在生成过程中识别关键 token,按需触发有针对性的检索,在这些决定性位置提供仓库上下文;并为稠密检索器设计了一种位置感知加权方法,以优先考虑对生成更具信息量的上下文。ACToR identifies critical tokens during generation and triggers targeted retrieval on demand to provide repository context at these decisive positions, and designs a position-aware weighting method for dense retrievers to prioritize context that is more informative for generation.