研究库 论文知识库
Papers · organized/paper_cards

论文

1094 张论文卡片 · 方法 · OA 绿色

开放获取 全部 绿色 · 1640
RECAP-Forcing: Retaining Content Appearances for Long Video Generation
RECAP-Forcing:面向长视频生成的内容外观保持方法
arXiv:2608.26671 多模态 方法 OA · 绿色 被引 2 · S2

本文提出 RECAP-Forcing,一种无需训练的推理方法,不增加任何可学习参数,在多个强基线上稳定提升视觉质量与语义保真度,并优于现有记忆方法。This work proposes RECAP-Forcing, a training-free inference method with no additional learnable parameters that consistently improves visual quality and semantic fidelity across multiple strong baselines and outperforms existing memory methods.

SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers
SMELT:面向计算匹配的 MoE 循环 Transformer 的扩展定律
arXiv:2609.01343 RAG 检索增强 方法 OA · 绿色 被引 13 · S2

结果表明,即便在算力预算匹配的前提下,循环(looping)仍可提升 Transformer,提供了一种将深度复用转化为可衡量增益的实用方案。Results show that looping can improve Transformers even under budget matching, offering a practical recipe that turns depth reuse into measurable gains.

Safin-1: Safety from Within through Memory-Native State Evolution
Safin-1:通过记忆原生状态演化实现内在安全
arXiv:2609.00092 安全与风险 方法 OA · 绿色 被引 0 · S2 + OpenAlex

路由状态接口在模型的原生计算中统一了上下文记忆与持久的能力适配,将记忆从对历史上下文的被动记录重塑为维持与演化模型行为的主动基质。The routed-state interface unifies contextual memory and persistent capability adaptation within the model's native computation, reframing memory from a passive record of prior context into an active substrate for maintaining and evolving model behavior.

DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory
DiagEvo:通过分层错误记忆实现的诊断引导自我演化
arXiv:2609.00768 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

本文提出 DagEvo,通过利用 self-play 中求解器的失败历史来引导问题生成,无需外部任务资源,并表明混合生成、跨状态拼接的记忆状态更新以及双置信度过滤共同贡献了这些性能提升。DagEvo is introduced, which guides question generation using the solver's failure history from self-play, without external task resources, and shows that mixed generation, memory-state updates with cross-state stitching, and double-confidence filtering contribute to these gains.

6. LLM Research Papers: The 2026 List (Jan–May) — Sebastian Raschka
6. LLM Research Papers:2026 清单(1月–5月) — Sebastian Raschka
arXiv:2603.15569 LLM 基础设施 方法 OA · 绿色 被引 99 · S2

本工作借鉴线性模型的 state space model(SSM)视角,提出三项核心方法改进并组合形成更具表达力的递推结构:源自 SSM 离散化的递推式、用于更丰富状态追踪的复数值状态更新规则,以及在不增加 decode 延迟前提下提升模型性能的多输入多输出(MIMO)建模。This work introduces three core methodological improvements inspired by the state space model (SSM) viewpoint of linear models that combine to form a more expressive recurrence derived from SSM discretization, a complex-valued state update rule that enables richer state tracking, and a multi-input, multi-output (MIMO) formulation for better model performance without increasing decode latency.

Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement
Harness-of-Harness:具备持续改进能力的多日自主软件开发
arXiv:2609.01481 评测基准 方法 OA · 绿色 被引 2 · S2

在持续多天、超过 70 轮迭代的部署中,HoH 自主开发出一款第一人称射击游戏,具备完整的主线剧情、完整实现的核心机制、可供人类游玩的体验、精美的画面与集成的音效。In a multi-day deployment with more than 70 iterations, HoH autonomously develops a first-person-shooter game, featuring a coherent storyline, fully implemented core mechanics, human-playable experience, polished visuals and integrated audio.

ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training
ZimaBlue:通过可扩展视频预训练演化出可泛化的通用世界动作模型
arXiv:2609.00188 多模态 方法 OA · 绿色 被引 2 · S2

本文提出 ZimaBlue,一个可扩展的框架,用于从大规模视频中学习可泛化的 World Action Models (WAMs),并采用异步 Slow-Fast 双系统架构,使生成式 WAM 具备面向实时控制的实用性。This work introduces ZimaBlue, a scalable framework for learning generalizable World Action Models (WAMs) from large-scale video, and adopts an asynchronous Slow-Fast dual-system architecture to make generative WAMs practical for real-time control.

Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving
Qwen-Drive-1.0:迈向自动驾驶视觉-语言基础模型的初步探索
arXiv:2609.00111 多模态 方法 OA · 绿色 被引 7 · S2

实验表明,该方法在大幅保留通用视觉-语言能力的同时,具备出色的 3D 感知与驾驶场景理解能力;在开环、伪闭环与闭环设定下的综合评估进一步显示其运动规划性能具有很强的竞争力。Experiments demonstrate strong 3D perception and driving scene understanding while largely preserving general vision-language capability and comprehensive evaluations across open-loop, pseudo-closed-loop, and closed-loop settings further show highly competitive motion-planning performance.

Adaptive Critical Token-Aware Retrieval for Repository-Level Code Generation
自适应关键 token 感知的仓库级代码生成检索
arXiv:2609.01601 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

ACToR 在生成过程中识别关键 token,按需触发有针对性的检索,在这些决定性位置提供仓库上下文;并为稠密检索器设计了一种位置感知加权方法,以优先考虑对生成更具信息量的上下文。ACToR identifies critical tokens during generation and triggers targeted retrieval on demand to provide repository context at these decisive positions, and designs a position-aware weighting method for dense retrievers to prioritize context that is more informative for generation.

Agent Memory Is a Surface for Endogenous Authorization Laundering
Agent 内存是内生授权洗钱的表层载体
arXiv:2609.01836 Agent 智能体 方法 OA · 绿色 被引 3 · S2

本工作在采购、网络安全与金融场景下评估了 5 个 LLM 作为记忆写入者、2 个 LLM 作为执行者,并提出 EAL-Bench,用于衡量持久记忆在多大程度上准确保留不断演化的授权状态,以及错误是否会向下游传播为未授权操作。This work evaluates five LLMs as memory writers and two as executors across procurement, cybersecurity, and finance and introduces EAL-Bench, which measures how accurately persistent memory preserves evolving authorization state and whether errors propagate to downstream unauthorized actions.

Does Imitation Learning Preserve Temporal Robustness in Dexterous Manipulation? An Expert-Learner Comparison Across Task Execution Speeds
模仿学习是否保留灵巧操作的时间鲁棒性?跨任务执行速度的专家-学习者对比
arXiv:2609.01453 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

任务名义成功率相同并不意味着跨执行速度保留专家性能;本文在相同任务条件、初始条件采样与加速倍率下对比了专家与学习者。Equal nominal task success does not imply preservation of expert performance across execution speeds, and an expert and learner under the same task conditions, initial-condition draws, and speedup factors is compared.

6. LLM Research Papers: The 2026 List (Jan–May) — Sebastian Raschka
6. LLM Research Papers:2026 清单(1月–5月) — Sebastian Raschka
arXiv:2604.12374 LLM 基础设施 方法 OA · 绿色 被引 37 · S2

Nemotron 3 Super 是 Nemotron 3 系列中首个采用 NVFP4 进行预训练的模型,借助 LatentMoE(一种同时优化精度 per FLOP 与精度 per parameter 的新型 Mixture-of-Experts 架构),并集成 MTP 层以通过原生 speculative decoding 加速推理。Nemotron 3 Super is the first model in the Nemotron 3 family to be pre-trained in NVFP4, leverage LatentMoE, a new Mixture-of-Experts architecture that optimizes for both accuracy per FLOP and accuracy per parameter, and include MTP layers for inference acceleration through native speculative decoding.

Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
中期训练阶段的知识蒸馏更偏向推理而非事实记忆
arXiv:2609.01532 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Switch Distillation,一种简单的 mid-training 目标:以教师预测熵作为轻量路由信号,仅在教师置信的 token 上蒸馏,其余回退到交叉熵;在不同教师规模下均稳定优于现有蒸馏目标。Switch Distillation is proposed, a simple mid-training objective that distills on tokens where the teacher is confident, using teacher predictive entropy as a lightweight routing signal, and otherwise falls back to cross-entropy, which consistently outperforms existing distillation objectives across teacher sizes.

Learning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal Geometry
学习结果变化的位置:面向多模态几何的可寻址信用推理
arXiv:2608.30457 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 credit-addressable reasoning:推理时暴露的语义单元同时定义学习阶段比较候选与分配 credit 的位置;并实例化为 Code-CoT,保留图示、将视觉关系表示为行可寻址的可执行代码,并将推理组织为类型化事件。This work introduces credit-addressable reasoning, in which the semantic units exposed during inference also define where learning compares alternatives and assigns credit, and instantiates Code-CoT, which retains the diagram, represents visual relations as line-addressable executable code, and organizes reasoning into typed events.

VibeVoice-ASR-Streaming Technical Report
VibeVoice-ASR-Streaming 技术报告
arXiv:2609.02812 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

VibeVoice-ASR-Streaming 是首批基于 LLM 的端到端流式说话人归属 ASR 方法之一,无需独立 diarization 阶段即可在语音到达时输出"谁说了什么"。VibeVoice-ASR-Streaming is one of the first LLM-based end-to-end approaches to streaming speaker-attributed ASR, allowing the model to produce''who said what''as speech arrives, without a separate diarization stage.

CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing
CRISP:面向结构-质量驱动的悬崖感知输入自适应稀疏预填充
arXiv:2609.01925 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文用 C_struct 替代 Jensen-Shannon Divergence 路由:C_struct 是一种结构化代理,通过度量 Vertical-Slash 兼容位置上的 mass 来复现 JSD 的路由决策,同时消除池化 matmul 与后续 KL 散度开销。This work replaces the Jensen-Shannon Divergence routing with C_struct, a structural proxy that measures mass at Vertical-Slash compatible positions and reproduces JSD's routing decisions while eliminating both the pooled matmul and subsequent KL divergence overhead.

Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers
机构报纸流水线:从历史报纸中提炼数十亿高质量 token
arXiv:2608.18972 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Institutional Newspapers Pipeline,一个模块化系统,旨在从历史报纸扫描件中提取高质量、结构化的数据集;其架构设计使每个步骤都保持可解释和可定制,并使整个 pipeline 在计算上足够精简,可在工作站级硬件上运行。The Institutional Newspapers Pipeline is presented, a modular system designed to extract high-quality, structured datasets from historical newspaper scans that was architected so that each step remains interpretable and customizable, and so that the pipeline as a whole remains computationally frugal enough to run on workstation-level hardware.

ViSAR: Training-Free Adaptive-$k$ Retrieval for Visual Document Question Answering
ViSAR:面向视觉文档问答的无训练自适应 $k$ 检索
arXiv:2609.02486 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 ViSAR(Visual Semantic Activation Retrieval),一种面向 late-interaction 视觉文档检索的无训练自适应 k 检索方法,并表明相似度矩阵结构与答案准确率相关,为面向检索质量感知的文档理解指明了未来方向。ViSAR (Visual Semantic Activation Retrieval), a training-free adaptive-$k$ retrieval method for late-interaction visual document retrieval, is introduced and it is shown that the similarity matrix structure correlates with answer accuracy, suggesting future directions for retrieval quality-aware document understanding.

NE-R1: Enhancing Named Entity Recognition Model via Reinforcement Learning
NE-R1:通过强化学习增强命名实体识别模型
arXiv:2609.02366 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 NE-R1,一种面向自适应检索增强 NER 的新框架,在多个基准上达到 SOTA 性能,域内评估平均 F1 提升 2.52%,零样本跨域评估平均 F1 提升 1.18%。This paper proposes NE-R1, a novel framework for adaptive retrieval-augmented NER, which achieves state-of-the-art performance on various benchmarks, with an average F1 score gain of 2.52% in in-domain evaluation and 1.18% in zero-shot cross-domain evaluation.

NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference
NeoMME:用于高效微调与推理的单塔多模态原生多语言基础编码器
arXiv:2609.01657 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 NeoMME,一个 260M 与 800M 参数的多模态多语言双向编码器系列,可在单个双向 Transformer encoder 中处理多语言文本与原始图像 patch。This work introduces NeoMME, a family of 260M and 800M-parameter Multimodal and Multilingual bidirectional Encoders that process multilingual text and raw image patches in a single bidirectional Transformer encoder.

FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos
FoldingAgent:从演示视频中推断参数化折纸过程
arXiv:2609.00377 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 FoldingAgent,一个从折纸演示视频中推断显式参数化折纸程序的 Agent 框架,利用预训练 Vision-Language Model 的推理能力,并配备可模拟几何变换、验证物理合理性、检索与比较视觉内容以及评估自身预测的专用工具集。FoldingAgent is presented, an agentic framework for inferring explicit parametric folding programs directly from origami demonstration videos that leverages the reasoning power of a pre-trained Vision-Language Model equipped with a suite of specialized tools that enable the agent to simulate geometric transitions, verify physical plausibility, retrieve and compare visual content, and evaluate its own predictions.

5. GraphRAG / LLMs+Graphs 综合研究
arXiv:2606.11560 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本教程综合了推动这些汇聚方向的算法、系统与设计原则,为数据科学与数据挖掘研究者提供统一视角,涵盖将 LLM、图数据管理、图挖掘、图 ML 与 agentic 计算融合到下一代 graph-native AI 系统中。This tutorial synthesizes the algorithms, systems, and design principles driving these converging directions, offering data science and data mining researchers a unified perspective on integrating LLMs, graph data management, graph mining, graph ML, and agentic computation into next-generation graph-native AI systems.

Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering
超越视觉相似性:面向知识库视觉问答的实体对齐检索
arXiv:2608.21450 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 KBMR,首个面向 KB-VQA 的基于 MLLM 的 embedding retriever,并引入一个基于 MLLM 的语义判别器以生成连续的实体一致性权重,应对维基百科规模检索中的噪声监督挑战。KBMR is proposed, the first MLLM-based embedding retriever tailored for KB-VQA, and an MLLM-based semantic discriminator that generates continuous entity-consistency weights is introduced to tackle the challenge of noisy supervision in Wikipedia-scale retrieval.

Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens
Sparse Readout Prism:用特征而非 token 解释 Logit-Lens 分数
arXiv:2609.01936 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Sparse Readout Prism (SRP) 仅使用 readout 的权重对其进行分解,将任意 token logit 或 logit 差表示为来自稀疏 readout 特征贡献之和,揭示了 readout 特征作为 lens 解读新单元的价值,暴露出 token 身份可能掩盖的结构,并支持跨 token、上下文、层与 lens 的比较。Sarse Readout Prism (SRP), which decomposes the readout using only its weights and expresses any token logit or logit difference as a sum of contributions from sparse readout features, reveals readout features as a new unit of analysis for lens readings, exposing structure that token identities can obscure and enabling comparisons across tokens, contexts, layers, and lenses.

WHALE: A Simple Recipe for Joint Harness-Weight Optimization
WHALE:联合 Harness-权重优化的简洁方案
arXiv:2609.00196 评测基准 方法 OA · 绿色 被引 8 · S2

本文提出 Weight-Harness Alternating LEarning (WHALE),一种简单的方法,交替进行两个阶段:先在当前 harness 下更新模型,再在更新后的模型下通过在线拒绝采样微调与 Meta-Harness 搜索更优的 harness。Weight-Harness Alternating LEarning (WHALE), a simple recipe that alternates two phases: updating the model under the current harness, then searching for a better harness under the updated model with online rejection-sampling fine-tuning and Meta-Harness, is proposed.

Small Language Models as Judges for Rubric-Based Reinforcement Learning
小语言模型作为基于评分标准的强化学习评判器
arXiv:2608.30005 评测基准 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文研究较小的语言模型能否作为高效且可靠的 rubric 评分器,并比较了从小模型中提取逐项判断的三种方式:生成式判定、Yes/No Logprob 边际以及探针判别器。This work studies whether smaller language models can serve as efficient and reliable rubric-based judges, and compares three ways of extracting criterion-level judgments from small models: Generative verdicts, Yes/No Logprob margins, and Probe judges.

An Empirical Study on Zero-Data Bootstrapping for Conversational Recommender Systems
面向对话推荐系统的零数据引导的实证研究
arXiv:2504.15476 多模态 方法 OA · 绿色 被引 7 · S2

结果表明:领域驱动的合成数据一致优于零样本提示与朴素合成基线;主动选择相比随机采样提升了数据效率;元数据与协同过滤信号各自提升选择质量;在低资源场景下,合成数据可优于稀缺的真实对话,并进一步对真实对话形成补充。The results show that domain-grounded synthetic data consistently outperforms zero-shot prompting and naive synthetic baselines; active selection improves data efficiency over random sampling; metadata and collaborative filtering signals each improve selection quality; and, in low-resource settings, synthetic data can outperform scarce real dialogues while further complementing them.

Extending concurrent separation logic to the hardware level to verify the xv6 OS kernel on RISC-V with AI agents
将并发分离逻辑扩展到硬件层面,借助 AI Agent 在 RISC-V 上验证 xv6 OS 内核
arXiv:2609.04043 Agent 智能体 方法 OA · 绿色 被引 0 · S2 + OpenAlex

证明了一个应用层定理:若用户在 UART 控制台输入 echo hello world,系统唯一能产生的输出即为 hello world;这证明了基于 LLM 的 Agent 能够对如此底层的细节进行推理。An application-level theorem is proved: if the user types echo hello world as input on the UART console, the only output the system can produce is hello world, which proves LLM-based agents are capable of reasoning about such low-level details.

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States
Puffin-World:以原生 3D 世界状态扩展统一多模态模型
arXiv:2609.04196 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Puffin-World,一种统一的多模态架构,集成物理理解、空间仿真与 3D 世界生成重建,无需依赖外部离线模块,可支持需要多任务协同的交错式闭环应用。Puffin-World is proposed, a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation and reconstruction without relying on external offline modules and enables interleaved closed-loop applications requiring synergy across multiple tasks.

Compile by Training: Turning Natural-Language Specifications into Local Neural Functions
Compile by Training:将自然语言规约转化为本地神经函数
arXiv:2609.04199 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

许多重复出现的文本函数易于描述却难以用规则实现;而为每个输入调用大型远程模型会带来重复开销、延迟与对服务方的依赖。我们提出 compile by training,将自然语言规约转化为可复用的神经函数。在编译时,教师模型生成任务专属样本,用于为精简解释器训练一个小型适配器。生成的函数可在没有教师模型的情况下运行,并能像普通软件一样被存储、版本化管理与组合。在 FuzzyBench-Hard 这一子集上……Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remote model for every input introduces repeated cost, latency, and dependency on a provider. We present compile by training, which turns a natural-language specification into a reusable neural function. At compile time, teacher models generate task-specific examples that are used to train a small adapter for a compact interpreter. The resulting function runs without the teachers and can be stored, versioned, and composed like ordinary software. On FuzzyBench-Hard, a subset on which

Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding
超越检索:面向流式视频理解的渐进式潜在记忆演化
arXiv:2609.04131 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出 LatentStream,一种渐进式 latent working memory 框架,将流式记忆从"存储-检索"转变为"检索-内化",在现有在线和离线视频 benchmark 上取得新的 SOTA 结果。This work introduces LatentStream, a progressive latent working memory framework that shifts streaming memory from store-and-retrieve to retrieve-and-internalize, and achieves new state-of-the-art results on existing online and offline video benchmarks.

PACE: Towards Surfacing Hidden Conflicts in User Requests
PACE:揭示用户请求中的隐性冲突
arXiv:2609.03293 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 PACE 数据集,用于评估模型能否识别以自我中心知识或事件形式表达的潜在约束(这些约束使看似合理的用户请求变得不当),以及 PaceMaker 多 Agent 框架,其中专门 Agent 通过查询重构、多跳图遍历与冲突感知过滤进行协调,以检索上下文决定性证据。PACE is introduced, a dataset for evaluating whether models can identify latent constraints, expressed as egocentric knowledge or events, that render seemingly reasonable user requests inappropriate, and PaceMaker, a multi-agent framework in which specialized agents coordinate across query reformulation, multi-hop graph traversal, and conflict-aware filtering to retrieve contextually decisive evidence.

Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration
改变置信度,而非改变预测:面向后验校准的预测保持型修复
arXiv:2609.01072 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 CORD,第一个 post-fit adapter,通过从 calibrator 拟合中移除 preservation constraint 来修复完整校准概率向量,从而实现精确的预测保持,并将原始决策的精确恢复交由后续输出修复完成。CORD is proposed, the first post-fit adapter to impose exact prediction preservation by repairing the full calibrated probability vector by removing the preservation constraint from calibrator fitting and assigns exact recovery of the original decision to subsequent output repair.

Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training
知道何时不复用:自主 LLM 后训练中的条件经验迁移
arXiv:2608.26730 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

将条件经验迁移进行形式化,并提出 Boundary-Calibrated Intervention Transfer,一种在权重变化的训练之前即授权经验复用的方法,在相同预算下取得比所评估替代方案更高的最终模型质量。Conditional experience transfer is formulated as conditional experience transfer and Boundary-Calibrated Intervention Transfer is introduced, a method that authorizes experience reuse before weight-changing training and attains higher equal-budget final-model quality than the evaluated alternatives.

Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction
Scal3R:学习高效多相对位姿查询以实现可扩展的在线 3D 重建
arXiv:2609.04201 工程化 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该方法将在线重建重构为多参考相对位姿查询,在 Virtual KITTI、Sintel、TUM-Dynamic、ScanNet 和 7-Scenes 上取得 SOTA 性能,并在单 GPU 上 8 小时内收敛。This approach reformulates online reconstruction as multi-reference relative pose querying, which achieves state-of-the-art performance across Virtual KITTI, Sintel, TUM-Dynamic, ScanNet, and 7-Scenes and reaches convergence in 8 hours on a single GPU.

Percolation Dynamics in Optimization : Variance Cascades and Discrete Scale Invariance
优化中的渗流动力学:方差级联与离散尺度不变性
arXiv:2609.02373 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

通过将随机梯度流建模为渗流过程来研究随机梯度下降的动态,其中嵌套的架构对称性迫使子网络以离散块的形式合并,而非通过单边附着。The dynamics of Stochastic Gradient Descent is studied by modeling the stochastic gradient flow as a percolation process, in which nested architectural symmetries force subnetworks to merge in discrete blocks rather than by single-edge attachment.