研究库 论文知识库
Papers · organized/paper_cards

论文

1753 张论文卡片

开放获取 全部 绿色 · 1640
Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds
面向持久化故事与交互式世界的长时音视频生成
arXiv:2608.23383 多模态 方法 OA · 绿色 被引 2 · S2

结果表明,记忆、几何控制以及 rollout-aware 训练为生成连贯故事和持续演化的交互式世界提供了实用基础。Results indicate that memory, geometric control, and rollout-aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds.

Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data
Next-chunk reasoning RL 真的优于 SFT 吗?——在 no-CoT 数据下重新审视训练策略
arXiv:2608.23256 工程化 方法 OA · 绿色 被引 1 · S2

Mixed SFT 是一种单阶段监督微调,联合在 no-CoT 和 long-CoT 数据上训练,相比 next-chunk reasoning RL 取得了明显更高的 RLVR 后性能上限,同时训练算力开销减少超过 60 倍。Mixed SFT, a single supervised fine-tuning stage that jointly trains on no-CoT and long-CoT data, achieves a clearly higher post-RLVR performance ceiling than next-chunk reasoning RL while requiring over 60 times less training compute.

FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling
FIRM-Video:先核查再评分,实现可靠的 text-to-video 奖励建模
arXiv:2608.21839 多模态 方法 OA · 绿色 被引 1 · S2

本文提出 FHM-Video,一种基于 check-before-score 原则的、由 checklist 驱动的统一数据构建框架,在 FIRM-Video-Bench 上取得最佳综合 MAE,并在三种视频生成器的 Best-of-8 采样中始终获得最高的 VBench Total、Quality 和 Semantic Score。FHM-Video, a unified checklist-driven data construction framework based on a check-before-score principle, is introduced, which achieves the best overall MAE on FIRM-Video-Bench while consistently delivering the highest VBench Total, Quality, and Semantic Scores in Best-of-8 sampling across three video generators.

Real-TurnTurk: A Multimodal Turkish Corpus for Turn-Taking Prediction
Real-TurnTurk:用于话轮预测的多模态土耳其语语料库
arXiv:2608.22071 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文发布一个土耳其语多模态对话数据集,包含未脚本化的双人交互,并提供同步的前向视频、可归属到每位说话人的独立音频通道以及时间对齐的转写文本。A multimodal Turkish conversational dataset of unscripted dyadic interactions, comprising synchronized front-facing video, per-speaker audio channels that allow overlapping speech to be attributed to individual speakers, and time-aligned transcriptions is introduced.

7. LLM 压缩:联合剪枝 + 混合精度 PTQ
arXiv:2606.07819 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本工作提出一种新颖的混合精度 PTQ 策略,直接最小化整个模型的全局误差传播,而非孤立地处理逐层误差,并开发了一种新颖的联合优化方法,在统一搜索空间中同时学习结构化剪枝决策与混合精度量化策略。This work proposes a novel mixed-precision PTQ strategy that directly minimizes global error propagation across the entire model, rather than isolating layer-wise errors, and develops a novel joint optimization approach that simultaneously learns structural pruning decisions and mixed-precision quantization policies within a unified search space.

Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models
Stream4D:面向流式自回归扩散视频模型的 4D 一致性
arXiv:2608.19556 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文用前馈 4D 重建奖励替代静态 critic,显式建模场景动态,使连贯运动获得高一致性奖励,并加入对自然 scene-flow 幅值进行奖励同时抑制抖动与非刚性伪影的运动先验。This work replaces the static critic with a feed-forward 4D reconstruction reward that explicitly models scene dynamics, allowing coherent motion to receive high consistency rewards, and adds a motion prior that rewards natural scene-flow magnitude while penalizing jitter and non-rigid artifacts.

PlanSightRAG: A Visual-First Multimodal RAG for Automating Question Answering and Compliance Checking for Civil Standard Plans
PlanSightRAG:面向土木标准图自动化问答与合规审查的视觉优先多模态 RAG
arXiv:2608.26091 RAG 检索增强 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

PlanSightRAG 是一种 Visual-First 多模态 RAG 框架,直接对图纸图像建立索引并进行推理,集成了 ColNomic-3B 多向量检索、Agentic Planner-Retriever-Auditor-Synthesizer,并以 MaxSim 热力图作为证据链。A Visual-First Multimodal Retrieval-Augmented Generation (RAG) framework called PlanSightRAG, which indexes and reasons directly over plan imagery, integrates a ColNomic-3B multi-vector retrieval, an agentic Planner-Retriever-Auditor-Synthesizer, and MaxSim heatmaps as an evidence trail.

Multi-Granularity Context-Enhanced RAG over Multimodal Knowledge Graphs
多模态知识图谱上的多粒度上下文增强 RAG
arXiv:2608.25986 RAG 检索增强 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一种构建 Context-Enhanced MMKG (CEMMKG) 的新框架,能够有效利用上下文信息提升基于 MMKG 的 RAG 性能,并在不同基于 MMKG 的 RAG 方法上的有效性验证了其广泛适用性。A novel framework for constructing a Context-Enhanced MMKG (CEMMKG) is proposed, effective in leveraging contextual information to improve MMKG-based RAG performance and its effectiveness across different MMKG-based RAG methods demonstrates its broad applicability.

RetrievalRouter: Joint Modality and Architecture Selection for Document Retrieval
RetrievalRouter:面向文档检索的模态与架构联合选择
arXiv:2608.25625 RAG 检索增强 方法 被引 0 · S2

RetrievalRouter 是一种轻量级 query-aware router,仅依据 query 文本即可学习最匹配的检索 pipeline,在面向准确率的设置下 nDCG@5 显著更高,而在面向延迟的设置下,nDCG@5 和延迟均匹配或数值上优于基线。RetrievalRouter is a lightweight query-aware router that learns, from the query text alone, which retrieval pipeline best fits each query, and achieves significantly higher nDCG@5 across accuracy-oriented settings, while matching or numerically outperforming them on both nDCG@5 and latency in latency-oriented settings.

LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale
LibriBrain100:面向大规模神经语音解码的百小时广深 MEG 数据集
arXiv:2608.25204 多模态 方法 OA · 绿色 被引 4 · S2

本文发布 LibriBrain100,一个面向语音解码的大规模 MEG 数据集,从设计上保证可复现、标准化评估,并展示了广泛多被试数据的价值:对预训练模型进行有监督微调可大幅弥补单被试数据不足。LibriBrain100 is introduced, a large-scale MEG dataset for speech decoding designed from the ground up for reproducible, standardised evaluation and the value of broad multi-subject data is demonstrated: supervised finetuning of a pre-trained model can substantially compensate for limited per-subject data.

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?
SWE Refactor Bench:编码 Agent 能否完成长时序全仓库技术栈迁移?
arXiv:2608.23564 Agent 智能体 评测集 OA · 绿色 被引 3 · S2

本文发布 SWE Refactor Bench,一个包含 20 项全仓库迁移的基准,涵盖 4 类技术债务,作为开发面向可靠全仓库迁移的编码 Agent 的严格测试平台。SWE Refactor Bench is introduced, a benchmark comprising 20 whole-repository migrations, covering 4 kinds of technical debt, and SWE Refactor Bench is positioned as a rigorous testbed for developing coding agents for reliable whole-repository migrations.

Skill Issue: Are Skills Language-Invariant in LLMs?
Skill Issue:LLM 中的 Skill 是否具备语言无关性?
arXiv:2608.25832 评测基准 评测集 OA · 绿色 被引 1 · S2

本文通过多语言 self-play,正交于知识与综合基准性能对跨语言技能不一致性进行量化,表明技能差异是开发真正多语言模型过程中可衡量且主要的障碍。This work quantifies cross-lingual skill inconsistency orthogonally from knowledge and general benchmark performance via multilingual self-play, and shows that skill discrepancies are a measurable major roadblock in the development of truly multilingual models.

Prefix Sliding for efficient test-time scaling
Prefix Sliding:面向高效 test-time scaling 的前缀滑动方法
arXiv:2608.26070 LLM 基础设施 方法 OA · 绿色 被引 2 · S2

提出 Prefix Sliding,在推理过程中丢弃不属于前缀或最近几千 token 窗口的 token,从而实现高效的长时程 test-time scaling。Prefix Sliding is proposed, which discards tokens during reasoning that are not part of the prefix or the window of the last few thousand tokens, allowing for efficient long-horizon test-time scaling.

GUI-Primitives: Diagnosing Spatial Reasoning Failures in Vision-Language GUI Grounding
GUI-Primitives:诊断视觉语言 GUI grounding 中的空间推理失败
arXiv:2608.21832 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

计算机使用 Agent 将自然语言指令 grounding 到截图中以定位界面元素,但现有基准无法隔离模型是否将关系语言正确绑定到对应元素。我们提出 GUI-Primitives,一个包含 994 个条目的基准,由跨七种空间关系(左右、上下、包含、对齐、邻近、列表序数、遮挡)的对比指令对组成。每对保持截图和锚点不变,仅改变关系表达,使正确目标在两个指定候选之间切换。五位标注者……Computer-use agents ground natural-language instructions in screenshots to locate interface elements, yet existing benchmarks do not isolate whether models bind relational language to the correct element. We introduce GUI-Primitives, a 994-item benchmark of contrastive instruction pairs over seven spatial relations in graphical user interfaces (left/right, above/below, containment, alignment, proximity, list ordinal, occlusion). Each pair holds the screenshot and anchor fixed while changing the relation expression, so the correct target moves between two designated candidates. Five annotators

A Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans
一种面向 CT 扫描中空间关系验证的、可靠且可审计的模块化 Agent
arXiv:2608.21140 Agent 智能体 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一个模块化的医学影像 Agent,用于轴位 CT 切片中的二元空间关系验证,采用显式模块化空间验证阶段,并表明显式模块化空间验证可作为未来面向报告的医学影像 Agent 的有前景的构建模块。This work presents a modular medical imaging agent for binary spatial relation verification in axial CT slices using explicit modular spatial verification stages, and suggests that explicit modular spatial verification can serve as a promising building block for future report-oriented medical imaging agents.

6. QBugLM:量子软件调试多智能体框架
arXiv:2606.07314 Agent 智能体 方法 OA · 绿色 被引 2 · S2

本工作提出 QBugLM,一个多 Agent 框架,可自动化量子软件调试流水线,覆盖基于分类法的缺陷注入、基于 LLM 的检测与修复,直至基于仿真的验证,框架无关地支持 OpenQASM 3.0 程序。This work proposes QBugLM, a multi-agent framework that automates the quantum software debugging pipeline, from taxonomy-driven bug injection to LLM-based detection and repair, and finally to simulation-based validation, for framework-agnostic OpenQASM 3.0 programs.

TTPO: Test-Time Policy Optimization
TTPO:测试时策略优化。
arXiv:2608.27448 工程化 方法 OA · 绿色 被引 1 · S2

提出 Test-Time Policy Optimization,一种非对称目标,通过 OPSD 蒸馏一致性 rollout,并使用 Grouped RL 惩罚不一致的 rollout;进一步通过 token 级选择精炼两个分支:蒸馏降低已收敛位置的权重,而 RL 仅惩罚置信的错误。Test-Time Policy Optimization is proposed, an asymmetric objective that distills agreeing rollouts via OPSD and penalizes disagreeing rollouts with Grouped RL and Token-level selection further refines both branches: distillation down-weights already-converged positions, while RL penalizes only confident errors.

PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents
PILOT in the Loop:面向长时序 Agent 的实时自我改进。
arXiv:2608.26530 Agent 智能体 方法 OA · 绿色 被引 2 · S2

本文提出 PILOT,一种通过两种耦合机制实现实时自我改进的 supervisor-worker 框架:(1) live steering 允许独立的 supervisor 在执行期间重定向或中止当前 worker;(2) live self-evolution 将执行中发现的过程与失败模式提炼为可复用的 skills 与记忆。PILOT is presented, a supervisor-worker harness for live self-improvement through two coupled mechanisms: (1) live steering lets a separate supervisor redirect or abort the active worker during execution; and (2) live self-evolution distils procedures and failure modes revealed during execution into reusable skills and memory.

Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning
Thinking on Shots:基于 Agentic 推理的一致性多镜头视频编辑。
arXiv:2608.26809 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一个 Agentic 视频编辑框架,利用 LLM 与 VLM 的协同实现 shot 级视频解耦与精确指令解析,并构建 MMLVE-Bench,一个聚焦 MMLVE 的数据集,具有复杂的真实世界时空动态、高密度异构指令以及稀疏随机的实体分布。This work introduces an agentic editing framework that leverages the synergy of Large Language Models (LLMs) and Vision-Language Models (VLMs) to achieve shot-level video decoupling and precise instruction parsing, and constructs MMLVE-Bench, which is an MMLVE-focused dataset characterized by complex real-world spatiotemporal dynamics, high-density heterogeneous instructions, and sparse, random entity distributions.

Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning
Aphanta:诊断任务对齐的图像编辑中间态以服务多模态推理。
arXiv:2608.26993 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果将图像编辑定位为一种专门的视觉工作空间而非通用推理机制,并将 Aphanta 确立为可复用的协议,用于度量任务-表征对齐、编辑器实现及下游 pipeline 实用性。The results position image editing as a specialized visual workspace rather than a universal reasoning mechanism, and establish Aphanta as a reusable protocol for measuring task--representation alignment, editor realization, and downstream pipeline utility.

Procedura: Agentic 3D Modeling with Procedural Control
Procedura:基于过程化控制的 Agentic 三维建模
arXiv:2608.26238 Agent 智能体 观点 OA · 绿色 被引 1 · S2

本文探索"以代码表示 3D 形状"的范式,利用并放大 LLM 的编码能力进行 3D 建模,并提出 Procedura 框架——通过编写一个由命名零件构成、并通过类型化、可机器校验的连接关系装配而成的参数化程序来建模对象。The paradigm of 3D shape as code is explored, leveraging and scaling the coding ability of an LLM for 3D modeling, and Procedura, a novel 3D modeling agent framework that writes an object as a procedural assembly, a parametric program whose named parts are joined by typed, machine-checkable mates.

Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models
将 Agentic 游戏开发作为可扩展世界模型的可验证轨迹数据引擎
arXiv:2608.25518 Agent 智能体 方法 OA · 绿色 被引 1 · S2

本文提出 Reinforcement Learning with Human-Engine Verification(RLHEV),一种结合稠密引擎信号与开发过程中隐式人类接受反馈的后训练范式,用于支持强化学习后训练。This work proposes Reinforcement Learning with Human-Engine Verification (RLHEV), a post-training paradigm that combines dense engine signals with implicit human acceptance feedback from the development process to support RL post-training.

CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval
CaSKG:用于可扩展 Agent 技能检索的反事实-因果技能图
arXiv:2608.25500 RAG 检索增强 方法 OA · 绿色 被引 3 · S2

提出 CaSKG,一种反事实因果 Skill 图谱框架,在检索前校准程序关系,将边置信度校准定位为大规模紧凑且可执行的 Skill 检索的有效路径。CaSKG, a counterfactual-causal skill graph framework that calibrates procedural relations before retrieval, is proposed, position edge-confidence calibration as an effective route to compact and executable skill retrieval at scale.

GameWAM: A World Action Model for Video Games
GameWAM:面向视频游戏的世界动作模型
arXiv:2608.26200 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 GameWAM,据其所知是首个面向原生闭环游戏与 GUI 控制的 WAM,并发现 Low-Frequency Action Source Imprinting (LASI):在固定条件下,采样动作源的低频分量系统性地引导生成的粗粒度相机运动,揭示了生成式控制中的源敏感性失效模式。This work introduces GameWAM, to its knowledge the first WAM for native closed-loop gameplay and GUI control and uncovers Low-Frequency Action Source Imprinting (LASI), in which low-frequency components of the sampled action source systematically steer coarse generated camera motion under fixed conditioning, revealing a source-sensitivity failure mode in generative control.

Scaling Laws for Neural Language Models
神经语言模型的 Scaling Laws
arXiv:2001.08361 LLM 基础设施 方法 OA · 绿色 被引 9195 · S2

更大的模型显著更具样本效率,因此最优的算力高效训练方式是:在相对适中的数据量上训练非常大的模型,并在远未收敛时显著提前停止训练。Larger models are significantly more sample-efficient, such that optimally compute-efficient training involves training very large models on a relatively modest amount of data and stopping significantly before convergence.

CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes
CritICL:基于小语言模型失败模式的推理时弱到强泛化
arXiv:2608.27455 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

实验结果表明,CritICL 持续优于标准 in-context learning,并以显著更少的生成次数和更低的 token 成本,取得与 test-time scaling 方法相当或更优的性能。Experimental results show that CritICL consistently outperforms standard in-context learning and achieves performance competitive with or superior to test-time scaling methods, while requiring significantly fewer generations and lower token cost.

6. PLENA: Optimization Pathways for Long-Context Agentic LLM Inference
PLENA:面向长上下文 Agentic LLM 推理的优化路径
arXiv:2509.09505 Agent 智能体 方法 被引 14 · S2

PLENA 是一个软硬件协同设计的系统,采用三条核心优化路径,具备新颖的扁平化 systolic-array 架构以及支持非对称量化方案的高效计算与存储单元(路径 2)。PLENA is a hardware-software codesigned system that applies three core optimization pathways that features a novel flattened systolic-array architecture and efficient compute and memory units that support an asymmetric quantization scheme (Pathway 2).

EditaLive! Unified Character Video Editing for Live Streaming
EditaLive! 面向直播的统一人物视频编辑
arXiv:2608.27123 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

传统视频编辑主要关注场景级内容,而直播更强调人物主体。然而,直接将现有视频编辑方法应用于以人为中心的直播仍具挑战,因为它们可能引入面部表情不一致,且通常依赖多个离线推理步骤,难以满足实时交互需求。我们提出 EditaLive,一个用于实时流式角色视频编辑的新型框架。具体而言,我们从预训练图像动画模型(Wan-Animate)出发,该模型天然解耦...Conventional video editing primarily focuses on scene-level content, whereas live streaming places greater emphasis on the human subject. However, directly applying existing video-editing methods to human-centric live streaming remains challenging, as they may introduce facial-expression inconsistencies and typically depend on multiple offline inference steps, making them unsuitable for real-time interaction. We propose EditaLive, a novel framework for real-time streaming character video editing. In detail, we start from a pretrained image animation model (Wan-Animate), which naturally decoupl

TacForcing: Streaming Action Generation with Execution-Time Tactile Feedback
TacForcing:融合执行时间触觉反馈的流式动作生成
arXiv:2608.25798 多模态 方法 OA · 绿色 被引 1 · S2

提出 TacForcing,一种融合执行期触觉反馈的流式动作生成框架:按顺序生成动作块并保留未完成块的中间状态,并引入执行感知触觉注意力(EATA),将每次触觉更新的直接访问限制在下一个即将执行的块。TacForcing is introduced, a streaming action-generation framework incorporating execution-time tactile feedback that generates action blocks sequentially while preserving intermediate states of unfinished blocks and introduces Execution-Aware Tactile Attention (EATA), which restricts direct access to each tactile update to the next block scheduled for execution.

Luce: Relightable Gaussians for 3D Asset Generation
Luce:用于 3D 资产生成的可重光照高斯表示
arXiv:2608.23943 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Luce,一种三维表示方法,将几何与 PBR 材质统一到体素化的多模态高斯云中,分别使用专用高斯基元表示反照率、金属度-粗糙度与法线,在单图到三维生成任务上达到 SOTA。Luce is proposed, a 3D representation that unifies geometry and PBR materials within a voxelized multimodal Gaussian cloud, using dedicated Gaussian primitives for albedo, metallic-roughness, and surface normals to achieve state-of-the-art single-image-to-3D generation.

What Does an Evaluation License? A Commit-Bound Census of Claim-Relative Inference in Inspect Evals
评估授权了什么?Inspect Evals 中声明相对推理的提交绑定普查
arXiv:2608.19269 LLM 基础设施 观点 OA · 绿色 被引 0 · S2 + OpenAlex

冻结一个大规模评测集合,并尝试对其历史断言进行 replay,以使这一原本隐式的推理步骤变得显式且可执行。A large evaluation collection is frozen and an attempt to replay its historical claims is made to make this otherwise implicit inference step explicit and executable.

LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation
LayerRecall:面向视频生成长程一致性的状态条件记忆路由器。
arXiv:2608.28460 多模态 方法 OA · 绿色 被引 2 · S2

提出 LayerRecall,一种由当前状态条件化、按层选择的 memory router,仅从历史 K/V state 中检索相关状态并注入到 backbone 特定的 memory-sensitive 层中,其余层保留局部 attention;并提出 Cross-Horizon Prediction Matching (CHPM),借助特权的长上下文参考在预测空间中对有界 memory router 进行监督。This work introduces LayerRecall, a current-conditioned, layer-selective memory router that retrieves relevant historical K/V states and injects them only into backbone-specific memory-sensitive layers while preserving local attention elsewhere, and proposes Cross-Horizon Prediction Matching (CHPM), which uses a privileged long-context reference to supervise the bounded-memory router in prediction space.

CrabOS: An Operating System for Human-AI Co-inhabitation
CrabOS:面向人类与 AI 共生的操作系统。
arXiv:2608.28165 Agent 智能体 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

CrabOS 将人机交替主导复杂任务的支持从依赖桥接的应用层方案提升为原生操作系统能力,为开发与运行 AI agent 提供了新基础。CrabOS elevates support for complex tasks with alternating human and AI leadership from bridge-dependent application-level solutions to native operating-system capabilities, which provide a new foundation for developing and running AI agents.

J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data
J-Zero:从零数据出发的统一 Challenger--Solver--Judge 协同进化
arXiv:2608.26582 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Judge co-adaptation from Zero data (J-Zero),一个统一的 Challenger–Solver–Judge 自进化框架,支持在可验证与不可验证领域的自我提升,并识别出 Judge 协同进化是这一持续改进的关键驱动力。This work proposes Judge co-adaptation from Zero data (J-Zero), a unified Challenger--Solver--Judge self-evolution framework that supports self-improvement across both verifiable and unverifiable domains and identifies Judge co-adaptation as the key driver of this sustained improvement.

Revisiting Local Context for Long-Horizon Streaming 3D Reconstruction
重新审视长时序流式 3D 重建中的局部上下文
arXiv:2608.27529 LLM 基础设施 方法 被引 1 · S2

提出 ABot-Recon,一个简易的流式模型,仅缓存前 11 帧的 KV 特征,在当前相机坐标系下预测点图,并同时预测一个相对于相邻帧的相对位姿,该位姿在参考系变化下保持等变性。ABot-Recon is presented, a simple streaming model that caches KV features from only the preceding 11 frames and predicts a point map in the current camera coordinate system together with an adjacent-frame relative pose that remains equivariant under changes of reference frame.

Ring Forcing: Towards Precise Long-Term Memory for Autoregressive Video Diffusion
Ring Forcing:面向自回归视频扩散的精确长时记忆
arXiv:2608.26794 多模态 方法 OA · 绿色 被引 5 · S2

提出 Ring Forcing,一个自回归视频扩散框架,旨在稳健构建并精确利用长期记忆,并通过稀疏 RoPE 机制实现灵活、可扩展的记忆适配,同时充分利用预训练先验。Ring Forcing is presented, an autoregressive video diffusion framework designed to robustly construct and precisely utilize long-term memory and a sparse RoPE mechanism to enable flexible, scalable memory adaptation while fully exploiting pre-trained priors.