研究库 论文知识库
Papers · organized/paper_cards

论文

329 张论文卡片 · 多模态 · 方法

开放获取 全部 绿色 · 1640
PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation
PixelUMM:无编码器的统一图像与视频理解与生成
arXiv:2609.38597 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

统一多模态模型(UMM)通常依赖独立的视觉表示分别完成理解与生成,这增加了视觉上下文长度,并使其难以与既有视觉-语言预训练流程集成。近期 pixel-space modeling 的进展提供了一种无编码器的替代方案,但将该范式从图像扩展到视频并非易事:视频理解与生成采用不同的时间表示,统一视觉接口的设计仍是开放问题。本文提出 PixelUMM,一种用于统一图像与视频理解的无编码器模型……Unified Multimodal Models (UMMs) often rely on separate visual representations for understanding and generation, increasing visual context length and complicating integration with established vision-language pretraining pipelines. Recent advances in pixel-space modeling offer an encoder-free alternative, but extending this paradigm from images to videos is non-trivial: video understanding and generation adopt different temporal representations, leaving the design of a unified visual interface an open question. We present PixelUMM, an encoder-free model for unified image and video understanding

LOCI: Spatial Linear Memory for Streaming World Models
LOCI:面向流式世界模型的空间线性记忆
arXiv:2609.40222 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 LOCI,一种混合的空间记忆架构,同时维护键值缓存与循环记忆两种表示,在重访内容的复现上比代表性世界模型以及同配置的 full-softmax 模型都更为忠实。LOCI is introduced, a hybrid spatial-memory architecture that keeps both representations of key-value caches and recurrent memory that reproduces revisited content more faithfully than representative world models and a same-recipe full-softmax model.

Generative Adversarial Text to Image Synthesis
生成对抗式文本到图像合成
arXiv:1605.05396 多模态 方法 OA · 绿色 被引 3461 · S2

提出一种新颖的深度架构与 GAN 形式化方法,有效衔接文本与图像建模领域的进展,将视觉概念从字符转换为像素。A novel deep architecture and GAN formulation is developed to effectively bridge advances in text and image modeling, translating visual concepts from characters to pixels.

Honeycomb: Constant-Size Scene Memory Representation for Video World Models
Honeycomb:面向视频世界模型的恒定大小场景记忆表征
arXiv:2609.37690 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出了 Honeycomb,一种基于 HexMemory 构建的视频世界模型;HexMemory 是作者提出的低秩表示,用于在固定大小的内存中存储场景特征,共包含六个空间与时空平面。Honeycomb is introduced, a video world model built on HexMemory, the authors' proposed low-rank representation for storing scene features in a fixed-size memory with a total of six spatial and spatiotemporal planes.

ProAR: Learning Prospective Reasoning with Autoregressive Video Models
ProAR:基于自回归视频模型的前瞻性推理学习
arXiv:2610.03664 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

ProAR 引入两个关键组件:一是将生成锚定到长远结果,二是通过非对称注意力掩码将目标帧预测集成到自回归循环中,使预测的目标帧可指导中间状态的生成而不被其干扰。ProAR introduces two key components: to anchor generation to the long-range outcome, and to integrate goal-frame prediction into the autoregressive loop via an asymmetric attention mask, enabling the predicted goal frame to guide the generation of intermediate states without being disrupted by them.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
面向具身智能的 Mixture-of-Experts 视频预训练规模化
arXiv:2607.07675 多模态 方法 OA · 绿色 被引 15 · S2

提出 LingBot-Video,一种专为具身智能设计的基于 DiT 的视频预训练范式,并将其作为首个大规模开源 MoE 视频基础模型贡献给社区,致力于在数字创意与物理执行之间搭建桥梁。LingBot-Video is presented, a DiT-based video pretraining paradigm specifically tailored for embodied intelligence, and is contributed as the inaugural large-scale, open-source MoE video foundation model to the community, in a pioneering effort to bridge digital creativity and physical actuation.

MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation
MotorMind:为通用视觉语言模型搭建零样本机器人操控脚手架
arXiv:2609.38078 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出了 MotorMind,一个机器人操作框架,将 VLM 提出的中级动作与确定性机器人控制及反馈相连接,支持异步监控与后台记忆更新;研究表明,通用 VLM 在配备合适的中级动作表示和异步执行框架后,可执行有效的零样本机器人操作。MotorMind is introduced, a robot manipulation harness that connects VLM-proposed mid-level actions to deterministic robot control and feedback, with asynchronous monitoring and background memory updates, and shows that a general-purpose VLM, when equipped with an appropriate mid-level action representation and asynchronous execution harness, can perform effective zero-shot robotic manipulation.

World Action Modeling with Progressive Visual Planning
World Action Modeling with Progressive Visual Planning:基于渐进视觉规划的世界动作建模
arXiv:2610.02508 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 ProWAM,一种渐进式世界动作模型,联合预测动作和稀疏视觉子目标的有序序列,为任务执行过程中的动作生成提供显式视觉引导,展示了进度索引视觉前瞻对闭环控制的价值。ProWAM is presented, a progressive world action model that jointly predicts actions and an ordered sequence of sparse visual sub-goals, providing explicit visual guidance to anchor action generation throughout task execution, demonstrating the value of progress-indexed visual foresight for closed-loop control.

LVMT: Video Mask Transformer for Long-term Video Segmentation
LVMT:用于长期视频分割的视频掩码 Transformer
arXiv:2609.34895 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 TQP,一种训练策略,将视频按帧分块处理,块间传递被追踪对象信息,但仅在单个块内进行反向传播,从而在无显存溢出、推理开销或梯度消失问题的情况下实现更长时间范围的监督。TQP is introduced, a training strategy in which the model processes a video in chunks of frames, where information about tracked objects is propagated between chunks but backpropagation is only conducted in individual chunks, enabling longer temporal supervision without out-of-memory issues, inference overhead, or vanishing gradients.

Rollout-Marginal Distillation for Long-Horizon Autoregressive Video Generation
Rollout-Marginal Distillation for Long-Horizon Autoregressive Video Generation:面向长程自回归视频生成的 Rollout-Marginal 蒸馏
arXiv:2609.37925 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Rollout-Marginal Distillation,在 AR 预测中保留生成历史,但对每个 chunk 独立对照 chunk teacher 打分,确保其质量修正不受不完美的时序上下文影响。Rollout-Marginal Distillation is introduced, which retains the generated history for AR prediction but scores each chunk independently against a chunk teacher, ensuring its quality correction is not compromised by an imperfect temporal context.

PointWAM: 3D World Action Modeling for Dexterous Robotic Manipulation
PointWAM:面向灵巧机器人操作的 3D World Action Modeling
arXiv:2610.02840 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

一种 3D 世界动作模型,将世界分解为场景与手,并在共享时空坐标系内将二者联合预测为 3D 点轨迹,无需任务特定物体或关键点选择即可在大量人类演示视频上进行有效预训练。A 3D world action model that decomposes the world into a scene and hands and hands, and jointly forecasts both as 3D point trajectories within a shared space-time coordinate frame, which enables effective pre-training on large-scale human demonstration videos without requiring any task-specific object or keypoint selection.

EyeRobot 2.0: Active Gaze for Precise Manipulation without Wrist Cameras
EyeRobot 2.0:无需腕部摄像头的精确操作主动注视
arXiv:2610.03710 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

EyeRobot 2.0 通过旋转两个眼动视角将其注视中心对准场景中的 3D 注视点实现物理上的视觉聚焦,并将夹爪信息规范化到以注视点为参考的 SE(3) 坐标系,从而压缩待学习的动作分布规模。EyeRobot 2.0 physically attends to a 3D fixation point in the scene by swiveling two eye viewpoints to center their gaze on it, and takes advantage of fixation by canonicalizing gripper information into a fixation-relative SE(3) frame, which compacts the size of the action distribution to learn.

Infinite Worlds with Versatile Interactions
支持多样化交互的无限世界
arXiv:2607.07534 多模态 方法 OA · 绿色 被引 39 · S2

首次在 World Modeling 领域引入 Agentic Harness 框架:由 pilot agent 负责规划并执行角色行为,director agent 负责随着场景推进合成新的环境要素。The integration of an agentic harness within the domain of world modeling is pioneered, wherein a pilot agent is tasked with planning and executing character behaviors, while a director agent is responsible for synthesizing novel environmental elements as the scene progresses.

WM-VLM: Probing Internal World Models for Interleaved Visual-Textual Reasoning
WM-VLM:探查用于交错式视觉-文本推理的内部世界模型
arXiv:2609.34826 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 WM-VLM,为预训练 VLM 配备一个轻量级世界模型分支以生成中间视觉状态,表明内部世界模型为 VLM 在语言与视觉双重空间中的推理提供了一条有前景的路径。WM-VLM is introduced, which equips a pretrained VLM with a lightweight world model branch for generating intermediate visual states and suggests that internal world models offer a promising path toward VLMs that reason in both language and visual space.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation
面向机器人操作的视觉-语言-动作模型中的双潜在记忆
arXiv:2607.07608 多模态 方法 OA · 绿色 被引 3 · S2

提出 LaMem-VLA,一种以潜在记忆为核心框架的方法,将历史经验重建为潜在记忆 token,并直接与 VLA 推理交织,使记忆能够在有界上下文下直接参与 VLA 推理并引导动作生成。LaMem-VLA is introduced, a latent-memory-native framework that reconstructs historical experience into latent memory tokens and directly interweaves them with VLA reasoning, and enables memory to directly participate in VLA reasoning and guide action generation under a bounded context.

PerturBot: Breaking Shortcut Priors in Vision-Language-Action Models with Perturbative Training
PerturBot: 通过扰动训练打破视觉-语言-动作模型的捷径先验
arXiv:2610.04616 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Perturbot 和 GroundingFscore 提供了一套训练与评估框架,用于将 VLA 决策与捷径先验解耦,同时保持对任务相关证据的响应性。Perturbot and GroundingFscore provide a training-and-evaluation framework for disentangling VLA decisions from shortcut priors while preserving responsiveness to task-relevant evidence.

DiffGate: Difficulty-Gated Teacher Guidance for On-Policy Distillation
DiffGate:面向 On-Policy 蒸馏的难度门控教师指导
arXiv:2610.04596 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 DiffGate,一种将 GRPO 与选择性、有界教师指导相结合的结果门控目标,在四种模型–领域设置下均提升了 pass@8,表明在该评估协议下解的覆盖度得到改善。DiffGate is introduced, an outcome-gated objective that combines GRPO with selective, bounded teacher guidance, and improves pass@8 across all four model--domain settings, demonstrating improved solution coverage under the evaluation protocol.

Sensor-Language-Action Models
Sensor-Language-Action 模型(感知-语言-动作模型)
arXiv:2610.08244 多模态 方法

传感器不仅有助于理解世界,还有助于决定下一步动作。然而现有传感器模型大多止于感知层面:识别状态或预测结果,而动作则通过任务特定且通常封闭的标签空间单独建模。本文提出感知-语言-动作(SLA)建模,一个在统一模型中连接多模态传感器观测、自然语言与动作的框架。SLA 将语言用作感知与动作之间的语义接口,使异构动作得以表示、预测与解释……Sensors are useful not only for understanding the world but also for deciding what to do next. Existing sensor models however largely stop at perception: they recognize states or predict outcomes, leaving actions modeled separately through task-specific and often closed label spaces. We introduce Sensor-Language-Action (SLA) modeling, a framework that connects multimodal sensor observations, natural language, and actions within a unified model. SLA uses language as a semantic interface between sensing and acting, allowing heterogeneous actions to be represented, predicted, and explained while

DeCoPrune: Efficient KV-Cache Pruning for Autoregressive Video Diffusion via Denoising Consistency
DeCoPrune:通过去噪一致性实现自回归视频扩散的高效 KV-Cache 剪枝
arXiv:2609.39096 多模态 方法 OA · 绿色 被引 0 · OpenAlex

自回归视频扩散支持流式生成与交互控制,但其 KV cache 随生成历史持续增长。现有压缩策略要么采用固定窗口丢弃历史,要么依据局部注意力与相似度信号选择 token,均无法直接衡量当前块是否贡献了超出已保留上下文的信息。本文提出 DeCoPrune,一种将 cache 压缩视为去噪一致性问题的免训练方法。经验上发现,去噪难度可作为 token 价值的有用代理……Autoregressive video diffusion supports streaming generation and interactive control, but its KV cache grows continuously with the generated history. Existing compression strategies either discard history using fixed windows or select tokens through local attention and similarity signals, which do not directly measure whether the current chunk contributes information beyond the retained context. We introduce DeCoPrune, a training-free method that treats cache compression as a denoising-consistency problem. We find empirically that denoising difficulty provides a useful proxy for a token's valu

DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation
DMAD:面向快速视觉生成的分布匹配对抗蒸馏
arXiv:2610.02188 多模态 方法 OA · 绿色 被引 0 · OpenAlex

分布匹配蒸馏(DMD)通过分别估计的目标分数与学生分数之差训练少步学生模型,因而必须维护一个拟合学生动态分布的辅助扩散模型,带来额外的显存与计算开销。本文提出 DMAD——分布匹配对抗蒸馏,将分布匹配重构为分类任务,直接学习所需的 log-density 比。共享主干上的两个判别头分别区分真实数据与教师样本是否来自学生模型,并对其 logits 施加线性损失以训练……Distribution Matching Distillation (DMD) trains a few-step student from the difference between separately estimated target and student scores, so it must keep an auxiliary diffusion model fitted to the student's evolving distribution at extra memory and computation cost. We introduce DMAD, Distribution Matching as Adversarial Distillation, which recasts distribution matching as classification and learns the required log-density ratios directly. Two discriminator heads on a shared backbone distinguish real data and teacher samples from the student's, and linear losses on their logits train the

SheetSage2: Coherent Lead-Sheet Transcription with Synthetic Supervision
SheetSage2:基于合成监督的连贯主旋律谱转录
arXiv:2610.05336 多模态 方法 OA · 绿色 被引 0 · OpenAlex

将音乐转录为人类可读的乐谱需要对节奏、和声、旋律与曲式的整体理解。两大障碍限制了这一目标的实现:标注录音稀缺,以及精确的局部预测仍可能产生不一致的音乐序列。本文提出 SheetSage2,一个统一的音乐转录框架,结合合成数据、任务特定的结构化解码与自回归蒸馏。自动标注的 MIDI 渲染为音频后,可为各类音乐理解任务提供可扩展的监督。任务特定的结构化解码器整合互补的音乐……Transcribing music into a human-readable score requires a coherent understanding of rhythm, harmony, melody, and form. Two obstacles limit this goal: annotated recordings are scarce, and accurate local predictions can still produce inconsistent musical sequences. We present SheetSage2, a unified music transcription framework that combines synthetic data, task-specific structured decoding, and autoregressive distillation. Automatically annotated MIDI, rendered into audio, provides scalable supervision across music understanding tasks. Task-specific structured decoders integrate complementary mu

NAMVIS: Next-Scale Autoregressive Multi-View Image Synthesis
NAMVIS:下一尺度自回归多视图图像合成
arXiv:2610.04722 多模态 方法 OA · 绿色 被引 0 · OpenAlex

稀疏视图新视角合成是三维内容创作中的核心问题,但基于扩散的方法受迭代去噪限制,多视图生成在推理时开销高昂。本文提出 NAMVIS,一种无扩散框架,将多视图图像合成重新表述为几何条件下的下一尺度自回归过程。NAMVIS 不通过反复去噪生成目标视图,而是通过少量由粗到细的尺度步预测离散视觉 token,并在同一尺度内以及跨目标视图间并行采样所有 token。为锚定此过程……Sparse-view novel view synthesis is a central problem in 3D content creation, but diffusion-based approaches remain limited by iterative denoising, making multi-view generation expensive at inference time. We introduce NAMVIS, a diffusion-free framework that reformulates multi-view image synthesis as geometry-conditioned next-scale autoregression. Instead of generating target views through repeated denoising, NAMVIS predicts discrete visual tokens through a small number of coarse-to-fine scale steps, while sampling all tokens within each scale and across target views in parallel. To anchor thi

WildCity: A Real-World City-Scale Testbed for Rendering, Simulation, and Spatial Intelligence
WildCity:面向渲染、仿真与空间智能的真实城市规模测试平台
arXiv:2607.06838 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

WildCity 是一个由自动驾驶车队在复杂城市环境中采集的真实多模态数据集,旨在推动城市级渲染的进展,并更广泛地推动 AI 在空间感知、记忆与推理方面达到与人类认知相当规模的能力。WildCity, a real-world multimodal dataset collected by autonomous fleets traversing complex urban environments, aims to catalyze progress not only in city-scale rendering, but more broadly in the pursuit of AI that can perceive, remember, and reason across space at a scale comparable to human cognition.

OuroWorld: Bringing Any 3D World Alive as Diverse, Endlessly Looping 3D Cinemagraphs
OuroWorld:将任意 3D 世界激活为多样化、无缝循环的 3D Cinemagraph
arXiv:2610.12461 多模态 方法

现有的 3D 世界模型能够生成逼真、可探索的场景,但场景在时间维度上静止。OuroWorld 是一个无需掩码(mask-free)的框架,可将任意静态 3D Gaussian Splatting 场景转化为 3D Cinemagraph:从任意视角都能无缝循环、拥有生动多样运动的动态场景。视觉语言模型(vision-language model)推断合理的动态并引导视频模型合成参考视频,再将其提升并补全为多视图视频。为从这种不完美的监督中学习,我们提出 Inconsistency-Robust Periodic 4DGS:通过傅里叶级数形变场从构造上保证循环Recent 3D world models generate photorealistic, explorable scenes that remain frozen in time. OuroWorld is a mask-free framework that turns any static 3D Gaussian Splatting scene into a 3D cinemagraph: a dynamic scene with vivid, diverse motion looping seamlessly from any viewpoint. A vision-language model infers plausible dynamics and guides a video model to synthesize a reference video, which we lift and complete into multi-view videos. To learn from this imperfect supervision, we propose Inconsistency-Robust Periodic 4DGS: a Fourier-series deformation field guarantees looping by constructio

WorldGuide: Goal-Directed Video World Model for Procedural Task Execution
WorldGuide:面向过程化任务执行的目标导向视频世界模型
arXiv:2610.12459 多模态 方法

视频生成器和基于视频的世界模型能够合成合理的视觉轨迹,但长时程过程化任务要求生成过程能适应已实际生成的内容。模型必须根据其生成的状态决定下一步动作、执行该动作,并识别任务何时完成。开环(Open-loop)生成无法适应执行结果,而现有闭环(closed-loop)系统往往依赖预训练执行器或间接验证,导致动作决策与成功执行之间存在鸿沟。我们将过程化视频生成形式化为闭环Video generators and video-based world models can synthesize plausible visual trajectories, but long-horizon procedural tasks require generation to adapt to what has actually been produced. A model must determine the next action from its generated state, execute that action, and recognize when the task is complete. Open-loop generation cannot adapt to execution outcomes, while existing closed-loop systems often rely on pretrained executors or indirect verification. This leaves a gap between deciding an action and successfully realizing it. We formulate procedural video generation as closed-loo

One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
One Block, Multiple Depths:具有深度编程专家的循环视觉 Transformer
arXiv:2610.12448 多模态 方法

在本工作中,我们证明单个 Transformer block 循环应用即可在相近推理 FLOPs 下匹配全深度视觉编码器的精度,且无需中间特征蒸馏。reViT 通过将每个循环深度的 FFN 表示为小型共享专家库的凸组合来恢复深度特定的变换。一个连续归一化的深度坐标对该混合进行编程,在 FFN 参数空间中定义一条可重采样的轨迹。我们在两种场景下评估该设计:有监督 ImageNet-1k 训练以及从 DINOv2 教师模型蒸馏。在各规模下In this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank. A continuous normalized-depth coordinate programs this mixture, defining a resampleable trajectory through FFN parameter space. We evaluate this design in two regimes: supervised ImageNet-1k training and distillation from a DINOv2 teacher. Across b

MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing
MV-Forcing:通过 4D 几何基础的时空自强迫实现长多视角视频生成
arXiv:2607.05376 多模态 方法 OA · 绿色 被引 1 · S2

提出 MV-Forcing 框架,通过在顺序生成视角之间引入 4D 几何桥梁,在单一扩散模型中组合时间与视角自回归,并弥合时间与视角序列自回归中训练-推理的曝光偏差差距。MV-Forcing, a framework that composes temporal and view-wise autoregression within a single diffusion model by introducing a 4D geometric bridge between sequentially generated views, and closes the train-inference exposure bias gap for both temporal and view-sequential autoregression.

AI Wizards at EXIST 2026: Hierarchical Soft-Label Learning for Multimodal Sexism Identification in Memes
AI Wizards 参加 EXIST 2026:用于迷因中多模态性别歧视识别的分层软标签学习
arXiv:2607.04410 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作通过轻量级 Gated MLP 将固定的 Gemini Embedding 2 视觉-语言表征映射到目标空间,使用 KL 散度与同方差不确定性加权进行训练,并展示了 AI Wizards 参加 EXIST 2026 多模态迷因性别歧视识别任务的提交方案。This work maps fixed Gemini Embedding 2 vision-language representations through a lightweight Gated MLP trained with KL divergence and homoscedastic uncertainty weighting, and presents the AI Wizards submission to EXIST 2026 for multimodal sexism identification in memes.

Gemma 4 Technical Report
Gemma 4 技术报告
arXiv:2607.02770 多模态 方法 OA · 绿色 被引 206 · S2

本工作推出 Gemma 4——Gemma 模型系列中新一代开源权重、原生多模态的语言模型,在 STEM、多模态与长上下文基准上实现性能跃升,在人类评分任务上可与更大的前沿开源模型相媲美。This work introduces Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family that establishes a leap in performance across STEM, multimodal, and long-context benchmarks, and rivals larger, frontier open models in human-rated tasks.

OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers
OrbitQuant:面向图像与视频扩散 Transformer 的数据无关量化
arXiv:2607.02461 多模态 方法 OA · 绿色 被引 5 · S2

本文提出 OrbitQuant,一种数据无关的权重量化器,通过在归一化旋转基空间中进行量化以绕过范围估计,将图像扩散 Transformer 的 PTQ 推进到 W2A4 并保持可用生成质量。OrbitQuant is presented, a data-agnostic weight-activation quantizer that bypasses range estimation by quantizing in a normalized, rotated basis and pushes PTQ of image diffusion transformers to W2A4 with usable generation quality.

Generated Contents Enrichment
生成内容增强
arXiv:2405.03650 多模态 方法 OA · 绿色 被引 1 · S2

本文提出一个联合训练的对抗框架,通过建模对象语义和对象间关系来增强场景图,并在 Visual Genome 数据集上以代理场景图增强指标、图像质量比较、定性示例与用户研究进行评估。A jointly trained adversarial framework is proposed that enriches scene graphs by modeling object semantics and inter-object relations and is evaluated with proxy scene graph enrichment metrics, image-quality comparisons, qualitative examples, and user studies on the Visual Genome dataset.

VLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action Horizon
VLA-Corrector:面向自适应动作时域的轻量级检测-修正推理
arXiv:2607.01804 多模态 方法 OA · 绿色 被引 15 · S2

VLA-Corrector:面向动作分块 VLA 策略的轻量级修正推理框架;引入轻量的潜空间视觉监控器,持续比对预测与实际视觉特征演化,可在线检测视觉动态偏差,缓解静态时域在执行鲁棒性与策略调用频率之间的权衡。VLA-Corrector, a lightweight corrective inference framework for action-chunked VLA policies, introduces a lightweight Latent-space Vision Monitor that continuously compares predicted and actual visual feature evolution, enabling online detection of visual dynamics deviations and mitigates the trade-off imposed by static horizons between execution robustness and policy-call frequency.

From SRA to Self-Flow: Data Augmentation or Self-Supervision?
从 SRA 到 Self-Flow:数据增强还是自监督?
arXiv:2607.02508 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Attention Separation,在保留与 Self-Flow 相同的双时间步输入的同时,阻止被分配到不同噪声水平 token 之间的注意力,并表明 Attention Separation 本身通过将单张图像拆分为多个有效训练部分来扩充训练数据,从而带来增强效果。Attention Separation is introduced, which preserves the same dual-timestep input as Self-Flow while blocking attention between tokens assigned to different noise levels, and shows that Attention Separation itself provides an augmentation effect by splitting a single image into multiple effective training parts to expand the training data.

Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning
基于非对称互变分学习的多模态连续推理
arXiv:2607.00461 多模态 方法 OA · 绿色 被引 1 · S2

AMVL 在潜空间集成的 MLLM 中实例化,持续优于多种强离散和潜空间推理基线,在复杂 BLINK 基准上平均得分提升 +10.83,单个推理任务最高提升 +32.00,分析也确认了潜空间稳定性的改善。AMVL is instantiate in a latent-integrated MLLM and it consistently outperforms strong discrete and latent-reasoning baselines, improving the average score on the complex BLINK benchmark by +10.83 and achieving gains of up to +32.00 on individual reasoning tasks, with analyses confirming improved latent-space stability.

DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image Generation
DataEvolver:面向富文本图像生成的自演化多 Agent 数据构建
arXiv:2606.31537 多模态 方法 OA · 绿色 被引 2 · S2

在富文本图像生成基准上的实验表明,在数据预算匹配条件下,DataEvolver 比固定数据集基线生成更有用的训练数据,且结果表明被拒绝的样本可为改进富文本图像数据构建提供可操作的反馈信号。Experiments on text-rich image generation benchmarks show that DataEvolver produces more useful training data than fixed-dataset baselines under matched data budgets, and results suggest that rejected samples can provide actionable feedback for improving text-rich image data construction.

QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents
QVal:低成本评估面向长 horizon LLM Agent 的密集监督信号
arXiv:2606.32034 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 QVal,一个无需训练即可直接评估密集监督信号的测试平台,发现简单提示基线始终优于文献中近期提出的密集监督方法,且性能按模型家族高度聚类。QVal, a training-free testbed for directly evaluating dense supervision signals, is introduced, finding that simple prompting baselines consistently outperform recent dense supervision methods from the literature, and that performance clusters strongly by family.