Papers · organized/paper_cards

论文

189 张论文卡片 · 多模态 · 方法

开放获取 全部 绿色 · 724
From SRA to Self-Flow: Data Augmentation or Self-Supervision?
从 SRA 到 Self-Flow:数据增强还是自监督?
arXiv:2607.02508 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Attention Separation,在保留与 Self-Flow 相同的双时间步输入的同时,阻止被分配到不同噪声水平 token 之间的注意力,并表明 Attention Separation 本身通过将单张图像拆分为多个有效训练部分来扩充训练数据,从而带来增强效果。Attention Separation is introduced, which preserves the same dual-timestep input as Self-Flow while blocking attention between tokens assigned to different noise levels, and shows that Attention Separation itself provides an augmentation effect by splitting a single image into multiple effective training parts to expand the training data.

Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning
基于非对称互变分学习的多模态连续推理
arXiv:2607.00461 多模态 方法 OA · 绿色 被引 1 · S2

AMVL 在潜空间集成的 MLLM 中实例化,持续优于多种强离散和潜空间推理基线,在复杂 BLINK 基准上平均得分提升 +10.83,单个推理任务最高提升 +32.00,分析也确认了潜空间稳定性的改善。AMVL is instantiate in a latent-integrated MLLM and it consistently outperforms strong discrete and latent-reasoning baselines, improving the average score on the complex BLINK benchmark by +10.83 and achieving gains of up to +32.00 on individual reasoning tasks, with analyses confirming improved latent-space stability.

DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image Generation
DataEvolver:面向富文本图像生成的自演化多 Agent 数据构建
arXiv:2606.31537 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

在富文本图像生成基准上的实验表明,在数据预算匹配条件下,DataEvolver 比固定数据集基线生成更有用的训练数据,且结果表明被拒绝的样本可为改进富文本图像数据构建提供可操作的反馈信号。Experiments on text-rich image generation benchmarks show that DataEvolver produces more useful training data than fixed-dataset baselines under matched data budgets, and results suggest that rejected samples can provide actionable feedback for improving text-rich image data construction.

QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents
QVal:低成本评估面向长 horizon LLM Agent 的密集监督信号
arXiv:2606.32034 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 QVal,一个无需训练即可直接评估密集监督信号的测试平台,发现简单提示基线始终优于文献中近期提出的密集监督方法,且性能按模型家族高度聚类。QVal, a training-free testbed for directly evaluating dense supervision signals, is introduced, finding that simple prompting baselines consistently outperform recent dense supervision methods from the literature, and that performance clusters strongly by family.

MuSViT: A Foundation Vision Model for Sheet Music Representation
MuSViT:面向乐谱表示的基础视觉模型
arXiv:2606.31811 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本工作提出 MuSViT (Music Score Vision Transformer):首个面向乐谱表征的基础视觉模型——一个通过 Masked Autoencoders 在 IMSLP 970 万页数据上预训练的 ViT 编码器。This work introduces MuSViT (Music Score Vision Transformer): the first foundation vision model for sheet music representation -- a ViT encoder pre-trained via Masked Autoencoders on 9.7 million pages from the IMSLP.

Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation
照亮统一多模态模型:面向自由形式交错图文生成
arXiv:2606.30054 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 ILLUME-X,一种先进的统一多模态范式,通过提升多模态数据效率并稳定多模态训练过程,实现高质量、自由形式的交错图文生成。This paper introduces ILLUME-X, an advanced unified multimodal paradigm that enables high-quality, free-form interleaved text-image generation by improving multimodal data efficiency and stabilizing the multimodal training process.

Parameter-Efficient Quantum-Inspired Fast Weight Programmers for Traffic-Matrix Forecasting
用于流量矩阵预测的参数高效量子启发快速权重编程器
arXiv:2606.27821 多模态 方法 OA · 绿色 被引 3 · S2

本文将门控量子启发的 Kolmogorov-Arnold 网络快权重编程器用于直接多步 Abilene 流量矩阵预测,提出以经典慢速编程器搭配量子启发快速编程器的方案,作为面向资源受限场景的网络流量矩阵预测中一种兼顾精度与效率的有前景设计。This paper adapts gated quantum-inspired Kolmogorov-Arnold network fast-weight programmers to direct multi-step Abilene TM forecasting and identifies a classical slow programmer with a quantum-inspired fast programmer as a promising accuracy-efficiency design for resource-conscious network traffic-matrix forecasting.

PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation
PhysRAG:通过检索增强生成提升视频生成中的物理感知
arXiv:2606.26916 多模态 方法 OA · 绿色 被引 1 · S2

本文提出 PhysRAG——一条通过检索增强生成(RAG)提升视频生成物理感知的新流程,并基于 WISA-80K 数据集设计了两阶段数据过滤流程,最终筛选出 7K 高质量视频用于训练。This work introduces PhysRAG, a novel pipeline that enhances physical awareness in video generation through Retrieval-Augmented Generation (RAG), and designs a two-stage data filtering pipeline based on the WISA-80K dataset, resulting in a curated set of 7K high-quality videos for training.

Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)
Learning to Fold:LeHome Challenge 2026 获奖方案(线上第 1,线下第 2)
arXiv:2606.27163 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本工作通过强化学习循环改进视觉-语言-动作(VLA)策略,该循环预测成功、进展及若干任务相关的未来量,并驱动优势估计、实时失败检测与候选选择,在 LeHome Challenge 2026 中取得佳绩。The work improves a vision-language-action (VLA) policy with a reinforcement-learning loop that predicts success, progress, and a few task-relevant future quantities and drives advantage estimation, live failure detection, and candidate selection in the LeHome Challenge 2026.

EO-WM: A Physically Informed World Model for Probabilistic Earth Observation Forecasting
EO-WM:面向概率性地球观测预报的物理信息世界模型
arXiv:2606.27277 多模态 方法 OA · 绿色 被引 1 · S2

地球观测(EO)预报旨在依据变化的天气条件,从卫星观测预测未来地表动态。本文将其建模为部分可观测、天气驱动的世界建模问题,其中天气作为条件信号,而由于观测稀疏和未观测的陆面状态,预报本身具有不确定性。然而现有方法未能完整刻画这一设定:确定性模型将不确定性坍缩为单一未来预测,而基于扩散的方法通常将天气变量视作无条(原文此句截断)。Earth Observation (EO) forecasting aims to predict future Earth surface dynamics from satellite observations under changing meteorological conditions. In this paper, we view this task as a partially observed, weather-driven world modeling problem, in which weather acts as a conditioning signal, while forecasting remains uncertain due to sparse observations and unobserved land-surface states. However, existing methods do not fully capture this setting: deterministic models collapse uncertainty into a single future prediction, while diffusion-based methods typically treat weather variables as un

Dziri Voicebot: An End-to-End Low-Resource Speech-to-Speech Conversational System for Algerian Dialect
Dziri Voicebot:面向阿尔及利亚方言的端到端低资源语音对话系统
arXiv:2606.26003 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

所提系统为阿尔及利亚方言的端到端对话建模提供了可复现基线;实验结果显示各组件均表现优异:ASR 词错误率低,NLU 意图分类与实体识别得分高,语音合成质量稳定。The proposed system provides a reproducible baseline for end-to-end conversational modeling in Algerian Dialect, and experimental results show strong performance across all components, including low word error rate for ASR, high intent classification and entity recognition scores for NLU, and stable speech synthesis quality.

Distribution-Aware Diffusion-LLM for Robust Ultra-Long-Term Time Series Forecasting
用于鲁棒超长期时间序列预测的分布感知扩散 LLM
arXiv:2606.23391 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出新框架 Diffusion-LLM,将条件扩散模型集成到基于 LLM 的预测流水线中,展示了分布感知正则化在提升时间序列 LLM 的鲁棒性与泛化能力方面的价值。This work proposes a new framework Diffusion-LLM that integrates a conditional diffusion model into an LLM-based forecasting pipeline, and demonstrates the value of distribution-aware regularization for enhancing robustness and generalization in time series LLMs.

RaysUp: Ultra-light Universal Feature Upsampling via Geometry-Aware Ray Representation
RaysUp:基于几何感知光线表示的超轻量通用特征上采样
arXiv:2606.22749 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 RaysUp,一个超轻量级、任务无关且与 VFM 无关的特征上采样框架,可在任意分辨率下重建高分辨率特征图,仅使用 AnyUp 16% 的参数即达到 SOTA 性能,推理速度提升约 7 倍。RaysUp is proposed, an ultra-lightweight, task-agnostic, and VFM-agnostic feature upsampling framework that reconstructs high-resolution feature maps at arbitrary resolutions that achieves state-of-the-art performance while using only 16% of the parameters of AnyUp and delivering approximately 7x faster inference.

MedRLM: Recursive Multimodal Health Intelligence for Long-Context Clinical Reasoning, Sensor-Guided Screening, Evidence-Grounded Decision Support, and Community-to-Tertiary Referral Optimization
MedRLM:用于长上下文临床推理、传感器引导筛查、循证决策支持和社区到三级转诊优化的递归多模态健康智能
arXiv:2606.20164 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

MedRLM 旨在将医疗 AI 从静态问答转向可审计、多模态且工作流感知的临床决策支持,并引入临床证据图记忆,将患者特定观察与检索到的证据相连接。MedRLM aims to move medical AI from static question answering toward auditable, multimodal, and workflow-aware clinical decision support, and introduces a Clinical Evidence Graph Memory to connect patient-specific observations with retrieved evidence.

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models
CogniRoute:全模态模型中的社交证据路由学习
arXiv:2606.20970 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 CogniRoute,一种面向社交全模态推理的 schema 引导 Mixture-of-Experts(混合专家)框架,并引入路由感知强化学习,通过答案正确性、模态一致性推理与认知时序锚定等奖励联合优化 token 生成与专家分配。CogniRoute, a schema-guided Mixture-of-Experts framework for social omni reasoning, is introduced and route-aware reinforcement learning is introduced, which jointly optimizes token generation and expert allocation using rewards for answer correctness, modality-consistent reasoning, and cognitive temporal grounding.

RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures
RoboTALES:通过任务对齐的模拟未来学习推理引导的机器人策略
arXiv:2607.06018 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 RoboTALES,一个学习任务对齐模拟未来并据此训练机器人策略的单阶段框架,引入两项关键创新:基于 LLM 的分层规划器,将复杂任务拆解为子目标序列以引导模型的"想象";基于 VLM 的评判器,用于评估这些"想象"出的未来。This work proposes RoboTALES, a single-stage framework that learns task-aligned simulated futures and uses them to train robot policies and introduces two key innovations: a hierarchical LLM-based planner that breaks complex tasks into a sequence of subgoals to guide the model's imagination and a VLM-based critic that evaluates these ``imagined'' futures.

Token-Based Dual-view Fusion and Adaptation of Large Vision Models for Breast Cancer Classification
基于 Token 的双视图融合与适配用于乳腺癌分类的大视觉模型
arXiv:2607.06309 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

一种以 Token 为中心的双视图学习框架,在冻结的视觉 Transformer 主干中统一基于 prompt 的适配与跨视图融合,相比线性探测、仅 prompt 适配以及传统融合基线均取得稳定提升。A token-centric dual-view learning framework that unifies prompt-based adaptation and cross-view fusion within a frozen vision transformer backbone and demonstrates consistent improvements over linear probing, prompt-only adaptation, and conventional fusion baselines.

Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs
为触觉而醒!MLLM 中基于掩码隔离的触觉对齐学习
arXiv:2607.00302 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

[摘要] Splash 是一个面向 MLLMs 的掩码隔离触觉对齐学习框架,它量化每个预训练参数的重要性,并将参数空间划分为休眠子空间与关键子空间,从而有效防止灾难性遗忘,确保非破坏性的模态扩展。Splash is presented, a mask-isolated tactile alignment learning framework for MLLMs that quantifies the significance of each pretrained parameter, and partitions the parameter space into a dormant and critical subspace, which effectively prevents catastrophic forgetting and ensures non-destructive modality expansion.

CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generation
CineMobile:面向电影级相机运动生成的端侧图像到视频扩散模型
arXiv:2607.03803 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

CineMobile 采用三重优化策略,通过蒸馏引导的剪枝方法得到一个紧凑而高效的模型,保留实现电影级效果所需的核心视频生成能力,证明了其在移动端图像到视频创作中的实用可行性。CineMobile adopts a three-fold optimization strategy, leveraging a distillation-guided pruning approach to derive a compact yet efficient model that retains the essential video generation capabilities required for cinematic effects, demonstrating its practical applicability for mobile-based image-to-video creation.

LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models
LongE2V:基于视频扩散模型的长时间跨度事件驱动视频重建、预测与帧插值
arXiv:2607.08770 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 LongE2V,一种利用预训练视频扩散先验来联合处理基于事件视频重建、预测与帧间插值的新方法,并引入自回归展开与自适应上下文切换机制,以缓解超长序列中的时序漂移问题。This work proposes LongE2V, a novel approach that leverages pre-trained video diffusion priors to jointly handle event-based video reconstruction, prediction, and frame interpolation, and introduces Autoregressive Unrolling and Adaptive Context Switching to mitigate temporal drift in extremely long sequences.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation
SAM-MT:实时交互式多目标视频分割
arXiv:2607.08688 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

SAM-MT 成功将延迟与目标数量解耦,在保持 SAM2 鲁棒视频分割性能的同时,实现了与单目标基线相当的实时速度。SAM-MT successfully decouples latency from the number of targets, achieving real-time speed on par with single-target baselines while maintaining SAM2's robust video segmentation performance.

Phone Segmentation and Recognition through Phonological Activation Mapping
基于音韵激活映射的音素切分与识别
arXiv:2607.09020 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

论文主张语音结构已隐含在自监督语音模型(S3M)的表征中,只需对其进行引导即可同时完成切分与识别任务。It is argued that phonetic structure is already latent in the representations of self-supervised speech models (S3Ms), and one only needs to steer them to solve both segmentation and recognition tasks.

MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models
MedPMC:面向基础模型的高保真医学多模态数据规模化系统框架
arXiv:2607.07673 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 MedPMC——一种自动化、可持续更新的框架,可将宽松许可的文献转化为面向医学多模态模型的高保真基础设施,并公开发布该框架、语料库、基准与预训练模型。MedPMC, an automated, continuously updatable framework that transforms permissively licensed literature into high-fidelity infrastructure for medical multimodal models, is introduced and publicly release the framework, corpus, benchmarks, and pretrained models.

VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery
VaseMuseum:古希腊陶器数字智能博物馆
arXiv:2607.06374 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 VaseMuseum——一种面向古希腊陶器智能数字博物馆的轻量化、模块化多模态智能体框架,相比启用搜索的 VLM 基线,它提升了引用有效性,减少了知识密集型查询中的幻觉,并在含歧义场景下给出更中立的回答。VaseMuseum is proposed, a lightweight and modular multimodal agent framework for intelligent digital museums of ancient Greek pottery that improves citation validity, reduces hallucinations on knowledge-intensive queries, and produces more neutral answers under ambiguity compared with search-enabled VLM baselines.

4D Human-Scene Reconstruction from Low-Overlap Captures
低重叠度采集下的 4D 人体场景重建
arXiv:2607.09125 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 StudioRecon,一种通过解耦背景与人体、并利用视频扩散模型合成数百个相机可控新视角,从稀疏低重叠相机重建 4D 人体场景的流水线,在四个真实数据集上达到了 SOTA 的新视角合成效果StudioRecon is proposed, a pipeline that reconstructs 4D human scenes from sparse, low-overlap cameras by decoupling background and humans by synthesizing hundreds of camera-controlled novel views with a video diffusion model and achieves state-of-the-art novel view synthesis across four real-world datasets.

CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation
CtrlVTON:基于视觉实例提示分割的可控虚拟试穿
arXiv:2607.09362 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

推出 CtrlVTON,一个将试穿重构为图像编辑问题并引入分割掩码作为对服装布局(包括风格、尺寸与身体空间位置)像素级控制的可控 VTO 框架CtrlVTON is introduced, a controllable VTO framework that recasts try-on as an image editing problem and adds segmentation masks as pixel-level control over garment layout, including style, size, and spatial placement on the body.

LATO.2: Factorized 3D Mesh Generation with Vertex and Topology Flow
LATO.2:基于顶点流与拓扑流分解的 3D 网格生成
arXiv:2607.10623 多模态 方法 OA · 绿色 被引 2 · S2

提出 LATO.2,一个因子化 flow matching 框架,将网格生成分解为 vertex flow 和随后以已实现顶点为条件的 connectivity flow,在几何保真度和连通性质量上超越 SOTA 的拓扑感知网格生成方法。LATO.2, a factorized flow matching framework that decomposes mesh generation into a vertex flow followed by a connectivity flow conditioned on the realized vertices, is presented, which surpasses state-of-the-art topology-aware mesh generators in geometric fidelity and connectivity quality.

Latent-Identity Tuning in Text-to-Image Personalization Models
文本到图像个性化模型中的潜空间身份调优
arXiv:2607.11885 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文探索了一个预训练、冻结 encoder 的潜空间用于 text-to-image 个性化,并表明可以在该空间及由选定 token 定义的子空间中识别出有意义的编辑方向,从而实现局部化、细粒度且语义一致的编辑。This work explores the latent space of a pre-trained, frozen encoder for text-to-image personalization, and shows that meaningful directions can be identified within this space and within subspaces defined by selected tokens, enabling localized, fine-grained, and semantically coherent edits.

Evidence-Backed Video Question Answering
证据支撑的视频问答
arXiv:2607.11862 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 ST-Evidence,首个同时面向判别式和生成式像素级 grounding 的人工验证 benchmark,并开发可扩展的自动生成流程,构建了 16 万规模、衔接高层推理与细粒度 grounding 的数据集 ST-Evidence-Instruct。ST-Evidence is introduced, the first human-verified benchmark for both discriminative and generative pixel-level grounding, and scalable, automated generation pipelines are developed to create ST-Evidence-Instruct, a 160k-scale dataset bridging high-level reasoning with fine-grained grounding.

A Theory of Contrastive Learning with Natural Images
自然图像对比学习的一种理论
arXiv:2607.07470 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

针对一系列基本增广与任意具有平稳统计量的图像数据集,以解析方式根据对比损失计算最优表示,结果表明对于某些增广,最优解可由第一层滤波器为正弦函数的 CNN 实现。Analytically computing the optimal representation in terms of a contrastive loss for a range of basic augmentations and any image dataset with stationary statistics shows that for certain augmentations the optimum can be attained by a CNN whose first layer filters are sinusoids.

MAGIC: Transition-Aware Generation of Navigable Multi-Scene Game Worlds with Large Language Models
MAGIC:基于 LLM 的转换感知可导航多场景游戏世界生成
arXiv:2607.11594 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

MAGIC 是一个四阶段 pipeline,能将单一自然语言提示转化为可运行的多场景游戏项目,相比 LLM 基线与 Holodeck 可恢复更多真实 portal,并生成显著更可导航的布局。MAGIC is a four-stage pipeline that turns a single natural-language prompt into a runnable multi-scene game project that recovers more ground-truth portals and yields markedly more navigable layouts than an LLM baseline and Holodeck.

Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation
Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation
arXiv:2607.12752 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Hallo4D 提出"生成-检测-修正"范式,利用大型多模态语言模型(LMMs)从多视角与多帧渲染中识别并归纳时空不一致性,为一致性感知的内容生成提供了一种可扩展且可泛化的方案。Hallo4D introduces a generation-detection-correction paradigm that leverages large multimodal language models (LMMs) to identify and summarize spatial and temporal inconsistencies from multi-view and multi-frame renderings, providing a scalable and generalizable solution for consistency-aware content generation.

Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation
Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation
arXiv:2607.13125 多模态 方法 OA · 绿色 被引 1 · S2

研究表明,通过更强的多模态编码器、Agentic prompt 改写及相关技术来增强 Boogu-Image 系统的理解能力,并结合数据质量、训练流程和 Agentic 推理时扩展的改进,即使在计算预算极为受限的条件下,也能显著提升生成与编辑性能。It is demonstrated that strengthening the understanding capability of the Boogu-Image system, through a stronger multimodal encoder, agentic prompt rewriting, and related techniques, together with improvements in data quality, training pipelines, and agentic inference-time scaling, can substantially enhance generation and editing performance even under highly constrained compute budgets.

Registers Matter for Pixel-Space Diffusion Transformers
Registers 对像素空间 Diffusion Transformer 至关重要
arXiv:2605.16147 多模态 方法 OA · 绿色 被引 2 · S2

本研究表明 DiT 与 ViT 在一个关键方面存在差异:DiT 不会出现 patch-token 异常值,但仍能受益于 registers;并且 registers 在像素空间 DiT 中比在潜空间 DiT 中效果更显著。This work shows that DiTs differ from ViTs in a key respect: they do not exhibit patch-token outliers but still benefit from registers, and finds that registers are more effective in pixel-space DiTs than in latent-space DiTs.

AffectFlow-DINO: Uncertainty-Aware Multi-Task Affect Estimation via Conditional Rectified Flow
AffectFlow-DINO:基于条件 Rectified Flow 的不确定性感知多任务情感估计
arXiv:2607.13250 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

面向第 11 届 ABAW 挑战赛的多任务学习系统,在标准确定性架构基础上扩展条件 Rectified Flow 头,建模真实场景下面部行为固有的模糊性,借助蒙特卡洛采样实现不确定性感知的一对多预测。A multi-task learning system for the 11th ABAW challenge that extends a standard deterministic architecture with a conditional rectified-flow head to model the inherent ambiguity of in-the-wild facial behavior, enabling uncertainty-aware one-to-many predictions through Monte Carlo sampling.

Hierarchical Denoising For Multi-Step Visual Reasoning
面向多步视觉推理的分层去噪方法
arXiv:2607.15278 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 HDR (Hierarchical Denoising for Visual Reasoning),一个将层级潜变量集成到因果视频生成中以进行多步推理的统一框架,并引入一个含分布外情况的层级化多步视频推理基准。This work proposes HDR (Hierarchical Denoising for Visual Reasoning), a unified framework that integrates hierarchical latents into causal video generation for multi-step reasoning and introduces a level-stratified multi-step video reasoning benchmark with out-of-distribution cases.