研究库 论文知识库
Papers · organized/paper_cards

论文

313 张论文卡片 · 多模态 · 方法 · OA 绿色

开放获取 全部 绿色 · 1640
元信息
arXiv:2602.04476 多模态 方法 Open MIND OA · 绿色 被引 13 · S2

本文提出 Vision-aligned Latent Reasoning(VaLR),一种简洁而有效的推理框架,在每个 Chain of Thought 推理步骤之前动态生成视觉对齐的 latent token,引导模型在 latent space 中基于感知线索进行推理。Vision-aligned Latent Reasoning (VaLR) is introduced, a simple, yet effective reasoning framework that dynamically generates vision-aligned latent tokens before each Chain of Thought reasoning step, guiding the model to reason based on perceptual cues in the latent space.

论文信息
arXiv:2606.17053 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 ContextRL,一种上下文感知的强化学习方法,通过间接辅助目标提升长周期推理与多模态性能,并与将相同对比上下文复用作标准 query–context–answer 样本的数据增强基线进行对比。This work proposes ContextRL, a context-aware reinforcement learning (RL) method that improves long-horizon reasoning and multimodal performance through an indirect auxiliary objective, and compares against data-augmentation baselines that repurpose the same contrastive contexts as standard query--context--answer examples.

MMProLong:长上下文视觉语言模型的有效续训练(精读 · flyP)
arXiv:2605.13831 多模态 方法 OA · 绿色 被引 1 · S2

本研究建立了一套实用的 LongPT 方案,为推进长上下文 vision-language 模型奠定了经验基础,并提出 MMProLong,无需任务专属监督即可泛化至基于网页的多模态 needle 检索、长上下文图文压缩以及长视频理解等任务。This study establishes a practical LongPT recipe and an empirical foundation for advancing long-context vision-language models, and introduces MMProLong, which generalizes to webpage-based multimodal needle retrieval, long-context vision-text compression, and long-video understanding without task-specific supervision.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
LLaDA-V:基于视觉指令微调的大语言扩散模型
arXiv:2505.16933 多模态 方法 OA · 绿色 被引 154 · S2

本文提出 LLaDA-V,一种完全基于扩散范式的多模态大语言模型 (MLLM),将视觉指令微调与 masked diffusion 模型相结合,脱离了当前多模态方法中主流的自回归范式。LLaDA-V is introduced, a purely diffusion-based Multimodal Large Language Model (MLLM) that integrates visual instruction tuning with masked diffusion models, representing a departure from the autoregressive paradigms dominant in current multimodal approaches.

条目S1:To Data & Beyond — Important LLM Papers Week of 12-17 Jan 2026
条目S1:To Data & Beyond — Important LLM Papers Week of 12-17 Jan 2026
arXiv:2601.09668 多模态 方法 OA · 绿色 被引 27 · S2

本文推出 STEP3-VL-10B,一个面向"紧凑效率与前沿级多模态智能"权衡的轻量级开源基础模型,并发布完整模型套件,为社区提供强大、高效且可复现的 baseline。STEP3-VL-10B is presented, a lightweight open-source foundation model designed to redefine the trade-off between compact efficiency and frontier-level multimodal intelligence, and the full model suite is released to provide the community with a powerful, efficient, and reproducible baseline.

条目R1:MAGMaR 2026 Shared Task — 多模态增强生成的ACL 2026 Workshop(arXiv 2606.12295)
arXiv:2606.12295 多模态 方法 OA · 绿色 被引 2 · S2

本文概述了第二届 MAGMaR(Multimodal Retrieval 驱动的多模态增强生成)研讨会共享任务的成果,参赛系统聚焦于视频检索,或在给定检索视频的基础上进行有依据的文章生成。This overview paper presents the results of the shared task for the second workshop on Multimodal Augmented Generation via Multimodal Retrieval (MAGMaR), where participants submitted systems focused on either video retrieval or grounded generation of articles given retrieved videos.

MOSS-VL Technical Report
MOSS-VL 技术报告
arXiv:2608.15045 多模态 方法 OA · 绿色 被引 3 · S2

我们提出 MOSS-VL,一个开源视觉-语言模型系列,将实时交互(边说边看)视为一等能力。它贯穿整个栈进行协同设计:语言解码器仅通过门控交叉注意力访问视觉,因此模型在生成过程中可以自然地感知新输入帧;合成的交互语料用于监督何时说话、何时沉默、何时修正;分阶段课程将所有实时相关训练集中在一个轻量的最终阶段,基于强大的离线基础模型。在离线场景下,MOSS-VL-Instruct 在We present MOSS-VL, an open vision-language model family that treats real-time interaction -- perceiving while it speaks -- as a first-class capability. It is co-designed across the stack: the language decoder attends to vision only through gated cross-attention, so the model can naturally see incoming frames while generating; a synthesized interaction corpus supervises when to speak, when to stay silent, and when to revise; and a staged curriculum concentrates all real-time-specific training in one light final stage over a strong offline foundation. Offline, MOSS-VL-Instruct is competitive at

DeepSentiBank: Visual Sentiment Concept Classification with Deep Convolutional Neural Networks
DeepSentiBank: Visual Sentiment Concept Classification with Deep Convolutional Neural Networks
arXiv:1410.8586 多模态 方法 OA · 绿色 被引 316 · S2

性能评估显示,新训练的深度 CNN 模型 SentiBank 2.0(即 DeepSentiBank)在标注准确率与检索性能上较此前主要采用二分类 SVM 的版本有显著提升。Performance evaluation shows the newly trained deep CNNs model SentiBank 2.0 (or called DeepSentiBank) is significantly improved in both annotation accuracy and retrieval performance, compared to its predecessors which mainly use binary SVM classification models.

Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning
Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning
arXiv:2608.15869 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Internalized Visual Thinking (IVT),一种在无标签视频上联合优化文本预测与下一 embedding 预测的后训练框架,表明在推理时显式的像素级生成对有效的主动视频推理并非必要。Internalized Visual Thinking (IVT), a post-training framework that jointly optimizes textual prediction and next-embedding prediction over unlabeled videos, is introduced, suggesting that explicit pixel-space generation at inference time may not be necessary for effective proactive video reasoning.

EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing
EDITBRIDGE:面向忠实且高效的超高分辨率图像编辑
arXiv:2608.18063 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 EditBridge,一个用于高效超高分辨率编辑的扩散桥框架,可在最高 4K 分辨率下实现高保真编辑与卓越感知质量,在 2K 分辨率下带来 3.6–8.4 倍加速,并能在 61 秒内完成实用 4K 编辑。This work proposes EditBridge, a diffusion bridge framework for efficient ultra high-resolution editing that achieves high-fidelity editing with superior perceptual quality at resolutions up to 4K, delivering 3.6--8.4$\times$ speedup at 2K and enabling practical 4K editing in 61 seconds.

DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization
DiSCO:通过分布引导的对比提示优化保护文本到图像生成
arXiv:2608.17067 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 DiSCO,一种零样本、严格黑盒的防御方法,完全在提示层面以即插即用模块的形式运行,无需模型重训练、微调或访问模型内部,可直接应用于任何文本到图像系统,无需对模型本身进行任何修改。DiSCO is proposed, a zero-shot, strictly black-box defense that operates entirely at the prompt level as a plug-and-play module, requiring no model retraining, fine-tuning, or access to model internals, and can be readily applied to any text-to-image system without necessitating any changes to the model itself.

MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding
MoE-ViE:面向高效图像与视频理解的混合专家视觉编码器
arXiv:2608.17402 多模态 方法 OA · 绿色 被引 1 · S2

本文系统性地研究了视觉编码器扩展中的 MoE 设计,发现细粒度 MoE 拓扑相较于稠密与标准 MoE 基线均带来显著提升;提出了一种无辅助损失的均衡变体以改善专家利用率,并设计了专用 MoE kernel 以缓解推理时延开销。This work systematically study MoE designs for vision encoder scaling and finds that fine-grained MoE topologies yield substantial gains over both dense and standard MoE counterparts, and proposes an auxiliary-loss-free balancing variant for better expert utilization, and designs a specialized MoE kernel to mitigate inference latency overhead.

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist
OmniScientist:全模态跨学科 AI 科学家
arXiv:2608.13558 多模态 方法 OA · 绿色 被引 1 · S2

本文提出 OmniScientist,一种端到端、全模态 AI 科学家,可直接基于异构原始证据开展跨学科研究,表明全生命周期感知对于基于证据的科学发现至关重要,并为构建广泛适用的 AI 科学家提供了一条切实可行的路径OmniScientist is introduced, an end-to-end, omni-modal AI scientist that conducts multidisciplinary research directly from heterogeneous raw evidence and demonstrates that lifecycle-wide perception is essential for evidence-grounded scientific discovery and provides a practical path toward broadly capable AI scientists.

VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation
VA-Judger:基于人类偏好反馈的联合音视频生成奖励建模
arXiv:2608.18607 多模态 方法 OA · 绿色 被引 4 · S2

提出 VA-Judger,一种链式思考的通用奖励模型:从质量差距明显的样本对中学习以建立结构化输出和粗粒度偏好判别,再通过拒绝采样(以人类标注验证)蒸馏出可靠的偏好解释用于更难的质量相近样本比较,最后执行维度级强化学习,将人类反馈分解到各独立质量维度以获得更稠密的奖励信号。VA-Judger is proposed, a chain-of-thought omni-reward model that learns from pairs with clear quality gaps to establish structured output and coarse preference discrimination, then distills reliable preference explanations for harder near-quality comparisons via rejection sampling verified against human annotations, and finally performs dimension-wise reinforcement learning that decomposes human feedback into individual quality dimensions for denser reward signals.

Towards Real-Time and Adaptable LiDAR Scene Completion
迈向实时且自适应的 LiDAR 场景补全
arXiv:2608.16490 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 RapidLiDAR,一种将初始化本身视为可学习、数据驱动组件的 LiDAR 场景补全方法,在与 SOTA 相当的补全性能下,0.1 秒完成整个场景,比此前最快方法快 2.3 倍。RapidLiDAR is presented, a LiDAR scene completion method that treats the initialization itself as a learned, data-driven component and achieves completion performance on par with the state of the art while completing a full scene in 0.1 seconds, which is 2.3 times faster than the fastest prior method.

4DAnyone: Create Anyone in 4D from a Casual Monocular Video
4DAnyone:从随手单目视频创建 4D 人物
arXiv:2608.20335 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

4DAnyone 在新视角视频质量和下游 4DGS 重建上均优于先前方法,并具有稳健的野外泛化能力。4DAnyone outperforms prior methods in both novel-view video quality and downstream 4DGS reconstruction, with robust in-the-wild generalization.

τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation
τ_0-VLA:基于世界模型引导测试时计算的分层机器人基础模型
arXiv:2608.16885 多模态 方法 OA · 绿色 被引 3 · S2

在领域内和分布偏移设置下,增加测试时计算可显著提升下一子任务预测准确率,这些增益进一步转化为长时序机器人操作任务中更高的闭环成功率。Across in-domain and distribution-shifted settings, allocating additional test-time computation substantially improves next-subtask prediction accuracy, and these gains translate into higher closed-loop success on long-horizon robot manipulation tasks.

Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models
划分支撑集、重构残差:面向视频生成与世界模型的免训练稀疏注意力
arXiv:2608.18484 多模态 方法 OA · 绿色 被引 1 · S2

本文提出 SparsePR,一种无需训练的方法,将响应耦合划分(Response-Coupled Partitioning)与探针拟合残差重建(Probe-Fitted Residual Reconstruction)相结合,并发现划分方式既影响同一分组内查询偏好支持集之间的重叠程度,也影响稀疏输出的仿射函数对稠密与稀疏输出差异的拟合能力。This work introduces SparsePR, a training-free method combining Response-Coupled Partitioning with Probe-Fitted Residual Reconstruction, and finds that partition choice affects both the overlap among grouped queries' preferred supports and how well an affine function of the sparse output can represent the difference between dense and sparse outputs.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling
UniSpace:统一视觉表征与可扩展多模态建模
arXiv:2608.08676 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Patch Reparameterization,在保留原始语义通路的同时,添加一个面向重构的 patch embedding,为同一组冻结的 ViT 块提供细粒度视觉信息,在保持多模态理解能力的同时实现高保真图像重构,并取得有利的重构—生成权衡。Patch Reparameterization is introduced, which preserves the original semantic pathway while adding a reconstruction-aware patch embedding that provides fine-grained visual information to the same frozen ViT blocks, and preserves multimodal understanding while enabling high-fidelity image reconstruction and a favorable reconstruction--generation trade-off.

Block3D: Efficient Text-to-3D Generation via Block-Wise Diffusion
Block3D:通过块级扩散实现高效文本到 3D 生成
arXiv:2608.19567 多模态 方法 OA · 绿色 被引 1 · S2

本文提出 Block3D,一种块级扩散框架,将离散形状 token 序列划分为连续块,自回归地生成各块,并联合去噪当前块内的所有 token,同时引入置信度引导的块内修正机制,在每块定稿前对低置信度 token 进行修订。Block3D is proposed, a block-wise diffusion framework that partitions the discrete shape-token sequence into contiguous blocks, generates the blocks autoregressively, and jointly denoises all tokens within the current block and introduces confidence-guided intra-block correction, which revises low-confidence tokens before each block is finalized.

Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision
释放图像编辑潜力:概念缩放与密集监督
arXiv:2608.16812 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文建立了一个包含超过 1,000 个细粒度编辑概念的综合性层次化分类体系,并提出一种密集监督训练策略,将多个互不干扰的概念合成到单个图像对中,显著提升了训练效率和模型整体性能。A comprehensive hierarchical taxonomy featuring over 1,000 fine-grained edit concepts is established and a dense supervision training strategy that synthesizes multiple non-interfering concepts into single image pairs is proposed that significantly enhances both training efficiency and overall model performance.

EXPL-FR: Explaining Face Recognition Models via Vision-Language Alignment
EXPL-FR:通过视觉-语言对齐解释人脸识别模型
arXiv:2608.21486 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

覆盖 4 个 FR backbone 与 2 个 VLM 编码器;EXPL-FR 无需访问模型架构,支持身份级、单图及差异式解释,并在三种监督设置(人工标注、VLM 伪标签、完全 prompt 驱动的审计)下针对真实核验行为进行属性级审计基准测试。This work covers four FR backbones and two VLM encoders, EXPL-FR needs no architecture access, and supports identity-level, per-image, and differential explanations, and benchmark attribute-level auditing under three supervision settings, human labels, VLM pseudo-labels, and the authors' fully prompt-driven audit, against real verification behavior.

Length-Adaptive Decoding for Masked Diffusion Machine Translation
掩码扩散机器翻译的长度自适应解码
arXiv:2608.22274 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 Entropy-Valley (EV):一种无需训练的画布长度选择器,通过 all-mask 前向的预测平均熵对候选目标画布打分,并挑选出 backbone 最「准备好」填充的画布。Entropy-Valley (EV), a training-free length selector that scores candidate target canvases by mean predictive entropy from all-mask forward passes and selects the canvas the backbone is most prepared to fill, is introduced.

Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs
以标注作为 Rollout:面向视频 MLLMs 的高效可扩展强化学习
arXiv:2608.20492 多模态 方法 OA · 绿色 被引 1 · S2

本文研究视频 MLLM 的 RL 后训练样本效率与可扩展性,并提出 OraRL——一种随模型规模与数据规模共同 scaling 的解耦 advantage estimator,在 0.8B 到 9B backbone 上均超越其基线,并在 100k prompts 规模下超越 GRPO。The sample efficiency and scalability of RL post-training for video MLLMs and introduces OraRL, a decoupled advantage estimator that scales with model size and data, surpassing its backbone from 0.8B to 9B and GRPO up to 100k prompts.

MoTE: Mixture of Task Experts for Multi-Task Video Understanding
MoTE:面向多任务视频理解的 Task Expert 混合模型
arXiv:2608.24763 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 MoTE(Mixture of Task Experts),一种将大语言模型前馈网络转化为任务特定专家同时保持多模态 backbone 共享的 decoder 架构,并在五个 COIN 基准上使用显式任务路由进行评估。This work proposes MoTE (Mixture of Task Experts), a decoder architecture that converts large language model feed-forward networks into task-specific experts while keeping the multimodal backbone shared and evaluates it on five COIN benchmarks using explicit task routes.

GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture
GigaBrain-0.7:以三系统架构将具身基础模型扩展至涌现能力
arXiv:2608.15875 多模态 方法 OA · 绿色 被引 6 · S2

本文提出 GigaBrain-0.7,一种跨多种机器人 embodiment 泛化能力显著增强的 embodied foundation model,并引入一阶段对齐训练,联合优化 vision-language 理解和多 embodiment 动作生成。This work presents GigaBrain-0.7, an embodied foundation model with substantially improved generalization across diverse robot embodiments, and introduces one-stage alignment training that jointly optimizes vision-language understanding and multi-embodiment action generation.

Rubrics as Visual-Repair Context for Self-Evolving UI-to-Code Generation
以 Rubric 作为视觉修复上下文以实现自演化的 UI-to-Code 生成
arXiv:2608.24138 多模态 方法 OA · 绿色 被引 1 · S2

评估表明,RubSE 在终轮和最佳轮设置下均显著优于朴素 self-evolution,refinement 轨迹更稳定,且轨迹级性能上限更高。Evaluations demonstrate that RubSE substantially outperforms na\"ive self-evolution in final-round and best-round settings, achieving more stable refinement trajectories and a higher trajectory-level performance ceiling.

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds
面向持久化故事与交互式世界的长时音视频生成
arXiv:2608.23383 多模态 方法 OA · 绿色 被引 2 · S2

结果表明,记忆、几何控制以及 rollout-aware 训练为生成连贯故事和持续演化的交互式世界提供了实用基础。Results indicate that memory, geometric control, and rollout-aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds.

FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling
FIRM-Video:先核查再评分,实现可靠的 text-to-video 奖励建模
arXiv:2608.21839 多模态 方法 OA · 绿色 被引 1 · S2

本文提出 FHM-Video,一种基于 check-before-score 原则的、由 checklist 驱动的统一数据构建框架,在 FIRM-Video-Bench 上取得最佳综合 MAE,并在三种视频生成器的 Best-of-8 采样中始终获得最高的 VBench Total、Quality 和 Semantic Score。FHM-Video, a unified checklist-driven data construction framework based on a check-before-score principle, is introduced, which achieves the best overall MAE on FIRM-Video-Bench while consistently delivering the highest VBench Total, Quality, and Semantic Scores in Best-of-8 sampling across three video generators.

Real-TurnTurk: A Multimodal Turkish Corpus for Turn-Taking Prediction
Real-TurnTurk:用于话轮预测的多模态土耳其语语料库
arXiv:2608.22071 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文发布一个土耳其语多模态对话数据集,包含未脚本化的双人交互,并提供同步的前向视频、可归属到每位说话人的独立音频通道以及时间对齐的转写文本。A multimodal Turkish conversational dataset of unscripted dyadic interactions, comprising synchronized front-facing video, per-speaker audio channels that allow overlapping speech to be attributed to individual speakers, and time-aligned transcriptions is introduced.

Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models
Stream4D:面向流式自回归扩散视频模型的 4D 一致性
arXiv:2608.19556 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文用前馈 4D 重建奖励替代静态 critic,显式建模场景动态,使连贯运动获得高一致性奖励,并加入对自然 scene-flow 幅值进行奖励同时抑制抖动与非刚性伪影的运动先验。This work replaces the static critic with a feed-forward 4D reconstruction reward that explicitly models scene dynamics, allowing coherent motion to receive high consistency rewards, and adds a motion prior that rewards natural scene-flow magnitude while penalizing jitter and non-rigid artifacts.

LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale
LibriBrain100:面向大规模神经语音解码的百小时广深 MEG 数据集
arXiv:2608.25204 多模态 方法 OA · 绿色 被引 4 · S2

本文发布 LibriBrain100,一个面向语音解码的大规模 MEG 数据集,从设计上保证可复现、标准化评估,并展示了广泛多被试数据的价值:对预训练模型进行有监督微调可大幅弥补单被试数据不足。LibriBrain100 is introduced, a large-scale MEG dataset for speech decoding designed from the ground up for reproducible, standardised evaluation and the value of broad multi-subject data is demonstrated: supervised finetuning of a pre-trained model can substantially compensate for limited per-subject data.

Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning
Thinking on Shots:基于 Agentic 推理的一致性多镜头视频编辑。
arXiv:2608.26809 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一个 Agentic 视频编辑框架,利用 LLM 与 VLM 的协同实现 shot 级视频解耦与精确指令解析,并构建 MMLVE-Bench,一个聚焦 MMLVE 的数据集,具有复杂的真实世界时空动态、高密度异构指令以及稀疏随机的实体分布。This work introduces an agentic editing framework that leverages the synergy of Large Language Models (LLMs) and Vision-Language Models (VLMs) to achieve shot-level video decoupling and precise instruction parsing, and constructs MMLVE-Bench, which is an MMLVE-focused dataset characterized by complex real-world spatiotemporal dynamics, high-density heterogeneous instructions, and sparse, random entity distributions.

Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning
Aphanta:诊断任务对齐的图像编辑中间态以服务多模态推理。
arXiv:2608.26993 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果将图像编辑定位为一种专门的视觉工作空间而非通用推理机制,并将 Aphanta 确立为可复用的协议,用于度量任务-表征对齐、编辑器实现及下游 pipeline 实用性。The results position image editing as a specialized visual workspace rather than a universal reasoning mechanism, and establish Aphanta as a reusable protocol for measuring task--representation alignment, editor realization, and downstream pipeline utility.

GameWAM: A World Action Model for Video Games
GameWAM:面向视频游戏的世界动作模型
arXiv:2608.26200 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 GameWAM,据其所知是首个面向原生闭环游戏与 GUI 控制的 WAM,并发现 Low-Frequency Action Source Imprinting (LASI):在固定条件下,采样动作源的低频分量系统性地引导生成的粗粒度相机运动,揭示了生成式控制中的源敏感性失效模式。This work introduces GameWAM, to its knowledge the first WAM for native closed-loop gameplay and GUI control and uncovers Low-Frequency Action Source Imprinting (LASI), in which low-frequency components of the sampled action source systematically steer coarse generated camera motion under fixed conditioning, revealing a source-sensitivity failure mode in generative control.

EditaLive! Unified Character Video Editing for Live Streaming
EditaLive! 面向直播的统一人物视频编辑
arXiv:2608.27123 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

传统视频编辑主要关注场景级内容,而直播更强调人物主体。然而,直接将现有视频编辑方法应用于以人为中心的直播仍具挑战,因为它们可能引入面部表情不一致,且通常依赖多个离线推理步骤,难以满足实时交互需求。我们提出 EditaLive,一个用于实时流式角色视频编辑的新型框架。具体而言,我们从预训练图像动画模型(Wan-Animate)出发,该模型天然解耦...Conventional video editing primarily focuses on scene-level content, whereas live streaming places greater emphasis on the human subject. However, directly applying existing video-editing methods to human-centric live streaming remains challenging, as they may introduce facial-expression inconsistencies and typically depend on multiple offline inference steps, making them unsuitable for real-time interaction. We propose EditaLive, a novel framework for real-time streaming character video editing. In detail, we start from a pretrained image animation model (Wan-Animate), which naturally decoupl