Papers · organized/paper_cards

论文

258 张论文卡片 · 多模态

开放获取 全部 绿色 · 724
Self-Supervised Visual On-Policy Distillation
自监督视觉 On-Policy Distillation
arXiv:2608.14144 多模态 方法 被引 0 · S2

提出 Self-Supervised Visual On-Policy Distillation(S²VOPD),一种简单有效的方法,通过非对称增强视图构建 on-policy 学习信号,系统地探索了视觉增强的广阔设计空间,并发现非对称性至关重要。Self-Supervised Visual On-Policy Distillation (S$^2$VOPD), a simple yet effective method that constructs on-policy learning signals from asymmetric augmented views, systematically explores a broad design space of visual augmentations and uncover that asymmetry matters.

UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations
UniProbe:一种利用多结构内部表示的可学习 token 级 VLM 幻觉检测器
arXiv:2608.10835 多模态 方法 被引 0 · S2

提出 UniProbe,一个轻量、统一、可学习的检测器,通过单次前向传播建模冻结 LVLM 的异构计算 trace,在 token 级和物体级幻觉检测上达到 SOTA。UniProbe is introduced, a lightweight, unified, learnable detector that models a frozen LVLM's heterogeneous computational trace from a single forward pass and achieves state-of-the-art token-level and object-hallucination detection.

AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model
AnyTalk:基于视频生成模型的任意角色语音动画
arXiv:2608.16143 多模态 方法 被引 0 · S2

AnyTalk 能够在多样化的人脸网格和 blendshape 配置下生成唇形同步动画,显著减少人工工作和数据需求;并通过将 AnyTalk 蒸馏为精简网络 $\text{AnyTalk}_{RT}$ 来提升可用性,从而实现实时性能。AnyTalk enables lip-synced animations across diverse face meshes and blendshape configurations, significantly reducing manual effort and data requirements and enhances usability by distilling AnyTalk into a streamlined network, $\text{AnyTalk}_{RT}$, thereby enabling real-time performance.

GRNEdit: Efficient General Video Editing from a New Binary-Evidence Perspective in Generative Refinement Networks
GRNEdit:基于生成式 Refinement 网络的新二元证据视角下的高效通用视频编辑
arXiv:2608.16328 多模态 方法 被引 0 · S2

GRNEdit 是一个轻量级的两阶段指令驱动通用视频编辑框架,性能优于多个 14B 开源编辑器,同时其 8B 模型与领先的开源编辑器表现相当。GRNEdit, a lightweight two-stage framework for instruction-based general video editing that outperforms multiple 14B open-source editors, while the 8B model performs on par with leading open-source editors.

TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation
TRACE-Bench:多参考图像生成的分解与诊断
arXiv:2608.16765 多模态 评测集 被引 0 · S2

认识到多样化的多参考任务共享一组共同的原子操作,本文形式化了四个算子:Anchor、Disentangle、Apply 和 Compose,并构建了 TRACE-Bench,包含约 1,600 个跨 slot 数量 1–8 的评估用例。Recognizing that diverse multi-reference tasks share a common set of atomic operations, this work formalizes four operators: Anchor, Disentangle, Disentangle, Apply, and Compose, and constructs TRACE-Bench, comprising approximately 1,600 evaluation cases across slot counts 1--8.

WorldRover: A Scalable Synthetic Video Data Engine for World Exploration with Rich Annotations
WorldRover:一个用于世界探索的、可扩展的、带有丰富标注的合成视频数据引擎
arXiv:2608.15659 多模态 方法 被引 0 · S2

WorldRover 将长视野世界探索转化为可扩展的数据生成问题,为需要在可探索世界中构建、维护并重访一致表征的模型提供监督信号。WorldRover turns long-horizon world exploration into a scalable data-generation problem, providing supervision for models that must build, maintain, and revisit coherent representations of an explorable world.