Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision
- 类型:arxiv
- 标识:2608.16812
- 链接:https://arxiv.org/abs/2608.16812
- 主分类:multimodal
- 形态:method
- TLDR:Existing image editing frameworks predominantly follow the training paradigm of text-to-image diffusion models. However, extending this paradigm to image editing highlights two inherent discrepancies, specifically, the insufficient attention to edit concept granularity and the training inefficiency caused by sparse supervision signals. To address these issues, we establish a comprehensive hierarchical taxonomy featuring over 1,000 fine-grained edit concepts and build ConceptEdit-12M, a massive dataset of 12 million high-quality editing pairs via an improved synthesis framework. This library-dr
- 副分类:engineering
- 待LLM分类:否
- 标题中文:释放图像编辑潜力:概念缩放与密集监督
- TLDR中文:现有图像编辑框架主要沿用 text-to-image 扩散模型的训练范式。然而将该范式扩展到图像编辑时暴露出两个固有差异:一是对编辑概念粒度的关注不足,二是稀疏监督信号导致的训练低效。为应对这些问题,我们建立了一个包含超过 1,000 种细粒度编辑概念的综合分层分类体系,并构建了 ConceptEdit-12M——通过改进的合成框架生成的包含 1,200 万高质量编辑对的超大规模数据集。该 library-dr
- 来源文件:
- /inbox/tom/_candidates/2026-08-25-agent-rag-longcontext-candidates.json