Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation
- 类型:arxiv
- 标识:2606.30054
- 链接:https://arxiv.org/abs/2606.30054
- 主分类:multimodal
- 形态:method
- 被引:0
- 被引来源:Semantic Scholar + OpenAlex
- S2被引:0
- OpenAlex被引:0
- 影响力被引:0
- TLDR:This paper introduces ILLUME-X, an advanced unified multimodal paradigm that enables high-quality, free-form interleaved text-image generation by improving multimodal data efficiency and stabilizing the multimodal training process.
- OpenAlex ID:W7166664512
- OpenAlex DOI:10.48550/arxiv.2606.30054
- DOI:10.48550/arxiv.2606.30054
- DOI来源:OpenAlex
- 开放获取:green
- 开放获取链接:https://doi.org/10.48550/arxiv.2606.30054
- OpenAlex更新:2026-07-19
- 待LLM分类:否
- 标题中文:照亮统一多模态模型:面向自由形式交错图文生成
- TLDR中文:本文提出 ILLUME-X,一种先进的统一多模态范式,通过提升多模态数据效率并稳定多模态训练过程,实现高质量、自由形式的交错图文生成。
- 来源文件:
- /inbox/tom/_candidates/2026-06-30-agent-rag-longcontext-candidates.json
- [S2 enrich]
- [OpenAlex backfill]