NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference

  • 类型:arxiv
  • 标识:2609.01657
  • 链接:https://arxiv.org/abs/2609.01657
  • 主分类:multimodal
  • 形态:method
  • TLDR:Multimodal models often build on architectures designed for generative vision-language modeling, typically combining separately pretrained vision encoders with causal language models. Visual document retrievers such as ColPali repurpose these models as encoders, carrying over the parameter and compute overhead of a VLM for a non-generative task. We introduce NeoMME, a family of 260M and 800M-parameter Multimodal and Multilingual bidirectional Encoders that process multilingual text and raw image patches in a single bidirectional Transformer encoder. Both models are pretrained from scratch with
  • 副分类:engineering
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-03-agent-rag-longcontext-candidates.json