All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
- 类型:arxiv
- 标识:2609.24058
- 链接:https://arxiv.org/abs/2609.24058
- 主分类:multimodal
- 形态:application
- TLDR:Multilingual scene text recognition (STR) remains challenging due to the scarcity of training data for most languages and the difficulty of serving diverse scripts within a single model. Existing solutions either deploy one recognizer per language, inflating cost and introducing error accumulation, or rely on massive vision-language models (VLMs) that are expensive and still inaccurate on many scripts. In this work, we pursue an all-in-one multilingual recognizer that is simpler than per-language experts, lighter than VLMs, and more accurate than both. First, we construct TextMuSS-10M, a large
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-09-23-agent-rag-longcontext-candidates.json