WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing
- 类型:arxiv
- 标识:2609.20423
- 链接:https://arxiv.org/abs/2609.20423
- 主分类:llm-infra
- 形态:method
- TLDR:Document parsing converts document images into structured content and requires reliable performance across diverse layouts and acquisition conditions. Yet training corpora are biased toward common document types and clean digital pages, while expanding coverage alone does not specify how to address a parser's remaining weaknesses. We present WeVisDoc, a two-stage data-centric framework for robust end-to-end document parsing. Stage I broadens semantic, structural, and appearance coverage through heterogeneous data and structure-preserving degradation synthesis. Stage II uses a held-out probe to
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-09-18-agent-rag-longcontext-candidates.json