Flow-Based Page Unique Semantic Mapping Architecture for Document Visual Question Answering
Haosen Wang, Jing Xiao, Chaochao Du, Xiaowang Zhang, Zhiyong Feng
Abstract
Document Visual Question Answering (DocVQA) aims to generate answers by jointly understanding the textual, layout, and visual elements within document images. Although end-to-end vision-based generative methods have reduced dependency on OCR, they still struggle to achieve precise evidence localization when page semantics are complex and highly similar. However, existing research lacks an in-depth theoretical analysis of the question-driven semantic representation space, failing to fundamentally address the distinguishability problem among semantically similar pages. To fill this theoretical gap, we propose and prove that, given a specific question, each page possesses a unique semantic representation, and there exists a bijective mapping between the page and its unique semantics. Based on this theoretical foundation, we introduce the Flow-Based Page Unique Semantic Mapping Architecture (FUMA), which reconstructs evidence localization from similarity-based retrieval into precise selection on unique semantics. FUMA employs fine-grained cross-modal attention to extract discriminative cues and utilizes flow-based reversible transformations with likelihood regularization to learn bijective mappings, ensuring that each page obtains a unique semantic representation. Moreover, a multi-expert collaboration mechanism complementarily models fine-grained multimodal information within each page, achieving robust answer generation. Experimental results demonstrate that FUMA significantly outperforms existing methods in both evidence localization and answer generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e83c0194-709f-448c-9df6-bbfb38adbd35Builds on19
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingYupan Huang, Tengchao Lv, Lei Cui, Yutong Lu et al.ACM MM 2022 · 606 citations
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang et al.KDD 2020 · 575 citations
- Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingKenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu et al.ICML 2023 · 426 citations
Related papers
- UniDocVLM: Enhancing Visual Reasoning and Document Understanding for VLM via Reinforcement LearningZongsheng Cao, Anran Liu, Jun Xie, Lang Chen et al.KDD 2026
- CREAM: Coarse-to-Fine Retrieval and Multi-modal Efficient Tuning for Document VQAJinxu Zhang, Yongqi Yu, Yu ZhangACM MM 2024 · 4 citations
- Doc-V^*: Coarse-to-Fine Interactive Visual Reasoning for Multi-Page Document VQAYuanlei Zheng, Pei Fu, Hang Li, Ziyang Wang et al.ACL 2026 · 2 citations
- DREAM: Integrating Hierarchical Multimodal Retrieval with Multi-page Multimodal Language Model for Documents VQAJinxu Zhang, Qiyuan Fan, Yongqi Yu, Yu ZhangACM MM 2025
- Attention as Selector: Unlocking VLM Attention for Long Document Page RetrievalMinfeng Zhu, Linxin Bao, Wei Chen, Linchao ZhuACL 2026
