M3Grounder: Mask-Based Multi-Span and Multi-Granular Grounding for Document QA
Venkata Kesav Venna, Sai Madhusudan Gunda, Jyothi Swaroopa Jinka, Hrithik Sagar Rachakonda, Anirudh Srinivasan, Ravi Kiran Sarvadevabhatla
Abstract
Document QA requires not only accurate answers but also identifying where each answer is grounded on the page. Most approaches treat the task as textonly generation, while existing answer grounding methods generate coarse bounding boxes that fail to capture curved text. We introduce M3Grounder, a hybrid vision-language and segmentation architecture that formulates document grounding as pixel-level segmentation. It produces fine-grained evidence masks refined by a bleed-suppression loss to prevent spillover.
M3Grounder autoregressively generates answer text interleaved with [GROUND] tokens that link individual answer spans to their corresponding evidence regions. Also, M3Grounder grounds evidence hierarchically across phrase, line, and block levels using an enclosure loss that enforces spatial containment. We release Ground-ingDocQA dataset (200K documents, 2M multi-span and multi-granular QA pairs with pixel-level grounding masks), built through a data engine that handles complex layouts, curved-text and graphic-rich documents. We also release GroundingDocQA-Bench, a diverse and challenging human-verified benchmark. M3Grounder sets a new state-of-the-art in grounded DocVQA, advancing from coarse boxes to hierarchical, fine-grained and contextually grounded mask evidence. Code, dataset and models will be released.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2fff6863-6703-4b67-8de7-bcc23b0277c3Builds on15
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Grounding Multimodal Large Language Models to the WorldZhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao et al.ICLR 2024 · 1,170 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang et al.ICLR 2020 · 674 citations
Related papers
- UniDocVLM: Enhancing Visual Reasoning and Document Understanding for VLM via Reinforcement LearningZongsheng Cao, Anran Liu, Jun Xie, Lang Chen et al.KDD 2026
- VideoGLaMM : A Large Multimodal Model for Pixel-Level Visual Grounding in VideosShehan Munasinghe, Hanan Gani, Wenqi Zhu, Jiale Cao et al.CVPR 2025
- MGDoc: Pre-training with Multi-granular Hierarchy for Document Image UnderstandingZilong Wang, Jiuxiang Gu, Chris Tensmeyer, Nikolaos Barmpalios et al.EMNLP 2022 · 6 citations
- Groundhog Grounding Large Language Models to Holistic SegmentationYichi Zhang, Ziqiao Ma, Xiaofeng Gao, Suhaila Shakiah et al.CVPR 2024 · 24 citations
- Grounded 3D-Aware Spatial Vision-Language ModelingAn-Chieh Cheng, Yang Fu, Yatai Ji, Ligeng Zhu et al.CVPR 2026
