LayoutFormer: Hierarchical Text Detection Towards Scene Text Understanding
Min Liang, Jia-Wei Ma, Xiaobin Zhu, Jingyan Qin, Xu-Cheng Yin
Abstract
Existing scene text detectors generally focus on accurately detecting single-level (i.e., word-level, line-level, or paragraph-level) text entities without exploring the relationships among different levels of text entities. To comprehensively understand scene texts, detecting multi-level texts while exploring their contextual information is critical. To this end, we propose a unified framework (dubbed LayoutFormer) for hierarchical text detection, which simultaneously conducts multi-level text detection and predicts the geometric layouts for promoting scene text understanding. In LayoutFormer, WordDecoder, LineDecoder, and Pa-raDecoder are proposed to be responsible for word-level text prediction, line-level text prediction, and paragraphlevel text prediction, respectively. Meanwhile, WordDecoder and ParaDecoder adaptively learn word-line and line-paragraph relationships, respectively. In addition, we propose a Prior Location Sampler to be used on multi-scale features to adaptively select a few representative foreground features for updating text queries. It can improve hierarchical detection performance while significantly reducing the computational cost. Comprehensive experiments verify that our method achieves state-of-the-art performance on single-level and hierarchical text detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8081934e-5fc0-4b35-aabe-30951cb6188cCited by top-tier papers2
- When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and UnderstandingYan Shu, Hangui Lin, Yexin Liu, Yan Zhang et al.NeurIPS 2025 · 17 citations
- Explicit Relational Reasoning Network for Scene Text DetectionYuchen Su, Zhineng Chen, Yongkun Du, Zhilong Ji et al.AAAI 2025
Builds on20
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Real-Time Scene Text Detection with Differentiable BinarizationMinghui Liao, Zhaoyi Wan, Cong Yao, Kai Chen et al.AAAI 2020 · 818 citations
- Efficient and Accurate Arbitrary-Shaped Text Detection With Pixel Aggregation NetworkWenhai Wang, Enze Xie, Xiaoge Song, Yuhang Zang et al.ICCV 2019 · 490 citations
- Scene Text Visual Question AnsweringAli Furkan Biten, Rubèn Tito, Andrés Mafla, Lluís Gómez i Bigorda et al.ICCV 2019 · 482 citations
- Rethinking Transformer-based Set Prediction for Object DetectionZhiqing Sun, Shengcao Cao, Yiming Yang, Kris KitaniICCV 2021 · 381 citations
Related papers
- Towards End-to-End Unified Scene Text Detection and Layout AnalysisShangbang Long, Siyang Qin, Dmitry Panteleev, Alessandro Bissacco et al.CVPR 2022 · 86 citations
- Text Grouping Adapter: Adapting Pre-Trained Text Detector for Layout AnalysisTianci Bi, Xiaoyi Zhang, Zhizheng Zhang, Wenxuan Xie et al.CVPR 2024
- PBFormer: Capturing Complex Scene Text Shape with Polynomial Band TransformerRuijin Liu, Ning Lu, Dapeng Chen, Cheng Li et al.ACM MM 2023 · 2 citations
- Few Could Be Better Than All: Feature Sampling and Grouping for Scene Text DetectionJingqun Tang, Wenqing Zhang, Hongye Liu, Mingkun Yang et al.CVPR 2022 · 103 citations
- Towards Unified Multi-granularity Text Detection with Interactive AttentionXingyu Wan, Chengquan Zhang, Pengyuan Lyu, Sen Fan et al.ICML 2024 · 4 citations
