Towards End-to-End Unified Scene Text Detection and Layout Analysis
Shangbang Long, Siyang Qin, Dmitry Panteleev, Alessandro Bissacco, Yasuhisa Fujii, Michalis Raptis
摘要
Scene text detection and document layout analysis have long been treated as two separate tasks in different image domains. In this paper, we bring them together and introduce the task of unified scene text detection and layout analysis. The first hierarchical scene text dataset is introduced to enable this novel research task. We also propose a novel method that is able to simultaneously detect scene text and form text clusters in a unified way. Comprehensive experiments show that our unified model achieves better performance than multiple well-designed baseline methods. Additionally, this model achieves state-of-the-art results on multiple scene text detection datasets without the need of complex post-processing. Dataset and code: https://github.com/google-research-datasets/hiertext.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper33
- SAM 3: Segment Anything with ConceptsNicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath 等ICLR 2026 · 被引用 1,103 次
- VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language TasksJiannan Wu, Muyan Zhong, Sen Xing, Zeqiang Lai 等NeurIPS 2024 · 被引用 179 次
- Detect Anything via Next Point PredictionQing Jiang, Junan Huo, Xingyu Chen, Yuda Xiong 等CVPR 2026 · 被引用 79 次
- Revisiting Scene Text Recognition: A Data PerspectiveQing Jiang, Jiapeng Wang, Dezhi Peng, Chongyu Liu 等ICCV 2023 · 被引用 70 次
- YOLOE: Real-Time Seeing AnythiAo Wang, Lihao Liu, Hui Chen, Zijia Lin 等ICCV 2025 · 被引用 52 次
它引用的顶会 Paper11
- Scene Text Visual Question AnsweringAli Furkan Biten, Rubèn Tito, Andrés Mafla, Lluís Gómez i Bigorda 等ICCV 2019 · 被引用 482 次
- DocFormer: End-to-End Transformer for Document UnderstandingSrikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie 等ICCV 2021 · 被引用 392 次
- Towards Unconstrained End-to-End Text SpottingSiyang Qin, Alessandro Bissacco, Michalis Raptis, Yasuhisa Fujii 等ICCV 2019 · 被引用 138 次
- Adaptive Boundary Proposal Network for Arbitrary Shape Text DetectionShi-Xue Zhang, Xiaobin Zhu, Chun Yang, Hongfa Wang 等ICCV 2021 · 被引用 112 次
- CentripetalText: An Efficient Text Instance Representation for Scene Text DetectionTao Sheng, Jie Chen, Zhouhui LianNeurIPS 2021 · 被引用 30 次
相关 Paper
- LayoutFormer: Hierarchical Text Detection Towards Scene Text UnderstandingMin Liang, Jia-Wei Ma, Xiaobin Zhu, Jingyan Qin 等CVPR 2024
- Text Grouping Adapter: Adapting Pre-Trained Text Detector for Layout AnalysisTianci Bi, Xiaoyi Zhang, Zhizheng Zhang, Wenxuan Xie 等CVPR 2024
- Towards Unified Multi-granularity Text Detection with Interactive AttentionXingyu Wan, Chengquan Zhang, Pengyuan Lyu, Sen Fan 等ICML 2024 · 被引用 4 次
- Vision-Language Pre-Training for Boosting Scene Text DetectorsSibo Song, Jianqiang Wan, Zhibo Yang, Jun Tang 等CVPR 2022 · 被引用 38 次
- Scene Text Retrieval via Joint Text Detection and Similarity LearningHao Wang, Xiang Bai, Mingkun Yang, Shenggao Zhu 等CVPR 2021
