Mask is All You Need: Rethinking Mask R-CNN for Dense and Arbitrary-Shaped Scene Text Detection
Xugong Qin, Yu Zhou, Youhui Guo, Dayan Wu, Zhihong Tian, Ning Jiang, Hongbin Wang, Weiping Wang
摘要
Due to the large success in object detection and instance segmentation, Mask R-CNN attracts great attention and is widely adopted as a strong baseline for arbitrary-shaped scene text detection and spotting. However, two issues remain to be settled. The first is dense text case, which is easy to be neglected but quite practical. There may exist multiple instances in one proposal, which makes it difficult for the mask head to distinguish different instances and degrades the performance. In this work, we argue that the performance degradation results from the learning confusion issue in the mask head. We propose to use an MLP decoder instead of the "deconv-conv" decoder in the mask head, which alleviates the issue and promotes robustness significantly. And we propose instanceaware mask learning in which the mask head learns to predict the shape of the whole instance rather than classify each pixel to text or non-text. With instance-aware mask learning, the mask branch can learn separated and compact masks. The second is that due to large variations in scale and aspect ratio, RPN needs complicated anchor settings, making it hard to maintain and transfer across different datasets. To settle this issue, we propose an adaptive label assignment in which all instances especially those with extreme aspect ratios are guaranteed to be associated with enough anchors. Equipped with these components, the proposed method named MAYOR 1 achieves state-of-the-art performance on five benchmarks including DAST1500, MSRA-TD500, ICDAR2015, CTW1500, and Total-Text.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Exploring Stroke-Level Modifications for Scene Text EditingYadong Qu, Qingfeng Tan, Hongtao Xie, Jianjun Xu 等AAAI 2023 · 被引用 51 次
- TPSNet: Reverse Thinking of Thin Plate Splines for Arbitrary Shape Scene Text RepresentationWei Wang, Yu Zhou, Jiahao Lyu, Dayan Wu 等ACM MM 2022 · 被引用 35 次
- Towards Robust Real-Time Scene Text Detection: From Semantic to Instance Representation LearningXugong Qin, Pengyuan Lyu, Chengquan Zhang, Yu Zhou 等ACM MM 2023 · 被引用 21 次
- Focus, Distinguish, and Prompt: Unleashing CLIP for Efficient and Flexible Scene Text RetrievalGangyan Zeng, Yuan Zhang, Jin Wei, Dongbao Yang 等ACM MM 2024 · 被引用 8 次
- PBFormer: Capturing Complex Scene Text Shape with Polynomial Band TransformerRuijin Liu, Ning Lu, Dapeng Chen, Cheng Li 等ACM MM 2023 · 被引用 2 次
它引用的顶会 Paper14
- Real-Time Scene Text Detection with Differentiable BinarizationMinghui Liao, Zhaoyi Wan, Cong Yao, Kai Chen 等AAAI 2020 · 被引用 818 次
- Efficient and Accurate Arbitrary-Shaped Text Detection With Pixel Aggregation NetworkWenhai Wang, Enze Xie, Xiaoge Song, Yuhang Zang 等ICCV 2019 · 被引用 490 次
- TextDragon: An End-to-End Framework for Arbitrary Shaped Text SpottingWei Feng, Wenhao He, Fei Yin, Xu-Yao Zhang 等ICCV 2019 · 被引用 212 次
- Video Cloze Procedure for Self-Supervised Spatio-Temporal LearningDezhao Luo, Chang Liu, Yu Zhou, Dongbao Yang 等AAAI 2020 · 被引用 167 次
- Towards Unconstrained End-to-End Text SpottingSiyang Qin, Alessandro Bissacco, Michalis Raptis, Yasuhisa Fujii 等ICCV 2019 · 被引用 138 次
相关 Paper
- MANGO: A Mask Attention Guided One-Stage Scene Text SpotterLiang Qiao, Ying Chen, Zhanzhan Cheng, Yunlu Xu 等AAAI 2021 · 被引用 91 次
- CRNet: A Center-aware Representation for Detecting Text of Arbitrary ShapesYu Zhou, Hongtao Xie, Shancheng Fang, Yan Li 等ACM MM 2020 · 被引用 31 次
- CentripetalText: An Efficient Text Instance Representation for Scene Text DetectionTao Sheng, Jie Chen, Zhouhui LianNeurIPS 2021 · 被引用 30 次
- ContourNet: Taking a Further Step Toward Accurate Arbitrary-Shaped Scene Text DetectionYuxin Wang, Hongtao Xie, Zheng-Jun Zha, Mengting Xing 等CVPR 2020
- SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text RecognitionMingxin Huang, Yuliang Liu, Zhenghao Peng, Chongyu Liu 等CVPR 2022 · 被引用 150 次
