GLOMA: Global Video Text Spotting with Morphological Association
Han Wang, Yanjie Wang, Yang Li, Can Huang
Abstract
Video Text Spotting (VTS) is a fundamental visual task that aims to predict the trajectories and content of texts in a video. Previous works usually conduct local associations and apply IoU-based distance and complex post-processing procedures to boost performance, ignoring the abundant temporal information and the morphological characteristics in VTS. In this paper, we propose to model the tracking problem as global associations and utilize the Gaussian Wasserstein distance to guide the morphological correlation between frames. Our main contributions can be summarized as three folds. 1). We propose a Transformer-based global tracking method for VTS and associate multiple frames simultaneously. 2). We introduce a Wasserstein distance-based method to conduct positional associations between frames. 3). We conduct extensive experiments on public datasets. On the ICDAR2015 video dataset, achieves 56.0 MOTA with 4.6 absolute improvement compared with the previous SOTA method and outperforms the previous Transformer-based method by a significant 8.3 MOTA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1ca4effb-9331-487a-b80b-5415490e030bCited by top-tier papers1
Ask how each one uses itBuilds on8
- Rethinking Rotated Object Detection with Gaussian Wasserstein Distance LossXue Yang, Junchi Yan, Qi Ming, Wentao Wang et al.ICML 2021 · 572 citations
- Efficient and Accurate Arbitrary-Shaped Text Detection With Pixel Aggregation NetworkWenhai Wang, Enze Xie, Xiaoge Song, Yuhang Zang et al.ICCV 2019 · 490 citations
- Sequence Level Semantics Aggregation for Video Object DetectionHaiping Wu, Yuntao Chen, Naiyan Wang, Zhaoxiang ZhangICCV 2019 · 236 citations
- Relation Distillation Networks for Video Object DetectionJiajun Deng, Yingwei Pan, Ting Yao, Wengang Zhou et al.ICCV 2019 · 211 citations
- Global Tracking TransformersXingyi Zhou, Tianwei Yin, Vladlen Koltun, Philipp KrähenbühlCVPR 2022 · 180 citations
Related papers
- GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term MatchingHaibin He, Maoyuan Ye, Jing Zhang, Juhua Liu et al.NeurIPS 2024 · 16 citations
- Embracing Consistency: A One-Stage Approach for Spatio-Temporal Video GroundingYang Jin, Yongzhi Li, Zehuan Yuan, Yadong MuNeurIPS 2022 · 69 citations
- End-to-End Referring Video Object Segmentation with Multimodal TransformersAdam Botach, Evgenii Zheltonozhskii, Chaim BaskinCVPR 2022 · 150 citations
- Semantic-Aware Video Text DetectionWei Feng, Fei Yin, Xu-Yao Zhang, Cheng-Lin LiuCVPR 2021
- TubeDETR: Spatio-Temporal Video Grounding with TransformersAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev et al.CVPR 2022 · 87 citations
