MosaicOS: A Simple and Effective Use of Object-Centric Images for Long-Tailed Object Detection
Cheng Zhang, Tai-Yu Pan, Yandong Li, Hexiang Hu, Dong Xuan, Soravit Changpinyo, Boqing Gong, Wei-Lun Chao
Abstract
Many objects do not appear frequently enough in complex scenes (e.g., certain handbags in living rooms) for training an accurate object detector, but are often found frequently by themselves (e.g., in product images). Yet, these object-centric images are not effectively leveraged for improving object detection in scene-centric images. In this paper, we propose Mosaic of Object-centric images as Scene-centric images (MOSAICOS), a simple and novel framework that is surprisingly effective at tackling the challenges of long-tailed object detection. Keys to our approach are three-fold: (i) pseudo scene-centric image construction from object-centric images for mitigating domain differences, (ii) high-quality bounding box imputation using the object-centric images' class labels, and (iii) a multi-stage training procedure. On LVIS object detection (and instance segmentation), MOSAICOS leads to a massive 60% (and 23%) relative improvement in average precision for rare object categories. We also show that our framework can be compatibly used with other existing approaches to achieve even further gains. Our pre-trained models are publicly available at https:// github.com/czhang0528/MosaicOS/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fb423fbc-7bf8-4009-9755-1899c19df831Cited by top-tier papers14
- On Model Calibration for Long-Tailed Object Detection and Instance SegmentationTai-Yu Pan, Cheng Zhang, Yandong Li, Hexiang Hu et al.NeurIPS 2021 · 56 citations
- CLIM: Contrastive Language-Image Mosaic for Region RepresentationSize Wu, Wenwei Zhang, Lumin Xu, Sheng Jin et al.AAAI 2024 · 30 citations
- Learning from Rich Semantics and Coarse Locations for Long-tailed Object DetectionLingchen Meng, Xiyang Dai, Jianwei Yang, Dongdong Chen et al.NeurIPS 2023 · 23 citations
- Relieving Long-tailed Instance Segmentation via Pairwise Class BalanceYin-Yin He, Peizhen Zhang, Xiu-Shen Wei, Xiangyu Zhang et al.CVPR 2022 · 18 citations
- DiffuLT: Diffusion for Long-tail Recognition Without External KnowledgeJie Shao, Ke Zhu, Hanxiao Zhang, Jianxin WuNeurIPS 2024 · 17 citations
Builds on16
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan et al.ICLR 2020 · 1,496 citations
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 1,188 citations
- Objects365: A Large-Scale, High-Quality Dataset for Object DetectionShuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng et al.ICCV 2019 · 1,018 citations
- Balanced Meta-Softmax for Long-Tailed Visual RecognitionJiawei Ren, Cunjun Yu, Shunan Sheng, Xiao Ma et al.NeurIPS 2020 · 861 citations
- Frustratingly Simple Few-Shot Object DetectionXin Wang, Thomas E. Huang, Joseph Gonzalez, Trevor Darrell et al.ICML 2020 · 723 citations
Related papers
- Boosting Long-tailed Object Detection via Step-wise Learning on Smooth-tail DataNa Dong, Yongqiang Zhang, Mingli Ding, Gim Hee LeeICCV 2023 · 8 citations
- SimLTD: Simple Supervised and Semi-Supervised Long-Tailed Object DetectionPhi Vu TranCVPR 2025
- Overcoming Classifier Imbalance for Long-Tail Object Detection With Balanced Group SoftmaxYu Li, Tao Wang, Bingyi Kang, Sheng Tang et al.CVPR 2020
- Adaptive Hierarchical Representation Learning for Long-Tailed Object DetectionBanghuai LiCVPR 2022 · 16 citations
- DropLoss for Long-Tail Instance SegmentationTing-I Hsieh, Esther Robb, Hwann-Tzong Chen, Jia-Bin HuangAAAI 2021 · 53 citations
