Learning from Rich Semantics and Coarse Locations for Long-tailed Object Detection
Lingchen Meng, Xiyang Dai, Jianwei Yang, Dongdong Chen, Yinpeng Chen, Mengchen Liu, Yi-Ling Chen, Zuxuan Wu, Lu Yuan, Yu-Gang Jiang
Abstract
Long-tailed object detection (LTOD) aims to handle the extreme data imbalance in real-world datasets, where many tail classes have scarce instances. One popular strategy is to explore extra data with image-level labels, yet it produces limited results due to (1) semantic ambiguity -- an image-level label only captures a salient part of the image, ignoring the remaining rich semantics within the image; and (2) location sensitivity -- the label highly depends on the locations and crops of the original image, which may change after data transformations like random cropping. To remedy this, we propose RichSem, a simple but effective method, which is robust to learn rich semantics from coarse locations without the need of accurate bounding boxes. RichSem leverages rich semantics from images, which are then served as additional soft supervision for training detectors. Specifically, we add a semantic branch to our detector to learn these soft semantics and enhance feature representations for long-tailed object detection. The semantic branch is only used for training and is removed during inference. RichSem achieves consistent improvements on both overall and rare-category of LVIS under different backbones and detectors. Our method achieves state-of-the-art performance without requiring complex training and testing procedures. Moreover, we show the effectiveness of our method on other long-tailed datasets with additional experiments. Code is available at https://github.com/MengLcool/RichSem.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b1943992-5b0d-463f-88d4-0c33b2b22707Cited by top-tier papers9
- Improved Balanced Classification with Theoretically Grounded Loss FunctionsCorinna Cortes, Mehryar Mohri, Yutao ZhongNeurIPS 2025 · 19 citations
- FOCUS: Towards Universal Foreground SegmentationZuyao You, Lingyu Kong, Lingchen Meng, Zuxuan WuAAAI 2025 · 9 citations
- Optimized Deferral for Imbalanced SettingsCorinna Cortes, Anqi Mao, Mehryar Mohri, Yutao ZhongICML 2026 · 7 citations
- SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object DetectionHao Vo, Khoa Vo, Thinh Phan, Ngo Xuan Cuong et al.CVPR 2026 · 1 citation
- CQ-DINO: Mitigating Gradient Dilution via Category Queries for Vast Vocabulary Object DetectionZhichao Sun, Huazhang Hu, Yidong Ma, Gang Liu et al.NeurIPS 2025
Builds on37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
Related papers
- SimLTD: Simple Supervised and Semi-Supervised Long-Tailed Object DetectionPhi Vu TranCVPR 2025
- Boosting Long-tailed Object Detection via Step-wise Learning on Smooth-tail DataNa Dong, Yongqiang Zhang, Mingli Ding, Gim Hee LeeICCV 2023 · 8 citations
- Adaptive Hierarchical Representation Learning for Long-Tailed Object DetectionBanghuai LiCVPR 2022 · 16 citations
- MosaicOS: A Simple and Effective Use of Object-Centric Images for Long-Tailed Object DetectionCheng Zhang, Tai-Yu Pan, Yandong Li, Hexiang Hu et al.ICCV 2021 · 50 citations
- Overcoming Classifier Imbalance for Long-Tail Object Detection With Balanced Group SoftmaxYu Li, Tao Wang, Bingyi Kang, Sheng Tang et al.CVPR 2020
