CAT: LoCalization and IdentificAtion Cascade Detection Transformer for Open-World Object Detection
Shuailei Ma, Yuefeng Wang, Ying Wei, Jiaqi Fan, Thomas H. Li, Hongli Liu, Fanbing Lv
Abstract
Open-world object detection (OWOD), as a more general and challenging goal, requires the model trained from data on known objects to detect both known and unknown objects and incrementally learn to identify these unknown objects. The existing works which employ standard detection framework and fixed pseudo-labelling mechanism (PLM) have the following problems: (𝑖) The inclusion of detecting unknown objects substantially reduces the model's ability to detect known ones. (𝑖𝑖) The PLM does not adequately utilize the priori knowledge of inputs. (𝑖𝑖𝑖) The fixed selection manner of PLM cannot guarantee that the model is trained in the right direction. We observe that humans subconsciously prefer to focus on all foreground objects and then identify each one in detail, rather than localize and identify a single object simultaneously, for alleviating the confusion. This motivates us to propose a novel solution called CAT: LoCalization and IdentificAtion Cascade Detection Transformer which decouples the detection process via the shared decoder in the cascade decoding way. In the meanwhile, we propose the self-adaptive pseudo-labelling mechanism which combines the modeldriven with input-driven PLM and self-adaptively generates robust pseudo-labels for unknown objects, significantly improving the ability of CAT to retrieve unknown objects. Experiments on two benchmarks, 𝑖.𝑒., MS-COCO and PAS-CAL VOC, show that our model outperforms the state-ofthe-art methods. The code is publicly available at https: //github.com/xiaomabufei/CAT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers18
- Learning Vision-and-Language Navigation from YouTube VideosKunyang Lin, Peihao Chen, Diwei Huang, Thomas H. Li et al.ICCV 2023 · 57 citations
- UMB: Understanding Model Behavior for Open-World Object DetectionXing Xi, Yangyang Huang, Zhijie Zhong, Ronghua LuoNeurIPS 2024 · 10 citations
- 3D Indoor Instance Segmentation in an Open-WorldMohamed El Amine Boudjoghra, Salwa K. Al Khatib, Jean Lahoud, Hisham Cholakkal et al.NeurIPS 2023 · 9 citations
- Looking Beyond the Known: Towards a Data Discovery Guided Open-World Object DetectionAnay Majee, Amitesh Gangrade, Rishabh IyerNeurIPS 2025 · 5 citations
- Knowing the Unknown: Interpretable Open-World Object Detection via Concept Decomposition ModelXueqiang Lv, Shizhou Zhang, Yinghui Xing, di xu et al.ICML 2026 · 2 citations
Builds on9
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Few-Shot Object Detection via Feature ReweightingBingyi Kang, Zhuang Liu, Xin Wang, Fisher Yu et al.ICCV 2019 · 835 citations
- An End-to-End Transformer Model for 3D Object DetectionIshan Misra, Rohit Girdhar, Armand JoulinICCV 2021 · 602 citations
- Dynamic DETR: End-to-End Object Detection with Dynamic AttentionXiyang Dai, Yinpeng Chen, Jianwei Yang, Pengchuan Zhang et al.ICCV 2021 · 429 citations
Related papers
- OW-DETR: Open-world Detection TransformerAkshita Gupta, Sanath Narayan, K. J. Joseph, Salman Khan et al.CVPR 2022 · 209 citations
- PROB: Probabilistic Objectness for Open World Object DetectionOrr Zohar, Kuan-Chieh Wang, Serena YeungCVPR 2023
- Semi-supervised Open-World Object DetectionSahal Shaji Mullappilly, Abhishek Singh Gehlot, Rao Muhammad Anwer, Fahad Shahbaz Khan et al.AAAI 2024 · 19 citations
- OW-Adapter: Human-Assisted Open-World Object Detection with a Few ExamplesSuphanut Jamonnak, Jiajing Guo, Wenbin He, Liang Gou et al.IEEE VIS 2023 · 6 citations
- Annealing-based Label-Transfer Learning for Open World Object DetectionYuqing Ma, Hainan Li, Zhange Zhang, Jinyang Guo et al.CVPR 2023
