CAT: LoCalization and IdentificAtion Cascade Detection Transformer for Open-World Object Detection
Shuailei Ma, Yuefeng Wang, Ying Wei, Jiaqi Fan, Thomas H. Li, Hongli Liu, Fanbing Lv
摘要
Open-world object detection (OWOD), as a more general and challenging goal, requires the model trained from data on known objects to detect both known and unknown objects and incrementally learn to identify these unknown objects. The existing works which employ standard detection framework and fixed pseudo-labelling mechanism (PLM) have the following problems: (𝑖) The inclusion of detecting unknown objects substantially reduces the model's ability to detect known ones. (𝑖𝑖) The PLM does not adequately utilize the priori knowledge of inputs. (𝑖𝑖𝑖) The fixed selection manner of PLM cannot guarantee that the model is trained in the right direction. We observe that humans subconsciously prefer to focus on all foreground objects and then identify each one in detail, rather than localize and identify a single object simultaneously, for alleviating the confusion. This motivates us to propose a novel solution called CAT: LoCalization and IdentificAtion Cascade Detection Transformer which decouples the detection process via the shared decoder in the cascade decoding way. In the meanwhile, we propose the self-adaptive pseudo-labelling mechanism which combines the modeldriven with input-driven PLM and self-adaptively generates robust pseudo-labels for unknown objects, significantly improving the ability of CAT to retrieve unknown objects. Experiments on two benchmarks, 𝑖.𝑒., MS-COCO and PAS-CAL VOC, show that our model outperforms the state-ofthe-art methods. The code is publicly available at https: //github.com/xiaomabufei/CAT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Learning Vision-and-Language Navigation from YouTube VideosKunyang Lin, Peihao Chen, Diwei Huang, Thomas H. Li 等ICCV 2023 · 被引用 57 次
- UMB: Understanding Model Behavior for Open-World Object DetectionXing Xi, Yangyang Huang, Zhijie Zhong, Ronghua LuoNeurIPS 2024 · 被引用 10 次
- 3D Indoor Instance Segmentation in an Open-WorldMohamed El Amine Boudjoghra, Salwa K. Al Khatib, Jean Lahoud, Hisham Cholakkal 等NeurIPS 2023 · 被引用 9 次
- Looking Beyond the Known: Towards a Data Discovery Guided Open-World Object DetectionAnay Majee, Amitesh Gangrade, Rishabh IyerNeurIPS 2025 · 被引用 5 次
- Knowing the Unknown: Interpretable Open-World Object Detection via Concept Decomposition ModelXueqiang Lv, Shizhou Zhang, Yinghui Xing, di xu 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper9
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Few-Shot Object Detection via Feature ReweightingBingyi Kang, Zhuang Liu, Xin Wang, Fisher Yu 等ICCV 2019 · 被引用 835 次
- An End-to-End Transformer Model for 3D Object DetectionIshan Misra, Rohit Girdhar, Armand JoulinICCV 2021 · 被引用 602 次
- Dynamic DETR: End-to-End Object Detection with Dynamic AttentionXiyang Dai, Yinpeng Chen, Jianwei Yang, Pengchuan Zhang 等ICCV 2021 · 被引用 429 次
相关 Paper
- OW-DETR: Open-world Detection TransformerAkshita Gupta, Sanath Narayan, K. J. Joseph, Salman Khan 等CVPR 2022 · 被引用 209 次
- PROB: Probabilistic Objectness for Open World Object DetectionOrr Zohar, Kuan-Chieh Wang, Serena YeungCVPR 2023
- Semi-supervised Open-World Object DetectionSahal Shaji Mullappilly, Abhishek Singh Gehlot, Rao Muhammad Anwer, Fahad Shahbaz Khan 等AAAI 2024 · 被引用 19 次
- OW-Adapter: Human-Assisted Open-World Object Detection with a Few ExamplesSuphanut Jamonnak, Jiajing Guo, Wenbin He, Liang Gou 等IEEE VIS 2023 · 被引用 6 次
- Annealing-based Label-Transfer Learning for Open World Object DetectionYuqing Ma, Hainan Li, Zhange Zhang, Jinyang Guo 等CVPR 2023
