OW-DETR: Open-world Detection Transformer
Akshita Gupta, Sanath Narayan, K. J. Joseph, Salman Khan, Fahad Shahbaz Khan, Mubarak Shah
Abstract
Open-world object detection (OWOD) is a challenging computer vision problem, where the task is to detect a known set of object categories while simultaneously identifying unknown objects. Additionally, the model must incrementally learn new classes that become known in the next training episodes. Distinct from standard object detection, the OWOD setting poses significant challenges for generating quality candidate proposals on potentially unknown objects, separating the unknown objects from the background and detecting diverse unknown objects. Here, we introduce a novel end-to-end transformer-based framework, OW-DETR, for open-world object detection. The proposed OW-DETR comprises three dedicated components namely, attention-driven pseudo-labeling, novelty classification and objectness scoring to explicitly address the aforementioned OWOD challenges. Our OW-DETR explicitly encodes multi-scale contextual information, possesses less inductive bias, enables knowledge transfer from known classes to the unknown class and can better discriminate between unknown objects and background. Comprehensive experiments are performed on two benchmarks: MS-COCO and PASCAL VOC. The extensive ablations reveal the merits of our proposed contributions. Further, our model out-performs the recently introduced OWOD approach, ORE, with absolute gains ranging from 1.8% to 3.3% in terms of unknown recall on MS-COCO. In the case of incremental object detection, OW-DETR outperforms the state-of-the-art for all settings on PASCAL VOC. Our code is available at https://github.com/akshitac8/OW-DEtr.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2fa35e39-562d-4fa7-993e-5f49a89c295aCited by top-tier papers69
- SIREN: Shaping Representations for Detecting Out-of-Distribution ObjectsXuefeng Du, Gabriel Gozum, Yifei Ming, Yixuan LiNeurIPS 2022 · 105 citations
- CoDA: Collaborative Novel Box Discovery and Cross-modal Alignment for Open-vocabulary 3D Object DetectionYang Cao, Yihan Zeng, Hang Xu, Dan XuNeurIPS 2023 · 69 citations
- End-to-End Zero-Shot HOI Detection via Vision and Language Knowledge DistillationMingrui Wu, Jiaxin Gu, Yunhang Shen, Mingbao Lin et al.AAAI 2023 · 64 citations
- Augmented Box Replay: Overcoming Foreground Shift for Incremental Object DetectionYuyang Liu, Yang Cong, Dipam Goswami, Xialei Liu et al.ICCV 2023 · 52 citations
- A Hard-to-Beat Baseline for Training-free CLIP-based AdaptationZhengbo Wang, Jian Liang, Lijun Sheng, Ran He et al.ICLR 2024 · 51 citations
Builds on7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Intriguing Properties of Vision TransformersMuzammal Naseer, Kanchana Ranasinghe, Salman Khan, Munawar Hayat et al.NeurIPS 2021 · 863 citations
- Frustratingly Simple Few-Shot Object DetectionXin Wang, Thomas E. Huang, Joseph Gonzalez, Trevor Darrell et al.ICML 2020 · 723 citations
Related papers
- PROB: Probabilistic Objectness for Open World Object DetectionOrr Zohar, Kuan-Chieh Wang, Serena YeungCVPR 2023
- CAT: LoCalization and IdentificAtion Cascade Detection Transformer for Open-World Object DetectionShuailei Ma, Yuefeng Wang, Ying Wei, Jiaqi Fan et al.CVPR 2023
- Towards Open World Object DetectionK. J. Joseph, Salman H. Khan, Fahad Shahbaz Khan, Vineeth N. BalasubramanianCVPR 2021
- OW-Adapter: Human-Assisted Open-World Object Detection with a Few ExamplesSuphanut Jamonnak, Jiajing Guo, Wenbin He, Liang Gou et al.IEEE VIS 2023 · 6 citations
- Semi-supervised Open-World Object DetectionSahal Shaji Mullappilly, Abhishek Singh Gehlot, Rao Muhammad Anwer, Fahad Shahbaz Khan et al.AAAI 2024 · 19 citations
