Sniffing Threatening Open-World Objects in Autonomous Driving by Open-Vocabulary Models
Yulin He, Siqi Wang, Wei Chen, Tianci Xun, Yusong Tan
Abstract
Autonomous driving (AD) is a typical application that requires effectively exploiting multimedia information. For AD, it is critical to ensure safety by detecting unknown objects in an open world, driving the demand for open world object detection (OWOD). However, existing OWOD methods treat generic objects beyond known classes in the train set as unknown objects and prioritize recall in evaluation. This encourages excessive false positives and endangers safety of AD. To address this issue, we restrict the definition of unknown objects to threatening objects in AD, and introduce a new evaluation protocol, which is built upon a new metric named U-ARecall, to alleviate biased evaluation caused by neglecting false positives. Under the new evaluation protocol, we re-evaluate existing OWOD methods and discover that they typically perform poorly in AD. Then, we propose a novel OWOD paradigm for AD based on fine-tuning foundational open-vocabulary models (OVMs), as they can exploit rich linguistic and visual prior knowledge for OWOD. Following this new paradigm, we propose a brand-new OWOD solution, which effectively addresses two core challenges of fine-tuning OVMs via two novel techniques: 1) the maintenance of open-world generic knowledge by a dual-branch architecture; 2) the acquisition of scenario-specific knowledge by the visual-oriented contrastive learning scheme. Besides, a dual-branch prediction fusion module is proposed to avoid post-processing and hand-crafted heuristics. Extensive experiments show that our proposed method not only surpasses classic OWOD methods in unknown object detection by a large margin (∼× U-ARecall), but also notably outperforms OVMs without fine-tuning in known object detection (∼ 20% K-mAP). Our codes are available at https://github.com/harrylin-hyl/AD-OWOD.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 2f6bcda4-05ef-427f-8c0a-50dc6e0395fbCited by top-tier papers2
- NavigScene: Bridging Local Perception and Global Navigation for Beyond-Visual-Range Autonomous DrivingQucheng Peng, Chen Bai, Guoxiang Zhang, Bo Xu et al.ACM MM 2025 · 6 citations
- Attention to Threat-Relevant Objects: Reasoning Detection in Autonomous Driving via Multimodal Large Language ModelsYulin He, Wei Chen, Xinbiao Gan, Siqi Wang et al.AAAI 2026
Related papers
- OW-Adapter: Human-Assisted Open-World Object Detection with a Few ExamplesSuphanut Jamonnak, Jiajing Guo, Wenbin He, Liang Gou et al.IEEE VIS 2023 · 6 citations
- Rethinking Open-World Object Detection in Autonomous Driving ScenariosZeyu Ma, Yang Yang, Guoqing Wang, Xing Xu et al.ACM MM 2022 · 39 citations
- Hyp-OW: Exploiting Hierarchical Structure Learning with Hyperbolic Distance Enhances Open World Object DetectionThang Doan, Xin Li, Sima Behpour, Wenbin He et al.AAAI 2024 · 19 citations
- OW-DETR: Open-world Detection TransformerAkshita Gupta, Sanath Narayan, K. J. Joseph, Salman Khan et al.CVPR 2022 · 209 citations
- Open-Scenario Domain Adaptive Object Detection in Autonomous DrivingZeyu Ma, Ziqiang Zheng, Jiwei Wei, Xiaoyong Wei et al.ACM MM 2023 · 2 citations
