SearchAD: Large-Scale Rare Image Retrieval Dataset for Autonomous Driving
Felix Embacher, Jonas Uhrig, Marius Cordts, Markus Enzweiler
Abstract
Retrieving rare and safety-critical driving scenarios from large-scale datasets is essential for building robust autonomous driving (AD) systems. As dataset sizes continue to grow, the key challenge shifts from collecting more data to efficiently identifying the most relevant samples. We introduce SearchAD, a large-scale rare image retrieval dataset for AD containing over 423k frames drawn from 11 established datasets. SearchAD provides high-quality manual annotations of more than 513k bounding boxes covering 90 rare categories. It specifically targets the needle-in-a-haystack problem of locating extremely rare classes, with some appearing fewer than 50 times across the entire dataset. Unlike existing benchmarks, which focused on instance-level retrieval, SearchAD emphasizes semantic image retrieval with a well-defined data split, enabling text-to-image and image-to-image retrieval, few-shot learning, and fine-tuning of multi-modal retrieval models. Comprehensive evaluations show that text-based methods outperform image-based ones due to stronger inherent semantic grounding. While models directly aligning spatial visual features with language achieve the best zero-shot results, and our fine-tuning baseline significantly improves performance, absolute retrieval capabilities remain unsatisfactory. With a held-out test set on a public benchmark server, SearchAD establishes the first large-scale dataset for retrieval-driven data curation and long-tail perception research in AD: https://iis-esslingen.github.io/searchad/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bba26c17-a099-4af4-ae95-6c8b5c1d47d5Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- ACDC: The Adverse Conditions Dataset with Correspondences for Semantic Driving Scene UnderstandingChristos Sakaridis, Dengxin Dai, Luc Van GoolICCV 2021 · 655 citations
Related papers
- Data Roaming and Quality Assessment for Composed Image RetrievalMatan Levy, Rami Ben-Ari, Nir Darshan, Dani LischinskiAAAI 2024 · 65 citations
- SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object DetectionHao Vo, Khoa Vo, Thinh Phan, Ngo Xuan Cuong et al.CVPR 2026 · 1 citation
- Zenseact Open Dataset: A large-scale and diverse multimodal dataset for autonomous drivingMina Alibeigi, William Ljungbergh, Adam Tonderski, Georg Hess et al.ICCV 2023 · 106 citations
- ILIAS: Instance-Level Image retrieval At ScaleGiorgos Kordopatis-Zilos, Vladan Stojnic, Anna Manko, Pavel Suma et al.CVPR 2025
- Generalized Predictive Model for Autonomous DrivingJiazhi Yang, Shenyuan Gao, Yihang Qiu, Li Chen et al.CVPR 2024 · 32 citations
