CARIM: Caption-Based Autonomous Driving Scene Retrieval via Inclusive Text Matching
Minjoo Ki, Daejung Kim, Kisung Kim, Seon Joo Kim, Jinhan Lee
Abstract
Text-to-video retrieval is a powerful tool for navigating vast video databases. This is especially useful in autonomous driving to retrieve scenes from a text query to simulate and evaluate a driving system in desired scenarios. However, traditional ranking-based retrieval methods often return partial matches that fail to satisfy all query conditions. To address this, we introduce Inclusive Text-to-Video Retrieval, which retrieves only videos that meet all specified conditions, regardless of additional irrelevant elements. We propose CARIM, a driving scene retrieval framework that employs inclusive text matching. By utilizing Vision-Language Model and Large Language Model to generate compressed captions for driving scenes, we reformulate text-to-video retrieval as a more efficient text-to-text retrieval problem, eliminating modality mismatch and heavy annotation cost. We present a novel positive and negative data curation strategy and an attention-based scoring mechanism tailored for driving scene retrieval. Experiments show that CARIM outperforms state-of-the-art retrieval methods, excelling in edge cases where traditional models fail.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e6923415-13d7-405e-990d-40fc64de18f3Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- HiT: Hierarchical Transformer with Momentum Contrast for Video-Text RetrievalSong Liu, Haoqi Fan, Shengsheng Qian, Yiru Chen et al.ICCV 2021 · 172 citations
- BEV-TSR: Text-Scene Retrieval in BEV Space for Autonomous DrivingTao Tang, Dafeng Wei, Zhengyu Jia, Tian Gao et al.AAAI 2025 · 21 citations
Related papers
- Bridging Information Asymmetry in Text-video Retrieval: A Data-centric ApproachZechen Bai, Tianjun Xiao, Tong He, Pichao Wang et al.ICLR 2025
- Narrating the Video: Boosting Text-Video Retrieval via Comprehensive Utilization of Frame-Level CaptionsChan Hur, Jeong-Hun Hong, Dong-hun Lee, Dabin Kang et al.CVPR 2025
- Multi-Modal Inductive Framework for Text-Video RetrievalQian Li, Yucheng Zhou, Cheng Ji, Feihong Lu et al.ACM MM 2024 · 8 citations
- ViLL-E: Video LLM Embeddings for RetrievalRohit Gupta, Jayakrishnan Unnikrishnan, Fan Fei, Sheng Liu et al.ACL 2026
- DGL: Dynamic Global-Local Prompt Tuning for Text-Video RetrievalXiangpeng Yang, Linchao Zhu, Xiaohan Wang, Yi YangAAAI 2024 · 53 citations
