Lune

ICCV2025Top-tier venue

CARIM: Caption-Based Autonomous Driving Scene Retrieval via Inclusive Text Matching

Minjoo Ki, Daejung Kim, Kisung Kim, Seon Joo Kim, Jinhan Lee

2025Year
2Citations
1Top-tier citations

Abstract

Text-to-video retrieval is a powerful tool for navigating vast video databases. This is especially useful in autonomous driving to retrieve scenes from a text query to simulate and evaluate a driving system in desired scenarios. However, traditional ranking-based retrieval methods often return partial matches that fail to satisfy all query conditions. To address this, we introduce Inclusive Text-to-Video Retrieval, which retrieves only videos that meet all specified conditions, regardless of additional irrelevant elements. We propose CARIM, a driving scene retrieval framework that employs inclusive text matching. By utilizing Vision-Language Model and Large Language Model to generate compressed captions for driving scenes, we reformulate text-to-video retrieval as a more efficient text-to-text retrieval problem, eliminating modality mismatch and heavy annotation cost. We present a novel positive and negative data curation strategy and an attention-based scoring mechanism tailored for driving scene retrieval. Experiments show that CARIM outperforms state-of-the-art retrieval methods, excelling in edge cases where traditional models fail.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext e6923415-13d7-405e-990d-40fc64de18f3

Cited by top-tier papers1

Ask how each one uses it

Builds on7

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines