Lune

ICCV2025顶会

CARIM: Caption-Based Autonomous Driving Scene Retrieval via Inclusive Text Matching

Minjoo Ki, Daejung Kim, Kisung Kim, Seon Joo Kim, Jinhan Lee

2025年份
2被引次数
1顶会引用

摘要

Text-to-video retrieval is a powerful tool for navigating vast video databases. This is especially useful in autonomous driving to retrieve scenes from a text query to simulate and evaluate a driving system in desired scenarios. However, traditional ranking-based retrieval methods often return partial matches that fail to satisfy all query conditions. To address this, we introduce Inclusive Text-to-Video Retrieval, which retrieves only videos that meet all specified conditions, regardless of additional irrelevant elements. We propose CARIM, a driving scene retrieval framework that employs inclusive text matching. By utilizing Vision-Language Model and Large Language Model to generate compressed captions for driving scenes, we reformulate text-to-video retrieval as a more efficient text-to-text retrieval problem, eliminating modality mismatch and heavy annotation cost. We present a novel positive and negative data curation strategy and an attention-based scoring mechanism tailored for driving scene retrieval. Experiments show that CARIM outperforms state-of-the-art retrieval methods, excelling in edge cases where traditional models fail.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper7

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖