Localizing Events in Videos with Multimodal Queries
Gengyuan Zhang, Mang Ling Ada Fok, Jialu Ma, Yan Xia, Daniel Cremers, Philip Torr, Volker Tresp, Jindong Gu
Abstract
Localizing events in videos based on semantic queries is a pivotal task in video understanding research and useroriented applications like video search. Yet, current research predominantly relies on natural language queries (NLQs), overlooking the potential of using multimodal queries (MQs) that incorporate images to flexibly represent semantic queries, particularly when it is difficult to express non-verbal or unfamiliar concepts in words. To bridge this gap, we introduce ICQ, a new benchmark designed for localizing events in videos with MQs, alongside an evaluation dataset ICQ-Highlight. To adapt and reevaluate existing video localization models for this new task, we propose 3 Multimodal Query Adaptation methods and a novel Surrogate Fine-tuning strategy, serving as strong baseline methods. ICQ systematically benchmarks 12 state-of-theart backbone models, spanning from specialized video localization models to Video Large Language Models. Our extensive experiments highlight the high potential of using MQs in real-world applications. We believe this is a first step toward video event localization with MQs 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3692a24f-d92a-4b9a-b9bf-d8fd74d85ba5Cited by top-tier papers2
- OVG-HQ: Online Video Grounding with Hybrid-Modal QueriesRunhao Zeng, Jiaqi Mao, Minghao Lai, Minh Hieu Phan et al.ICCV 2025 · 4 citations
- Invert4TVG: A Temporal Video Grounding Framework with Inversion Tasks Preserving Action Understanding AbilityChenzhaoyu, Hongnan Lin, Yongwei Nie, Fei Ma et al.ICLR 2026 · 3 citations
Builds on51
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Detecting Moments and Highlights in Videos via Natural Language QueriesJie Lei, Tamara L. Berg, Mohit BansalNeurIPS 2021 · 425 citations
- Self-Chained Image-Language Model for Video Localization and Question AnsweringShoubin Yu, Jaemin Cho, Prateek Yadav, Mohit BansalNeurIPS 2023 · 281 citations
- UniVTG: Towards Unified Video-Language Temporal GroundingKevin Qinghong Lin, Pengchuan Zhang, Joya Chen, Shraman Pramanick et al.ICCV 2023 · 221 citations
Related papers
- Transferable Video Moment Localization by Moment-Guided Query PromptingHao Jiang, Yang Yizhang, Yadong MuAAAI 2024 · 2 citations
- See, Rank, and Filter: Important Word-Aware Clip Filtering via Scene Understanding for Moment Retrieval and Highlight DetectionYuEun Lee, Jung Uk KimAAAI 2026
- Measure Twice, Cut Once: A Semantic-Oriented Approach to Video Temporal Localization with Video LLMsZongshang Pang, Mayu Otani, Yuta NakashimaICLR 2026 · 2 citations
- Span-based Localizing Network for Natural Language Video LocalizationHao Zhang, Aixin Sun, Wei Jing, Joey Tianyi ZhouACL 2020 · 279 citations
- ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded SegmentationAli Athar, Xueqing Deng, Liang-Chieh ChenCVPR 2025
