Single-Stage Visual Query Localization in Egocentric Videos
Hanwen Jiang, Santhosh Kumar Ramakrishnan, Kristen Grauman
Abstract
Visual Query Localization on long-form egocentric videos requires spatio-temporal search and localization of visually specified objects and is vital to build episodic memory systems. Prior work develops complex multi-stage pipelines that leverage well-established object detection and tracking methods to perform VQL. However, each stage is independently trained and the complexity of the pipeline results in slow inference speeds. We propose VQLoC, a novel single-stage VQL framework that is end-to-end trainable. Our key idea is to first build a holistic understanding of the query-video relationship and then perform spatio-temporal localization in a single shot manner. Specifically, we establish the query-video relationship by jointly considering query-to-frame correspondences between the query and each video frame and frame-to-frame correspondences between nearby video frames. Our experiments demonstrate that our approach outperforms prior VQL methods by 20% accuracy while obtaining a 10x improvement in inference speed. VQLoC is also the top entry on the Ego4D VQ2D challenge leaderboard. Project page: https://hwjiang1510.github.io/VQLoC/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f01d46c6-ff70-412f-8204-4b6c9c5fe38cCited by top-tier papers7
- Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal UnderstandingYunlong Tang, Daiki Shimada, Jing Bi, Mingqian Feng et al.AAAI 2025 · 29 citations
- REN: Fast and Efficient Region Encodings from Patch-Based Image EncodersSavya Khosla, Sethuraman TV, Barnett Lee, Alex Schwing et al.NeurIPS 2025 · 5 citations
- PRVQL: Progressive Knowledge-Guided Refinement for Robust Egocentric Visual Query LocalizationBing Fan, Yunhe Feng, Yapeng Tian, James Chenhao Liang et al.ICCV 2025 · 1 citation
- Towards Visual Query Localization in the 3D WorldLiang Peng, Bohan Tan, Zhipeng Zhang, Haobo Li et al.CVPR 2026 · 1 citation
- OmniGlue: Generalizable Feature Matching with Foundation Model GuidanceHanwen Jiang, Arjun Karpur, Bingyi Cao, Qixing Huang et al.CVPR 2024
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang et al.ICCV 2021 · 1,062 citations
Related papers
- EgoLoc: Revisiting 3D Object Localization from Egocentric Videos with Visual QueriesJinjie Mai, Abdullah Hamdi, Silvio Giancola, Chen Zhao et al.ICCV 2023 · 26 citations
- RELOCATE: A Simple Training-Free Baseline for Visual Query Localization Using Region-Based RepresentationsSavya Khosla, Sethuraman TV, Alexander G. Schwing, Derek HoiemCVPR 2025
- Where is my Wallet? Modeling Object Proposal Sets for Egocentric Visual Query LocalizationMengmeng Xu, Yanghao Li, Cheng-Yang Fu, Bernard Ghanem et al.CVPR 2023
- EAGLE: Episodic Appearance- and Geometry-aware Memory for Unified 2D-3D Visual Query Localization in Egocentric VisionYifei Cao, Yu Liu, Guolong Wang, Zhu Liu et al.AAAI 2026
- Helping Hands: An Object-Aware Ego-Centric Video Recognition ModelChuhan Zhang, Ankush Gupta, Andrew ZissermanICCV 2023 · 39 citations
