RELOCATE: A Simple Training-Free Baseline for Visual Query Localization Using Region-Based Representations
Savya Khosla, Sethuraman TV, Alexander G. Schwing, Derek Hoiem
摘要
We present RELOCATE, a simple training-free baseline designed to perform the challenging task of visual query localization in long videos. To eliminate the need for task-specific training and efficiently handle long videos, RELOCATE leverages a region-based representation derived from pretrained vision models. At a high level, it follows the classic object localization approach: (1) identify all objects in each video frame, (2) compare the objects with the given query and select the most similar ones, and (3) perform bidirectional tracking to get a spatio-temporal response. However, we propose some key enhancements to handle small objects, cluttered scenes, partial visibility, and varying appearances. Notably, we refine the selected objects for accurate localization and generate additional visual queries to capture visual variations. We evaluate RELOCATE on the challenging Ego4D Visual Query 2D Localization dataset, establishing a new baseline that outperforms prior task-specific methods by 49% (relative improvement) in spatio-temporal average precision.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- REN: Fast and Efficient Region Encodings from Patch-Based Image EncodersSavya Khosla, Sethuraman TV, Barnett Lee, Alex Schwing 等NeurIPS 2025 · 被引用 5 次
- Towards Visual Query Localization in the 3D WorldLiang Peng, Bohan Tan, Zhipeng Zhang, Haobo Li 等CVPR 2026 · 被引用 1 次
- PlanaReLoc: Camera Relocalization in 3D Planar Primitives via Region-Based Structure MatchingHanqiao Ye, Yuzhou Liu, Yangdong Liu, Shuhan ShenCVPR 2026
- EAGLE: Episodic Appearance- and Geometry-aware Memory for Unified 2D-3D Visual Query Localization in Egocentric VisionYifei Cao, Yu Liu, Guolong Wang, Zhu Liu 等AAAI 2026
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller 等ICML 2020 · 被引用 1,220 次
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang 等ICCV 2021 · 被引用 1,062 次
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 被引用 927 次
相关 Paper
- Single-Stage Visual Query Localization in Egocentric VideosHanwen Jiang, Santhosh Kumar Ramakrishnan, Kristen GraumanNeurIPS 2023 · 被引用 27 次
- PRVQL: Progressive Knowledge-Guided Refinement for Robust Egocentric Visual Query LocalizationBing Fan, Yunhe Feng, Yapeng Tian, James Chenhao Liang 等ICCV 2025 · 被引用 1 次
- Weakly-Supervised Video Re-Localization with Multiscale Attention ModelYung-Han Huang, Kuang-Jui Hsu, Shyh-Kang Jeng, Yen-Yu LinAAAI 2020 · 被引用 12 次
- QDETRv: Query-Guided DETR for One-Shot Object Localization in VideosYogesh Kumar, Saswat Mallick, Anand Mishra, Sowmya Rasipuram 等AAAI 2024 · 被引用 4 次
- EgoLoc: Revisiting 3D Object Localization from Egocentric Videos with Visual QueriesJinjie Mai, Abdullah Hamdi, Silvio Giancola, Chen Zhao 等ICCV 2023 · 被引用 26 次
