Video2BEV: Transforming Drone Videos to BEVs for Video-Based Geo-Localization
Hao Ju, Shaofei Huang, Si Liu, Zhedong Zheng
摘要
Existing approaches to drone visual geo-localization predominantly adopt the image-based setting, where a single drone-view snapshot is matched with images from other platforms. Such task formulation, however, underutilizes the inherent video output of the drone and is sensitive to occlusions and viewpoint disparity. To address these limitations, we formulate a new video-based drone geolocalization task and propose the Video2BEV paradigm. This paradigm transforms the video into a Bird's Eye View (BEV), simplifying the subsequent inter-platform matching process. In particular, we employ Gaussian Splatting to reconstruct a 3D scene and obtain the BEV projection. Different from the existing transform methods, e.g., polar transform, our BEVs preserve more fine-grained details without significant distortion. To facilitate the discriminative intra-platform representation learning, our Video2BEV paradigm also incorporates a diffusion-based module for generating hard negative samples. To validate our approach, we introduce UniV, a new videobased geo-localization dataset that extends the image-based University-1652 dataset. UniV features flight paths at 30 • and 45 • elevation angles with increased frame rates of up to 10 frames per second (FPS). Extensive experiments on the UniV dataset show that our Video2BEV paradigm achieves competitive recall rates and outperforms conventional video-based methods. Compared to other competitive methods, our proposed approach exhibits robustness at lower elevations with more occlusions. The code is available at: https://github.com/HaoDot/Video2BEV-Open.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-LocalizationJiahao Wen, Hang Yu, Zhedong ZhengNeurIPS 2025 · 被引用 11 次
- Invert4TVG: A Temporal Video Grounding Framework with Inversion Tasks Preserving Action Understanding AbilityChenzhaoyu, Hongnan Lin, Yongwei Nie, Fei Ma 等ICLR 2026 · 被引用 3 次
- Parse, Search, and Confirmation: Training-Free Aerial Vision-and-Dialog Navigation with Chain-of-Thought Reasoning and Structured Spatial MemoryYu Qi, Hongyu Li, Shaofei Huang, Tianrui Hui 等CVPR 2026 · 被引用 3 次
- UniGeoRS: A Unified Benchmark for Tri-view Geo-LocalizationXiao Liang, Huaizhi Tang, Feiyang Zhang, Shiji Yuan 等CVPR 2026
它引用的顶会 Paper31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
相关 Paper
- University-1652: A Multi-view Multi-source Benchmark for Drone-based Geo-localizationZhedong Zheng, Yunchao Wei, Yi YangACM MM 2020 · 被引用 390 次
- MMGeo: Multimodal Compositional Geo-Localization for UAVsYuxiang Ji, Boyong He, Zhuoyue Tan, Liaoni WuICCV 2025 · 被引用 5 次
- Game4Loc: A UAV Geo-Localization Benchmark from Game DataYuxiang Ji, Boyong He, Zhuoyue Tan, Liaoni WuAAAI 2025 · 被引用 35 次
- BEVPlace: Learning LiDAR-based Place Recognition using Bird's Eye View ImagesLun Luo, Shuhang Zheng, Yixuan Li, Yongzhi Fan 等ICCV 2023 · 被引用 97 次
- UniMODE: Unified Monocular 3D Object DetectionZhuoling Li, Xiaogang Xu, Ser-Nam Lim, Hengshuang ZhaoCVPR 2024
