Alleviating Spatial Misalignment and Motion Interference for UAV-based Video Recognition
Gege Shi, Xueyang Fu, Chengzhi Cao, Zheng-Jun Zha
摘要
Recognizing activities with Unmanned Aerial Vehicles (UAVs) is essential for many applications, while existing video recognition methods are mainly designed for ground cameras and do not account for UAV changing attitudes and fast motion. This creates spatial misalignment of small objects between frames, leading to inaccurate visual movement in drone videos. Additionally, camera motion relative to objects in the video causes relative movements that visually affect object motion and can result in misunderstandings of video content. To address these issues, we present a novel framework named Attentional Spatial and Adaptive Temporal Relations Modeling. First, to mitigate the spatial misalignment of small objects between frames, we design an Attentional Patch-level Spatial Enrichment (APSE) module that models dependencies among patches and enhances patch-level features. Then, we propose a Multi-scale Temporal and Spatial Mixer (MTSM) module that is capable of adapting to disturbances caused by the UAV flight and modeling various temporal clues. By integrating APSE and MTSM into a single model, our network can effectively and accurately capture spatiotemporal relations for UAV videos. Extensive experiments on several benchmarks demonstrate the superiority of our method over state-of-the-art approaches. For instance, our network achieves a classification accuracy of 68.1% with an absolute gain of 1.3% compared to FuTH-Net on the ERA dataset.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- HazeSpace2M: A Dataset for Haze Aware Single Image DehazingMd Tanvir Islam, Nasir Rahim, Saeed Anwar, Muhammad Saqib 等ACM MM 2024 · 被引用 19 次
- 3D-aware Select, Expand, and Squeeze Token for Aerial Action RecognitionLuying Peng, Xiangbo Shu, Yazhou Yao, Guo-Sen XieAAAI 2025 · 被引用 1 次
- Beyond the Horizon: Decoupling Multi-View UAV Action Recognition via Partial Order TransferWenxuan Liu, Zhuo Zhou, Xuemei Jia, Siyuan Yang 等AAAI 2026 · 被引用 1 次
相关 Paper
- Multi-Object Tracking Meets Moving UAVShuai Liu, Xin Li, Huchuan Lu, You HeCVPR 2022 · 被引用 112 次
- MOR-UAV: A Benchmark Dataset and Baselines for Moving Object Recognition in UAV VideosMurari Mandal, Lav Kush Kumar, Santosh Kumar VipparthiACM MM 2020 · 被引用 58 次
- Tracking the Unstable: Appearance-Guided Motion Modeling for Robust Multi-Object Tracking in UAV-Captured VideosJianbo Ma, Hui Luo, Qi Chen, Yuankai Qi 等AAAI 2026 · 被引用 2 次
- FOLT: Fast Multiple Object Tracking from UAV-captured Videos Based on Optical FlowMufeng Yao, Jiaqi Wang, Jinlong Peng, Mingmin Chi 等ACM MM 2023 · 被引用 28 次
- Spatio-Temporal Context Learning with Temporal Difference Convolution for Moving Infrared Small Target DetectionHouzhang Fang, Shukai Guo, Qiuhuan Chen, Yi Chang 等AAAI 2026
