Beyond the Horizon: Decoupling Multi-View UAV Action Recognition via Partial Order Transfer
Wenxuan Liu, Zhuo Zhou, Xuemei Jia, Siyuan Yang, Wenxin Huang, Xian Zhong, Chia-Wen Lin
Abstract
Action recognition in unmanned aerial vehicles (UAVs) poses unique challenges due to significant view variations along the vertical spatial axis. Unlike traditional ground-based settings, UAVs capture actions at a wide range of altitudes, resulting in considerable appearance discrepancies. We introduce a multiview formulation tailored to varying UAV altitudes and empirically observe a partial order among views, where recognition accuracy consistently decreases as altitude increases. This observation motivates a novel approach that explicitly models the hierarchical structure of UAV views to improve recognition performance across altitudes. To this end, we propose the Partial Order Guided Multi-View Network (POG-MVNet), designed to address drastic view variations by effectively leveraging view-dependent information across different altitude levels. The framework comprises three key components: a View Partition (VP) module, which uses the head-to-body ratio to group views by altitude; an Order-aware Feature Decoupling (OFD) module, which disentangles action-relevant and view-specific features under partial order guidance; and an Action Partial Order Guide (APOG), which uses the partial order to transfer informative knowledge from easier views to more challenging ones. We conduct experiments on DRONE-ACTION, MOD20, and UAV, demonstrating that POG-MVNet significantly outperforms competing methods. For example, POG-MVNet achieves a 4.7% improvement on DRONE-ACTION and a 3.5% improvement on UAV compared to state-of-the-art methods ASAT and FAR. Code will be released soon.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a5d40d8b-8c20-48d2-9991-155fbca577ffBuilds on14
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Gibbs Sampling with PeoplePeter M. C. Harrison, Raja Marjieh, Federico Adolfi, Pol van Rijn et al.NeurIPS 2020 · 81 citations
- Scalable Diverse Model Selection for Accessible Transfer LearningDaniel Bolya, Rohit Mittapalli, Judy HoffmanNeurIPS 2021 · 61 citations
- DVANet: Disentangling View and Action Features for Multi-View Action RecognitionNyle Siddiqui, Praveen Tirupattur, Mubarak ShahAAAI 2024 · 39 citations
- Stochastic Backpropagation: A Memory Efficient Strategy for Training Video ModelsFeng Cheng, Mingze Xu, Yuanjun Xiong, Hao Chen et al.CVPR 2022 · 11 citations
Related papers
- Alleviating Spatial Misalignment and Motion Interference for UAV-based Video RecognitionGege Shi, Xueyang Fu, Chengzhi Cao, Zheng-Jun ZhaACM MM 2023 · 8 citations
- Multi-Object Tracking Meets Moving UAVShuai Liu, Xin Li, Huchuan Lu, You HeCVPR 2022 · 112 citations
- MOR-UAV: A Benchmark Dataset and Baselines for Moving Object Recognition in UAV VideosMurari Mandal, Lav Kush Kumar, Santosh Kumar VipparthiACM MM 2020 · 58 citations
- IGIANet: Illumination Guided Implicit Alignment Network for Infrared-Visible UAV DetectionXiangqi Chen, Dawei Zhang, Li Zhao, Chengzhuan Yang et al.AAAI 2026
- Tracking the Unstable: Appearance-Guided Motion Modeling for Robust Multi-Object Tracking in UAV-Captured VideosJianbo Ma, Hui Luo, Qi Chen, Yuankai Qi et al.AAAI 2026 · 2 citations
