MUP: Multi-granularity Unified Perception for Panoramic Activity Recognition
Meiqi Cao, Rui Yan, Xiangbo Shu, Jiachao Zhang, Jinpeng Wang, Guo-Sen Xie
Abstract
Panoramic activity recognition is required to jointly identify multi-granularity human behaviors including individual actions, group activities, and global activities in multi-person videos. Previous methods encode these behaviors hierarchically through multiple stages, which disturb the inherent co-occurrence across multi-granularity behaviors in the same scene. To this end, we propose a novel Multi-granularity Unified Perception (MUP) framework that perceives different granularity behaviors universally to explore the co-occurrence motion pattern via the same parameters in an end-to-end fashion. To be specific, the proposed framework stacks three Unified Motion Encoding (UME) blocks for modeling multiple granularity behaviors with shared parameters. UME block mines intra-relevant and cross-relevant semantics synchronously from input feature sequences via Intra-granularity Motion Embedding (IME) and Cross-granularity Motion Prototyping (CMP). In particular, IME aims to model the interactions among visual features within each granularity based on the attention mechanism. CMP aims to aggregate features across different granularities (i.e., person to group) via several learnable prototypes. Extensive experiments demonstrate that MUP outperforms the state-of-the-art methods on JRDB-PAR and has satisfactory interpretability.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 237748ca-37d1-4563-ba98-01a8f57b9063Cited by top-tier papers3
- AdaFPP: Adapt-Focused Bi-Propagating Prototype Learning for Panoramic Activity RecognitionMeiqi Cao, Rui Yan, Xiangbo Shu, Guangzhao Dai et al.ACM MM 2024 · 3 citations
- 3D-aware Select, Expand, and Squeeze Token for Aerial Action RecognitionLuying Peng, Xiangbo Shu, Yazhou Yao, Guo-Sen XieAAAI 2025 · 1 citation
- Exploiting Frequency Dynamics for Enhanced Multimodal Event-Based Action RecognitionMeiqi Cao, Xiangbo Shu, Xin Jiang, Rui Yan et al.ICCV 2025 · 1 citation
Related papers
- Label Text-aided Hierarchical Semantics Mining for Panoramic Activity RecognitionTianshan Liu, Kin-Man Lam, Bing-Kun BaoACM MM 2024
- Multi-Granularity Reference-Aided Attentive Feature Aggregation for Video-Based Person Re-IdentificationZhizheng Zhang, Cuiling Lan, Wenjun Zeng, Zhibo ChenCVPR 2020
- You Need to Read Again: Multi-granularity Perception Network for Moment Retrieval in VideosXin Sun, Xuan Wang, Jialin Gao, Qiong Liu et al.SIGIR 2022 · 42 citations
- Advancing Visual Large Language Model for Multi-Granular Versatile PerceptionWentao Xiang, Haoxian Tan, Yujie Zhong, Cong Wei et al.ICCV 2025 · 1 citation
- Human Identification and Interaction Detection in Cross-View Multi-Person Videos with Wearable CamerasJiewen Zhao, Ruize Han, Yiyang Gan, Liang Wan et al.ACM MM 2020 · 32 citations
