EgoHierMask: Hierarchical Semantic-Prior Guided Masked Autoencoder for Egocentric Action Recognition
Jiang Shao, Xinbo Zhao, Xiaochun Zou, Xiaolin Ye
Abstract
Egocentric action recognition holds critical value in augmented reality, embodied AI, and human behavior analysis. While transformer-based masked autoencoders show potential in general video representation learning, their direct application to egocentric vision faces fundamental limitations --- random masking strategies disrupt crucial spatiotemporal features like hand-object interaction hierarchies and viewpoint dynamics by neglecting task-specific semantic priors. Through systematic analysis, this paper reveals three complementary semantic priors for egocentric video understanding: verb-centric motion patterns characterizing hand trajectories, noun-aware attention regions highlighting object contact points, and action-oriented global context integrating holistic semantics. These hierarchical cues address egocentric visual specificity through motion granularity, interaction locality, and semantic integrity. Building on this discovery, we propose EgoHierMask: a hierarchical semantic prior-guided masked autoencoder framework coordinating vision-language knowledge through differentiated masking strategies. The framework employs frozen vision-language teacher models to generate multi-level semantic attention maps, systematically guiding three specialized masking branches: a) dynamic motion masking preserves hand movement continuity through temporal verb attention, b) interaction-sensitive masking maintains object manipulation coherence via spatial noun saliency, and c) spatiotemporal joint masking encodes complete action semantics through global context alignment. Additionally, to enhance learning efficacy, we curate a distribution-balanced pretraining corpus and devise a unified architecture with dual-granularity supervision, combining pixel-level reconstruction with semantic-level distillation within the VideoMAE paradigm.Extensive experiments demonstrate state-of-the-art performance across major benchmarks, validating the crucial value of hierarchical prior injection for egocentric representation learning.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- Symbiotic Attention with Privileged Information for Egocentric Action RecognitionXiaohan Wang, Yu Wu, Linchao Zhu, Yi YangAAAI 2020 · 66 citations
- EgoHandICL: Egocentric 3D Hand Reconstruction with In-Context LearningBinzhu Xie, Shi Qiu, Sicheng Zhang, Yinqiao Wang et al.ICLR 2026 · 4 citations
- EgoPrompt: Prompt Learning for Egocentric Action RecognitionHuaihai Lyu, Chaofan Chen, Yuheng Ji, Changsheng XuACM MM 2025 · 3 citations
- Recurrent Video Masked AutoencodersDaniel Zoran, Nikhil Parthasarathy, Yi Yang, Drew A. Hudson et al.CVPR 2026 · 9 citations
- Video Language Model Pretraining with Spatio-temporal MaskingYue Wu, Zhaobo Qi, Junshu Sun, Yaowei Wang et al.CVPR 2025
