Symbiotic Attention with Privileged Information for Egocentric Action Recognition
Xiaohan Wang, Yu Wu, Linchao Zhu, Yi Yang
Abstract
Egocentric video recognition is a natural testbed for diverse interaction reasoning. Due to the large action vocabulary in egocentric video datasets, recent studies usually utilize a two-branch structure for action recognition, i.e., one branch for verb classification and the other branch for noun classification. However, correlation study between the verb and the noun branches have been largely ignored. Besides, the two branches fail to exploit local features due to the absence of position-aware attention mechanism. In this paper, we propose a novel Symbiotic Attention framework leveraging Privileged information (SAP) for egocentric video recognition. Finer position-aware object detection features can facilitate the understanding of actor's interaction with the object. We introduce these features in action recognition and regard them as privileged information. Our framework enables mutual communication among the verb branch, the noun branch, and the privileged information. This communication process not only injects local details into global features, but also exploits implicit guidance about the spatio-temporal position of an on-going action. We introduce a novel symbiotic attention (SA) to enable effective communication. It first normalizes the detection guided features on one branch to underline the action-relevant information from the other branch. SA adaptively enhances the interactions among the three sources. To further catalyze this communication, spatial relations are uncovered for the selection of most action-relevant information. It identifies the most valuable and discriminative feature for classification. We validate the effectiveness of our SAP quantitatively and qualitatively. Notably, it achieves the state-of-the-art on two large-scale egocentric video datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fd8a1f75-8967-44a4-8f1c-4975d06848d8Cited by top-tier papers15
- Interactive Prototype Learning for Egocentric Action RecognitionXiaohan Wang, Linchao Zhu, Heng Wang, Yi YangICCV 2021 · 78 citations
- Group Contextualization for Video RecognitionYanbin Hao, Hao Zhang, Chong-Wah Ngo, Xiangnan HeCVPR 2022 · 48 citations
- Uncertainty-aware State Space Transformer for Egocentric 3D Hand Trajectory ForecastingWentao Bao, Lele Chen, Libing Zeng, Zhong Li et al.ICCV 2023 · 34 citations
- Learning from Semantic Alignment between Unpaired Multiviews for Egocentric Video RecognitionQitong Wang, Long Zhao, Liangzhe Yuan, Ting Liu et al.ICCV 2023 · 26 citations
- A Hybrid Egocentric Activity Anticipation Framework via Memory-Augmented Recurrent and One-shot Representation ForecastingTianshan Liu, Kin-Man LamCVPR 2022 · 24 citations
Builds on1
Related papers
- EgoHierMask: Hierarchical Semantic-Prior Guided Masked Autoencoder for Egocentric Action RecognitionJiang Shao, Xinbo Zhao, Xiaochun Zou, Xiaolin YeACM MM 2025
- EgoPrompt: Prompt Learning for Egocentric Action RecognitionHuaihai Lyu, Chaofan Chen, Yuheng Ji, Changsheng XuACM MM 2025 · 3 citations
- Intention-Guided Cognitive Reasoning for Egocentric Long-Term Action AnticipationQiaohui Chu, Haoyu Zhang, Meng Liu, Yisen Feng et al.AAAI 2026 · 3 citations
- Slowfast Diversity-aware Prototype Learning for Egocentric Action RecognitionGuangzhao Dai, Xiangbo Shu, Rui Yan, Peng Huang et al.ACM MM 2023 · 3 citations
- The Audio-Visual Conversational Graph: From an Egocentric-Exocentric PerspectiveWenqi Jia, Miao Liu, Hao Jiang, Ishwarya Ananthabhotla et al.CVPR 2024
