Symbiotic Attention with Privileged Information for Egocentric Action Recognition
Xiaohan Wang, Yu Wu, Linchao Zhu, Yi Yang
摘要
Egocentric video recognition is a natural testbed for diverse interaction reasoning. Due to the large action vocabulary in egocentric video datasets, recent studies usually utilize a two-branch structure for action recognition, i.e., one branch for verb classification and the other branch for noun classification. However, correlation study between the verb and the noun branches have been largely ignored. Besides, the two branches fail to exploit local features due to the absence of position-aware attention mechanism. In this paper, we propose a novel Symbiotic Attention framework leveraging Privileged information (SAP) for egocentric video recognition. Finer position-aware object detection features can facilitate the understanding of actor's interaction with the object. We introduce these features in action recognition and regard them as privileged information. Our framework enables mutual communication among the verb branch, the noun branch, and the privileged information. This communication process not only injects local details into global features, but also exploits implicit guidance about the spatio-temporal position of an on-going action. We introduce a novel symbiotic attention (SA) to enable effective communication. It first normalizes the detection guided features on one branch to underline the action-relevant information from the other branch. SA adaptively enhances the interactions among the three sources. To further catalyze this communication, spatial relations are uncovered for the selection of most action-relevant information. It identifies the most valuable and discriminative feature for classification. We validate the effectiveness of our SAP quantitatively and qualitatively. Notably, it achieves the state-of-the-art on two large-scale egocentric video datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Interactive Prototype Learning for Egocentric Action RecognitionXiaohan Wang, Linchao Zhu, Heng Wang, Yi YangICCV 2021 · 被引用 78 次
- Group Contextualization for Video RecognitionYanbin Hao, Hao Zhang, Chong-Wah Ngo, Xiangnan HeCVPR 2022 · 被引用 48 次
- Uncertainty-aware State Space Transformer for Egocentric 3D Hand Trajectory ForecastingWentao Bao, Lele Chen, Libing Zeng, Zhong Li 等ICCV 2023 · 被引用 34 次
- Learning from Semantic Alignment between Unpaired Multiviews for Egocentric Video RecognitionQitong Wang, Long Zhao, Liangzhe Yuan, Ting Liu 等ICCV 2023 · 被引用 26 次
- A Hybrid Egocentric Activity Anticipation Framework via Memory-Augmented Recurrent and One-shot Representation ForecastingTianshan Liu, Kin-Man LamCVPR 2022 · 被引用 24 次
它引用的顶会 Paper1
相关 Paper
- EgoHierMask: Hierarchical Semantic-Prior Guided Masked Autoencoder for Egocentric Action RecognitionJiang Shao, Xinbo Zhao, Xiaochun Zou, Xiaolin YeACM MM 2025
- EgoPrompt: Prompt Learning for Egocentric Action RecognitionHuaihai Lyu, Chaofan Chen, Yuheng Ji, Changsheng XuACM MM 2025 · 被引用 3 次
- Intention-Guided Cognitive Reasoning for Egocentric Long-Term Action AnticipationQiaohui Chu, Haoyu Zhang, Meng Liu, Yisen Feng 等AAAI 2026 · 被引用 3 次
- Slowfast Diversity-aware Prototype Learning for Egocentric Action RecognitionGuangzhao Dai, Xiangbo Shu, Rui Yan, Peng Huang 等ACM MM 2023 · 被引用 3 次
- The Audio-Visual Conversational Graph: From an Egocentric-Exocentric PerspectiveWenqi Jia, Miao Liu, Hao Jiang, Ishwarya Ananthabhotla 等CVPR 2024
