Home Action Genome: Cooperative Compositional Action Understanding
Nishant Rai, Haofeng Chen, Jingwei Ji, Rishi Desai, Kazuki Kozuka, Shun Ishizaka, Ehsan Adeli, Juan Carlos Niebles
摘要
Existing research on action recognition treats activities as monolithic events occurring in videos. Recently, the benefits of formulating actions as a combination of atomicactions have shown promise in improving action understanding with the emergence of datasets containing such annotations, allowing us to learn representations capturing this information. However, there remains a lack of studies that extend action composition and leverage multiple viewpoints and multiple modalities of data for representation learning. To promote research in this direction, we introduce Home Action Genome (HOMAGE): a multi-view action dataset with multiple modalities and view-points supplemented with hierarchical activity and atomic action labels together with dense scene composition labels. Leveraging rich multi-modal and multi-view settings, we propose Cooperative Compositional Action Understanding (CCAU), a cooperative learning framework for hierarchical action recognition that is aware of compositional action elements. CCAU shows consistent performance improvements across all modalities. Furthermore, we demonstrate the utility of co-learning compositions in few-shot action recognition by achieving 28.6% mAP with just a single sample.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Audio-Adaptive Activity Recognition Across Video DomainsYunhua Zhang, Hazel Doughty, Ling Shao, Cees G. M. SnoekCVPR 2022 · 被引用 31 次
- Learning from Semantic Alignment between Unpaired Multiviews for Egocentric Video RecognitionQitong Wang, Long Zhao, Liangzhe Yuan, Ting Liu 等ICCV 2023 · 被引用 26 次
- Action Scene Graphs for Long-Form Understanding of Egocentric VideosIvan Rodin, Antonino Furnari, Kyle Min, Subarna Tripathi 等CVPR 2024 · 被引用 14 次
- EMHI: A Multimodal Egocentric Human Motion Dataset with HMD and Body-Worn IMUsZhen Fan, Peng Dai, Zhuo Su, Xu Gao 等AAAI 2025 · 被引用 13 次
- Adaptive Visual Scene Understanding: Incremental Scene Graph GenerationNaitik Khandelwal, Xiao Liu, Mengmi ZhangNeurIPS 2024 · 被引用 9 次
它引用的顶会 Paper6
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- Self-supervised Co-Training for Video Representation LearningTengda Han, Weidi Xie, Andrew ZissermanNeurIPS 2020 · 被引用 405 次
- MMAct: A Large-Scale Dataset for Cross Modal Human Action UnderstandingQuan Kong, Ziming Wu, Ziwei Deng, Martin Klinkigt 等ICCV 2019 · 被引用 108 次
- Action Genome: Actions As Compositions of Spatio-Temporal Scene GraphsJingwei Ji, Ranjay Krishna, Li Fei-Fei, Juan Carlos NieblesCVPR 2020
相关 Paper
- Action Motifs: Self-Supervised Hierarchical Representation of Human Body MovementsGenki Kinoshita, Shu Nakamura, Ryo Kawahara, Shohei Nobuhara 等CVPR 2026
- Beyond Label Semantics:Language-Guided Action Anatomy for Few-Shot Action RecognitionZefeng Qian, Xincheng Yao, Yifei Huang, Chongyang Zhang 等ICCV 2025 · 被引用 4 次
- MOMA: Multi-Object Multi-Actor Activity ParsingZelun Luo, Wanze Xie, Siddharth Kapoor, Yiyun Liang 等NeurIPS 2021 · 被引用 34 次
- Open-Ended Hierarchical Streaming Video Understanding with Vision Language ModelsHyolim Kang, Yunsu Park, Youngbeom Yoo, Yeeun Choi 等ICCV 2025 · 被引用 2 次
- Visual Knowledge Graph for Human Action Reasoning in VideosYue Ma, Yali Wang, Yue Wu, Ziyu Lyu 等ACM MM 2022 · 被引用 29 次
