MOMA: Multi-Object Multi-Actor Activity Parsing
Zelun Luo, Wanze Xie, Siddharth Kapoor, Yiyun Liang, Michael Cooper, Juan Carlos Niebles, Ehsan Adeli, Fei-Fei Li
Abstract
Abstract Complex activities often involve multiple humans utilizing different objects to complete actions (e.g., in healthcare settings, physicians, nurses, and patients interact with each other and various medical devices). Recognizing activities poses a challenge that requires a detailed understanding of actors' roles, objects' affordances, and their associated relationships. Furthermore, these purposeful activities comprise multiple achievable steps, including sub-activities and atomic actions, which jointly define a hierarchy of action parts. This paper introduces Activity Parsing as the overarching task of temporal segmentation and classification of activities, sub-activities, atomic actions, along with an instance-level understanding of actors, objects, and their relationships in videos. Involving multiple entities (actors and objects), we argue that traditional pair-wise relationships, often used in scene or action graphs, do not appropriately represent the dynamics between them. Hence, we introduce Action Hypergraph, a new representation of spatial-temporal graphs containing hyperedges (i.e., edges with higher-order relationships). In addition, we introduce Multi-Object Multi-Actor (MOMA), the first benchmark and dataset dedicated to activity parsing. Lastly, to parse a video, we propose the HyperGraph Activity Parsing (HGAP) network, which outperforms several baselines, including those based on regular graphs and raw video data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 48d6c6bd-4071-491e-b83f-d48879e3a8a7Cited by top-tier papers4
- Lecture Presentations Multimodal Dataset: Towards Understanding Multimodality in Educational VideosDong Won Lee, Chaitanya Ahuja, Paul Pu Liang, Sanika Natu et al.ICCV 2023 · 20 citations
- Task Breakpoint Generation using Origin-Centric Graph in Virtual Reality Recordings for Adaptive PlaybackSelin Choi, Dooyoung Kim, Taewook Ha, Seonji Kim et al.IEEE VR 2026
- Understanding Complexity in VideoQA via Visual Program GenerationCristóbal Eyzaguirre, Igor Vasiljevic, Achal Dave, Jiajun Wu et al.ICML 2025
- Panoptic Video Scene Graph GenerationJingkang Yang, Wenxuan Peng, Xiangtai Li, Zujin Guo et al.CVPR 2023
Builds on11
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding et al.ICCV 2019 · 709 citations
- HACS: Human Action Clips and Segments Dataset for Recognition and Temporal LocalizationHang Zhao, Antonio Torralba, Lorenzo Torresani, Zhicheng YanICCV 2019 · 298 citations
- Action Genome: Actions As Compositions of Spatio-Temporal Scene GraphsJingwei Ji, Ranjay Krishna, Li Fei-Fei, Juan Carlos NieblesCVPR 2020
- Spatio-Temporal Graph for Video Captioning With Knowledge DistillationBoxiao Pan, Haoye Cai, De-An Huang, Kuan-Hui Lee et al.CVPR 2020
Related papers
- Human-Object-Object Interaction: Towards Human-Centric Complex Interaction DetectionMingxuan Zhang, Xiao Wu, Zhaoquan Yuan, Qi He et al.ACM MM 2023 · 6 citations
- Intra- and Inter-Action Understanding via Temporal Action ParsingDian Shao, Yue Zhao, Bo Dai, Dahua LinCVPR 2020
- Prompt-guided Disentangled Representation for Action RecognitionTianci Wu, Guangming Zhu, Jiang Lu, Siyuan Wang et al.NeurIPS 2025 · 1 citation
- Spatio-Temporal Interaction Graph Parsing Networks for Human-Object Interaction RecognitionNing Wang, Guangming Zhu, Liang Zhang, Peiyi Shen et al.ACM MM 2021 · 32 citations
- Ordered Atomic Activity for Fine-grained Interactive Traffic Scenario UnderstandingNakul Agarwal, Yi-Ting ChenICCV 2023 · 7 citations
