Learning from Synthetic Human Group Activities
Che-Jui Chang, Danrui Li, Deep Patel, Parth Goel, Honglu Zhou, Seonghyeon Moon, Samuel S. Sohn, Sejong Yoon, Vladimir Pavlovic, Mubbasir Kapadia
Abstract
The study of complex human interactions and group activities has become a focal point in human-centric computer vision. However, progress in related tasks is often hindered by the challenges of obtaining large-scale labeled datasets from real-world scenarios. To address the limitation, we introduce M 3 Act, a synthetic data generator for multi-view multi-group multi-person human atomic actions and group activities. Powered by Unity Engine, M 3 Act features multiple semantic groups, highly diverse and photorealistic images, and a comprehensive set of annotations, which facilitates the learning of human-centered tasks across singleperson, multi-person, and multi-group conditions. We demonstrate the advantages of M 3 Act across three core experiments. The results suggest our synthetic dataset can significantly improve the performance of several downstream methods and replace real-world datasets to reduce cost. Notably, M 3 Act improves the state-of-the-art MOTRv2 on DanceTrack dataset, leading to a hop on the leaderboard from 10 th to 2 nd place. Moreover, M 3 Act opens new research for controllable 3D group activity generation. We define multiple metrics and propose a competitive baseline for the novel task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c5bd94b0-e42d-4838-8a53-21623498810dCited by top-tier papers2
- SVLTA: Benchmarking Vision-Language Temporal Alignment via Synthetic Video SituationHao Du, Bo Wu, Yan Lu, Zhendong MaoCVPR 2025
- Motions as Queries: One-Stage Multi-Person Holistic Human Motion CaptureKenkun Liu, Yurong Fu, Weihao Yuan, Jing Lin et al.CVPR 2025
Builds on18
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- Action2Motion: Conditioned Generation of 3D Human MotionsChuan Guo, Xinxin Zuo, Sen Wang, Shihao Zou et al.ACM MM 2020 · 394 citations
- DanceTrack: Multi-Object Tracking in Uniform Appearance and Diverse MotionPeize Sun, Jinkun Cao, Yi Jiang, Zehuan Yuan et al.CVPR 2022 · 305 citations
- SHIFT: A Synthetic Driving Dataset for Continuous Multi-Task Domain AdaptationTao Sun, Mattia Segù, Janis Postels, Yuxuan Wang et al.CVPR 2022 · 174 citations
- Human Motion Diffusion ModelGuy Tevet, Sigal Raab, Brian Gordon, Yonatan Shafir et al.ICLR 2023 · 167 citations
Related papers
- MOTSynth: How Can Synthetic Data Help Pedestrian Detection and Tracking?Matteo Fabbri, Guillem Brasó, Gianluca Maugeri, Orcun Cetintas et al.ICCV 2021 · 128 citations
- DNA-Rendering: A Diverse Neural Actor Repository for High-Fidelity Human-centric RenderingWei Cheng, Ruixiang Chen, Siming Fan, Wanqi Yin et al.ICCV 2023 · 106 citations
- ActFormer: A GAN-based Transformer towards General Action-Conditioned 3D Human Motion GenerationLiang Xu, Ziyang Song, Dongliang Wang, Jing Su et al.ICCV 2023 · 100 citations
- 3D Segmentation of Humans in Point Clouds with Synthetic DataAyça Takmaz, Jonas Schult, Irem Kaftan, Mertcan Akçay et al.ICCV 2023 · 31 citations
- MVHumanNet: A Large-Scale Dataset of Multi-View Daily Dressing Human CapturesZhangyang Xiong, Chenghong Li, Kenkun Liu, Hongjie Liao et al.CVPR 2024 · 16 citations
