FreeMan: Towards Benchmarking 3D Human Pose Estimation Under Real-World Conditions
Jiong Wang, Fengyu Yang, Bingliang Li, Wenbo Gou, Danqi Yan, Ailing Zeng, Yijun Gao, Junle Wang, Yanqing Jing, Ruimao Zhang
Abstract
Estimating the 3D structure of the human body from nat-ural scenes is afundamental aspect of visual perception. 3D human pose estimation is a vital step in advancing fields like AIGC and human-robot interaction, serving as a crucial tech-nique for understanding and interacting with human actions in real-world settings. However, the current datasets, often collected under single laboratory conditions using complex motion capture equipment and unvarying backgrounds, are insufficient. The absence of datasets on variable conditions is stalling the progress of this crucial task. To facilitate the development of 3D pose estimation, we present FreeMan, the first large-scale, multi-view dataset collected under the real-world conditions. FreeMan was captured by synchronizing 8 smartphones across diverse scenarios. It comprises 11M frames from 8000 sequences, viewed from different perspec-tives. These sequences cover 40 subjects across 10 different scenarios, each with varying lighting conditions. We have also established an semi-automated pipeline containing er-ror detection to reduce the workload of manual check and ensure precise annotation. We provide comprehensive eval-uation baselines for a range of tasks, underlining the sig-nificant challenges posed by FreeMan. Further evaluations of standard indoor/outdoor human sensing datasets reveal that FreeMan offers robust representation transferability in real and complex scenes. FreeMan is publicly available at https://wangjiongw.github.io/freeman.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on21
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- LightGlue: Local Feature Matching at Light SpeedPhilipp Lindenberger, Paul-Edouard Sarlin, Marc PollefeysICCV 2023 · 936 citations
Related papers
- M3GYM: A Large-Scale Multimodal Multi-view Multi-person Pose Dataset for Fitness Activity Understanding in Real-world SettingsQingzheng Xu, Ru Cao, Xin Shen, Heming Du et al.CVPR 2025
- Multi-Sensor Large-Scale Dataset for Multi-View 3D ReconstructionOleg Voynov, Gleb Bobrovskikh, Pavel A. Karpyshev, Saveliy Galochkin et al.CVPR 2023
- PoseSyn: Synthesizing Diverse 3D Pose Data from In-the-Wild 2D DataChangHee Yang, Hyeonseop Song, Seokhun Choi, Seungwoo Lee et al.ICCV 2025 · 1 citation
- LiveHPS: LiDAR-Based Scene-Level Human Pose and Shape Estimation in Free EnvironmentYiming Ren, Xiao Han, Chengfeng Zhao, Jingya Wang et al.CVPR 2024 · 14 citations
- Multi-View 3D Human Pose Estimation with Weakly Synchronized ImagesLing Li, Ruiwen Gu, Chongyang Wang, Junliang Xing et al.AAAI 2025 · 3 citations
