Listening Human Behavior: 3D Human Pose Estimation with Acoustic Signals
Yuto Shibata, Yutaka Kawashima, Mariko Isogawa, Go Irie, Akisato Kimura, Yoshimitsu Aoki
Abstract
Given only acoustic signals without any high-level information, such as voices or sounds of scenes/actions, how much can we infer about the behavior of humans? Unlike existing methods, which suffer from privacy issues because they use signals that include human speech or the sounds of specific actions, we explore how low-level acoustic signals can provide enough clues to estimate 3D human poses by active acoustic sensing with a single pair of microphones and loudspeakers (see Fig. 1). This is a challenging task since sound is much more diffractive than other signals and therefore covers up the shape of objects in a scene. Accordingly, we introduce a framework that encodes multichannel audio features into 3D human poses. Aiming to capture subtle sound changes to reveal detailed pose information, we explicitly extract phase features from the acoustic signals together with typical spectrum features and feed them into our human pose estimation network. Also, we show that reflected or diffracted sounds are easily influenced by subjects' physique differences e.g., height and muscularity, which deteriorates prediction accuracy. We reduce these gaps by using a subject discriminator to improve accuracy. Our experiments suggest that with the use of only low-dimensional acoustic information, our method outperforms baseline methods. The datasets and codes used in this project will be publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Towards 3D human pose construction using wifiWenjun Jiang, Hongfei Xue, Chenglin Miao, Shiyang Wang et al.MobiCom 2020 · 282 citations
- Audio2Gestures: Generating Diverse Gestures from Speech Audio with Conditional Variational AutoencodersJing Li, Di Kang, Wenjie Pei, Xuefei Zhe et al.ICCV 2021 · 144 citations
- Making the Invisible Visible: Action Recognition Through Walls and OcclusionsTianhong Li, Lijie Fan, Mingmin Zhao, Yingcheng Liu et al.ICCV 2019 · 126 citations
- Audio-Visual Floorplan ReconstructionSenthil Purushwalkam, Sebastià Vicenc Amengual Garí, Vamsi Krishna Ithapu, Carl Schissler et al.ICCV 2021 · 45 citations
- GLAVNet: Global-Local Audio-Visual Cues for Fine-Grained Material RecognitionFengmin Shi, Jie Guo, Haonan Zhang, Shan Yang et al.CVPR 2021
Related papers
- Sounding Bodies: Modeling 3D Spatial Sound of Humans Using Body Pose and AudioXudong Xu, Dejan Markovic, Jacob Sandakly, Todd Keebler et al.NeurIPS 2023 · 9 citations
- We Hear Your PACE: Passive Acoustic Localization of Multiple Walking PersonsChao Cai, Henglin Pu, Peng Wang, Zhe Chen et al.UbiComp 2021 · 59 citations
- LASense: Pushing the Limits of Fine-grained Activity Sensing Using Acoustic SignalsDong Li, Jialin Liu, Sunghoon Ivan Lee, Jie XiongUbiComp 2022 · 36 citations
- mmHolmes: Amodal Millimeter-wave Sensing by Understanding Human KineticsKun Liang, Yaxuan Li, He Hao, Hongli Zeng et al.UbiComp 2025 · 2 citations
- ExpresSense: Exploring a Standalone Smartphone to Sense Engagement of Users from Facial Expressions Using Acoustic SensingPragma Kar, Shyamvanshikumar Singh, Avijit Mandal, Samiran Chattopadhyay et al.CHI 2023 · 7 citations
