Listening Human Behavior: 3D Human Pose Estimation with Acoustic Signals
Yuto Shibata, Yutaka Kawashima, Mariko Isogawa, Go Irie, Akisato Kimura, Yoshimitsu Aoki
摘要
Given only acoustic signals without any high-level information, such as voices or sounds of scenes/actions, how much can we infer about the behavior of humans? Unlike existing methods, which suffer from privacy issues because they use signals that include human speech or the sounds of specific actions, we explore how low-level acoustic signals can provide enough clues to estimate 3D human poses by active acoustic sensing with a single pair of microphones and loudspeakers (see Fig. 1). This is a challenging task since sound is much more diffractive than other signals and therefore covers up the shape of objects in a scene. Accordingly, we introduce a framework that encodes multichannel audio features into 3D human poses. Aiming to capture subtle sound changes to reveal detailed pose information, we explicitly extract phase features from the acoustic signals together with typical spectrum features and feed them into our human pose estimation network. Also, we show that reflected or diffracted sounds are easily influenced by subjects' physique differences e.g., height and muscularity, which deteriorates prediction accuracy. We reduce these gaps by using a subject discriminator to improve accuracy. Our experiments suggest that with the use of only low-dimensional acoustic information, our method outperforms baseline methods. The datasets and codes used in this project will be publicly available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- Towards 3D human pose construction using wifiWenjun Jiang, Hongfei Xue, Chenglin Miao, Shiyang Wang 等MobiCom 2020 · 被引用 282 次
- Audio2Gestures: Generating Diverse Gestures from Speech Audio with Conditional Variational AutoencodersJing Li, Di Kang, Wenjie Pei, Xuefei Zhe 等ICCV 2021 · 被引用 144 次
- Making the Invisible Visible: Action Recognition Through Walls and OcclusionsTianhong Li, Lijie Fan, Mingmin Zhao, Yingcheng Liu 等ICCV 2019 · 被引用 126 次
- Audio-Visual Floorplan ReconstructionSenthil Purushwalkam, Sebastià Vicenc Amengual Garí, Vamsi Krishna Ithapu, Carl Schissler 等ICCV 2021 · 被引用 45 次
- GLAVNet: Global-Local Audio-Visual Cues for Fine-Grained Material RecognitionFengmin Shi, Jie Guo, Haonan Zhang, Shan Yang 等CVPR 2021
相关 Paper
- Sounding Bodies: Modeling 3D Spatial Sound of Humans Using Body Pose and AudioXudong Xu, Dejan Markovic, Jacob Sandakly, Todd Keebler 等NeurIPS 2023 · 被引用 9 次
- We Hear Your PACE: Passive Acoustic Localization of Multiple Walking PersonsChao Cai, Henglin Pu, Peng Wang, Zhe Chen 等UbiComp 2021 · 被引用 59 次
- LASense: Pushing the Limits of Fine-grained Activity Sensing Using Acoustic SignalsDong Li, Jialin Liu, Sunghoon Ivan Lee, Jie XiongUbiComp 2022 · 被引用 36 次
- mmHolmes: Amodal Millimeter-wave Sensing by Understanding Human KineticsKun Liang, Yaxuan Li, He Hao, Hongli Zeng 等UbiComp 2025 · 被引用 2 次
- ExpresSense: Exploring a Standalone Smartphone to Sense Engagement of Users from Facial Expressions Using Acoustic SensingPragma Kar, Shyamvanshikumar Singh, Avijit Mandal, Samiran Chattopadhyay 等CHI 2023 · 被引用 7 次
