Lune

ICCV2025顶会

Head2Body: Body Pose Generation from Multi-Sensory Head-Mounted Inputs

Minh Tran, Hongda Mao, Qingshuang Chen, Yelin Kim

2025年份
1被引次数

摘要

Generating body pose from head-mounted, egocentric inputs is essential for immersive VR/AR and assistive technologies, as it supports more natural interactions. However, the task is challenging due to limited visibility of body parts in first-person views and the sparseness of sensory data, with only a single device placed on the head. To address these challenges, we introduce Head2Body, a novel framework for body pose estimation that effectively combines head-IMU and egocentric visual data. First, we introduce a pretrained IMU encoder, trained on over 1,700 hours of Ego4D IMU data from head-mounted devices, to better capture detailed temporal motion cues given limited labeled egocentric pose data. For visual processing, we leverage large vision-language models (LVLMs) to segment body parts that appear sporadically in video frames to improve visual feature extraction. To better guide pose generation from sparse head-mounted signals, we incorporate a residual Vector Quantized Variational Autoencoder (VQ-VAE) to represent poses with discrete tokens, capturing high-frequency motion patterns and improving over direct continuous regression, which often lacks structure and temporal consistency. Our experiments demonstrate the effectiveness of the proposed approach, yielding 6-13% gains over state-of-the-art baselines on three datasets: AMASS, KinPoly, and EgoExo4D. By capturing subtle temporal dynamics and leveraging complementary sensory data, our approach advances accurate egocentric body pose estimation and sets a new benchmark for multi-modal, first-person motion tracking.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper23

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖