Context-Aware Head-and-Eye Motion Generation with Diffusion Model
Yuxin Shen, Manjie Xu, Wei Liang
Abstract
In humanity’s ongoing quest to craft natural and realistic avatars within virtual environments, the generation of authentic eye gaze behaviors stands paramount. Eye gaze not only serves as a primary non-verbal communication cue, but it also reflects cognitive processes, intent, and attentiveness, making it a crucial element in ensuring immersive interactions. However, automatically generating these intricate gaze behaviors presents significant challenges. Traditional methods can be both time-consuming and lack the precision to align gaze behaviors with the intricate nuances of the environment in which the avatar resides. To overcome these challenges, we introduce a novel two-stage approach to generate context-aware head-and-eye motions across diverse scenes. By harnessing the capabilities of advanced diffusion models, our approach adeptly produces contextually appropriate eye gaze points, further leading to the generation of natural head-and-eye movements. Utilizing Head-Mounted Display (HMD) eye-tracking technology, we also present a comprehensive dataset, which captures human eye gaze behaviors in tandem with associated scene features. We show that our approach consistently delivers intuitive and lifelike head-and-eye motions and demonstrates superior performance in terms of motion fluidity, alignment with contextual cues, and overall user satisfaction.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c0ead9f5-dc20-4029-a873-529299e729eaRelated papers
- TextGaze: Gaze-Controllable Face Generation with Natural LanguageHengfei Wang, Zhongqun Zhang, Yihua Cheng, Hyung Jin ChangACM MM 2024 · 3 citations
- The eyes have it: an integrated eye and face model for photorealistic facial animationGabriel Schwartz, Shih-En Wei, Te-Li Wang, Stephen Lombardi et al.SIGGRAPH 2020 · 54 citations
- Talking Together: Synthesizing Co-Located 3D Conversations from AudioMengyi Shan, Shouchieh Chang, Ziqian Bai, Shichen Liu et al.CVPR 2026
- EyeNeRF: a hybrid representation for photorealistic synthesis, animation and relighting of human eyesGengyan Li, Abhimitra Meka, Franziska Mueller, Marcel C. Bühler et al.SIGGRAPH 2022 · 39 citations
- PrivateEyes: Gaze-Preserving Anonymization for Data SharingSurabhi Gupta, Dinesh Prabhu Muthumariappan, Biplab Ch Das, Anoop Kolar Rajagopal et al.CVPR 2026
