S3: Speech, Script and Scene driven Head and Eye Animation
Yifang Pan, Rishabh Agrawal, Karan Singh
Abstract
We present S 3 , a novel approach to generating expressive, animator-centric 3D head and eye animation of characters in conversation. Given speech audio, a Directorial script and a cinematographic 3D scene as input, we automatically output the animated 3D rotation of each character's head and eyes. S 3 distills animation and psycho-linguistic insights into a novel modular framework for conversational gaze capturing: audio-driven rhythmic head motion; narrative script-driven emblematic head and eye gestures; and gaze trajectories computed from audio-driven gaze focus/aversion and 3D visual scene salience. Our evaluation is four-fold: we quantitatively validate our algorithm against ground truth data and baseline alternatives; we conduct a perceptual study showing our results to compare favourably to prior art; we present examples of animator control and critique of S 3 output; and present a large number of compelling and varied animations of conversational gaze.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8ac37fc9-256e-4bd7-b2b9-1c0089a4a5b1Cited by top-tier papers2
- xADA: Controllable and Expressive Audio-Driven AnimationSarah Taylor, Salvador Medina, Jonathan Windle, Erica Alcusa Sáez et al.SIGGRAPH 2025 · 2 citations
- Gaze and Speech in Multimodal Human-Computer Interaction: A Scoping ReviewAnam Ahmad Khan, Florian Weidner, Jungwoo Rhee, Yasmeen Abdrabou et al.CHI 2026 · 1 citation
Builds on2
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- MeshTalk: 3D Face Animation from Speech using Cross-Modality DisentanglementAlexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre et al.ICCV 2021 · 272 citations
Related papers
- Write-a-speaker: Text-based Emotional and Rhythmic Talking-head GenerationLincheng Li, Suzhen Wang, Zhimeng Zhang, Yu Ding et al.AAAI 2021 · 88 citations
- Talking Together: Synthesizing Co-Located 3D Conversations from AudioMengyi Shan, Shouchieh Chang, Ziqian Bai, Shichen Liu et al.CVPR 2026
- MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled EmbeddingChang Liu, Ye Pan, Chenyang Ding, Susanto Rahardja et al.ACM MM 2025 · 3 citations
- EyEar: Learning Audio Synchronized Human Gaze Trajectory Based on Physics-Informed DynamicsXiaochuan Liu, Xin Cheng, Yuchong Sun, Xiaoxue Wu et al.AAAI 2025 · 2 citations
- VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive InteractionShiying Li, Xingqun Qi, Bingkun Yang, Weile Chen et al.AAAI 2026 · 2 citations
