Text2Gestures: A Transformer-Based Network for Generating Emotive Body Gestures for Virtual Agents**This work has been supported in part by ARO Grants W911NF1910069 and W911NF1910315, and Intel. Code and additional materials available at: https: //gamma.umd.edu/t2g
Uttaran Bhattacharya, Nicholas Rewkowski, Abhishek Banerjee, Pooja Guhan, Aniket Bera, Dinesh Manocha
Abstract
We present Text2Gestures, a transformer-based learning method to interactively generate emotive full-body gestures for virtual agents aligned with natural language text inputs. Our method generates emotionally expressive gestures by utilizing the relevant biomechanical features for body expressions, also known as affective features. We also consider the intended task corresponding to the text and the target virtual agents' intended gender and handedness in our generation pipeline. We train and evaluate our network on the MPI Emotional Body Expressions Database and observe that our network produces state-of-the-art performance in generating gestures for virtual agents aligned with the text for narration or conversation. Our network can generate these gestures at interactive rates on a commodity GPU. We conduct a web-based user study and observe that around 91% of participants indicated our generated gestures to be at least plausible on a five-point Likert Scale. The emotions perceived by the participants from the gestures are also strongly positively correlated with the corresponding intended emotions, with a minimum Pearson coefficient of 0.77 in the valence dimension.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5d9f964c-7c3c-49fe-9b75-57f9b84fe4f9Cited by top-tier papers2
- When XR and AI Meet - A Scoping Review on Extended Reality and Artificial IntelligenceTeresa Hirzle, Florian Müller, Fiona Draxler, Martin Schmitz et al.CHI 2023 · 90 citations
- AgentHands: Generating Interactive Hand Gestures for Spatially Grounded Agent Conversations in XRZiyi Liu, David Li, Zhongyi Zhou, David Kim et al.CHI 2026 · 1 citation
Builds on4
- M3ER: Multiplicative Multimodal Emotion Recognition using Facial, Textual, and Speech CuesTrisha Mittal, Uttaran Bhattacharya, Rohan Chandra, Aniket Bera et al.AAAI 2020 · 282 citations
- Learning Unseen Emotions from Gestures via Semantically-Conditioned Zero-Shot Perception with Adversarial AutoencodersAbhishek Banerjee, Uttaran Bhattacharya, Aniket BeraAAAI 2022 · 19 citations
- EmotiCon: Context-Aware Multimodal Emotion Recognition Using Frege's PrincipleTrisha Mittal, Pooja Guhan, Uttaran Bhattacharya, Rohan Chandra et al.CVPR 2020
- STEP: Spatial Temporal Graph Convolutional Networks for Emotion Perception from GaitsUttaran Bhattacharya, Trisha Mittal, Rohan Chandra, Tanmay Randhavane et al.AAAI 2020
Related papers
- MIBURI: Towards Expressive Interactive Gesture SynthesisMuhammad Hamza Mughal, Rishabh Dabral, Vera Demberg, Christian TheobaltCVPR 2026 · 10 citations
- BodyFormer: Semantics-guided 3D Body Gesture Synthesis with TransformerKunkun Pang, Dafei Qin, Yingruo Fan, Julian Habekost et al.SIGGRAPH 2023 · 18 citations
- Speech2AffectiveGestures: Synthesizing Co-Speech Gestures with Generative Adversarial Affective Expression LearningUttaran Bhattacharya, Elizabeth Childs, Nicholas Rewkowski, Dinesh ManochaACM MM 2021 · 94 citations
- EMAGE: Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture ModelingHaiyang Liu, Zihao Zhu, Giorgio Becherini, Yichen Peng et al.CVPR 2024 · 55 citations
- Emotional Speech-Driven 3D Body Animation via Disentangled Latent DiffusionKiran Chhatre, Radek Danecek, Nikos Athanasiou, Giorgio Becherini et al.CVPR 2024
