Text2Gestures: A Transformer-Based Network for Generating Emotive Body Gestures for Virtual Agents**This work has been supported in part by ARO Grants W911NF1910069 and W911NF1910315, and Intel. Code and additional materials available at: https: //gamma.umd.edu/t2g
Uttaran Bhattacharya, Nicholas Rewkowski, Abhishek Banerjee, Pooja Guhan, Aniket Bera, Dinesh Manocha
摘要
We present Text2Gestures, a transformer-based learning method to interactively generate emotive full-body gestures for virtual agents aligned with natural language text inputs. Our method generates emotionally expressive gestures by utilizing the relevant biomechanical features for body expressions, also known as affective features. We also consider the intended task corresponding to the text and the target virtual agents' intended gender and handedness in our generation pipeline. We train and evaluate our network on the MPI Emotional Body Expressions Database and observe that our network produces state-of-the-art performance in generating gestures for virtual agents aligned with the text for narration or conversation. Our network can generate these gestures at interactive rates on a commodity GPU. We conduct a web-based user study and observe that around 91% of participants indicated our generated gestures to be at least plausible on a five-point Likert Scale. The emotions perceived by the participants from the gestures are also strongly positively correlated with the corresponding intended emotions, with a minimum Pearson coefficient of 0.77 in the valence dimension.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- When XR and AI Meet - A Scoping Review on Extended Reality and Artificial IntelligenceTeresa Hirzle, Florian Müller, Fiona Draxler, Martin Schmitz 等CHI 2023 · 被引用 90 次
- AgentHands: Generating Interactive Hand Gestures for Spatially Grounded Agent Conversations in XRZiyi Liu, David Li, Zhongyi Zhou, David Kim 等CHI 2026 · 被引用 1 次
它引用的顶会 Paper4
- M3ER: Multiplicative Multimodal Emotion Recognition using Facial, Textual, and Speech CuesTrisha Mittal, Uttaran Bhattacharya, Rohan Chandra, Aniket Bera 等AAAI 2020 · 被引用 282 次
- Learning Unseen Emotions from Gestures via Semantically-Conditioned Zero-Shot Perception with Adversarial AutoencodersAbhishek Banerjee, Uttaran Bhattacharya, Aniket BeraAAAI 2022 · 被引用 19 次
- EmotiCon: Context-Aware Multimodal Emotion Recognition Using Frege's PrincipleTrisha Mittal, Pooja Guhan, Uttaran Bhattacharya, Rohan Chandra 等CVPR 2020
- STEP: Spatial Temporal Graph Convolutional Networks for Emotion Perception from GaitsUttaran Bhattacharya, Trisha Mittal, Rohan Chandra, Tanmay Randhavane 等AAAI 2020
相关 Paper
- MIBURI: Towards Expressive Interactive Gesture SynthesisMuhammad Hamza Mughal, Rishabh Dabral, Vera Demberg, Christian TheobaltCVPR 2026 · 被引用 10 次
- BodyFormer: Semantics-guided 3D Body Gesture Synthesis with TransformerKunkun Pang, Dafei Qin, Yingruo Fan, Julian Habekost 等SIGGRAPH 2023 · 被引用 18 次
- Speech2AffectiveGestures: Synthesizing Co-Speech Gestures with Generative Adversarial Affective Expression LearningUttaran Bhattacharya, Elizabeth Childs, Nicholas Rewkowski, Dinesh ManochaACM MM 2021 · 被引用 94 次
- EMAGE: Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture ModelingHaiyang Liu, Zihao Zhu, Giorgio Becherini, Yichen Peng 等CVPR 2024 · 被引用 55 次
- Emotional Speech-Driven 3D Body Animation via Disentangled Latent DiffusionKiran Chhatre, Radek Danecek, Nikos Athanasiou, Giorgio Becherini 等CVPR 2024
