InstructAvatar: Text-Guided Emotion and Motion Control for Avatar Generation
Yuchi Wang, Junliang Guo, Jianhong Bai, Runyi Yu, Tianyu He, Xu Tan, Xu Sun, Jiang Bian
Abstract
Recent talking avatar generation models have made strides in achieving realistic and accurate lip synchronization with the audio, but often fall short in controlling and conveying detailed expressions and emotions of the avatar, making the generated video less vivid and controllable. In this paper, we propose a text-guided approach for generating emotionally expressive 2D avatars, offering fine-grained control, improved interactivity, and generalizability to the resulting video. Our framework, named InstructAvatar, leverages a natural language interface to control the emotion as well as the facial motion of avatars. Technically, we utilize GPT-4V to design an automatic annotation pipeline, constructing an instruction-video paired training dataset. This is combined with a novel two-branch diffusion-based generator to predict avatars using both audio and text instructions simultaneously. Experimental results demonstrate that InstructAvatar produces results that align well with both conditions, and outperforms existing methods in fine-grained emotion control, lip-sync quality, and naturalness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 966b95a2-f17a-46f8-aadc-36efa2b4afc1Cited by top-tier papers12
- MegActor-Sigma: Unlocking Flexible Mixed-Modal Control in Portrait Animation with Diffusion TransformerShurong Yang, Huadong Li, Juhao Wu, Minhao Jing et al.AAAI 2025 · 12 citations
- RadarLLM: Empowering Large Language Models to Understand Human Motion from Millimeter-wave Point Cloud SequenceZengyuan Lai, Jiarui Yang, Songpengcheng Xia, Lizhou Lin et al.AAAI 2026 · 5 citations
- AV-Flow: Transforming Text to Audio-Visual Human-Like InteractionsAggelina Chatziagapi, Louis-Philippe Morency, Hongyu Gong, Michael Zollhöfer et al.ICCV 2025 · 3 citations
- AUHead: Realistic Emotional Talking Head Generation via Action Units ControlJiayi Lyu, Leigang Qu, Wenjing Zhang, Hanyu Jiang et al.ICLR 2026 · 2 citations
- DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D PosesYatian Pang, Bin Zhu, Bin Lin, Mingzhe Zheng et al.ICCV 2025 · 2 citations
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled EmbeddingChang Liu, Ye Pan, Chenyang Ding, Susanto Rahardja et al.ACM MM 2025 · 3 citations
- ConsistentAvatar: Learning to Diffuse Fully Consistent Talking Head Avatar with Temporal GuidanceHaijie Yang, Zhenyu Zhang, Hao Tang, Jianjun Qian et al.ACM MM 2024 · 3 citations
- PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face GenerationBaiqin Wang, Xiangyu Zhu, Fan Shen, Hao Xu et al.CVPR 2026 · 8 citations
- IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-SpeechSiyi Zhou, Yiquan Zhou, Yi He, Xun Zhou et al.AAAI 2026 · 63 citations
- Style2Talker: High-Resolution Talking Head Generation with Emotion Style and Art StyleShuai Tan, Bin Ji, Ye PanAAAI 2024
