Continual Learning for Personalized Co-Speech Gesture Generation
Chaitanya Ahuja, Pratik Joshi, Ryo Ishii, Louis-Philippe Morency
Abstract
Co-speech gestures are a key channel of human communication, making them important for personalized chat agents to generate. In the past, gesture generation models assumed that data for each speaker is available all at once, and in large amounts. However in practical scenarios, speaker data comes sequentially and in small amounts as the agent personalizes with more speakers, akin to a continual learning paradigm. While more recent works have shown progress in adapting to low-resource data, they catastrophically forget the gesture styles of initial speakers they were trained on. Also, prior generative continual learning works are not multimodal, making this space less studied. In this paper, we explore this new paradigm and propose C-DiffGAN: an approach that continually learns new speaker gesture styles with only a few minutes of per-speaker data, while retaining previously learnt styles. Inspired by prior continual learning works, C-DiffGAN encourages knowledge retention by 1) generating reminiscences of previous low-resource speaker data, then 2) crossmodally aligning to them to mitigate catastrophic forgetting. We quantitatively demonstrate improved performance and reduced forgetting over strong baselines through standard continual learning measures, reinforced by a qualitative user study that shows that our method produces more natural, style-preserving gestures. Code and videos can be found at https://chahuja.com/cdiffgan
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Mitigating Error Accumulation in Co-Speech Motion Generation via Global Rotation Diffusion and Multi-Level ConstraintsXiangyue Zhang, Jianfang Li, Jianqiang Ren, Jiaxu ZhangAAAI 2026 · 7 citations
- SemGes: Semantics-Aware Co-Speech Gesture Generation Using Semantic Coherence and Relevance LearningLanmiao Liu, Esam Ghaleb, Asli Özyürek, Zerrin YumakICCV 2025 · 4 citations
- Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned RepresentationsSangmin Lee, Bolin Lai, Fiona Ryan, Bikram Boote et al.CVPR 2024
Builds on5
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine et al.NeurIPS 2020 · 2,345 citations
- Image Generation From Small Datasets via Batch Statistics AdaptationAtsuhiro Noguchi, Tatsuya HaradaICCV 2019 · 211 citations
- Low-Resource Adaptation for Personalized Co-Speech Gesture GenerationChaitanya Ahuja, Dong Won Lee, Louis-Philippe MorencyCVPR 2022 · 26 citations
- MineGAN: Effective Knowledge Transfer From GANs to Target Domains With Few ImagesYaxing Wang, Abel Gonzalez-Garcia, David Berga, Luis Herranz et al.CVPR 2020
Related papers
- GAN Memory with No ForgettingYulai Cong, Miaoyun Zhao, Jianqiao Li, Sijia Wang et al.NeurIPS 2020 · 156 citations
- Taming Diffusion Models for Audio-Driven Co-Speech Gesture GenerationLingting Zhu, Xian Liu, Xuanyu Liu, Rui Qian et al.CVPR 2023
- Bring Your Dreams to Life: Continual Text-to-Video CustomizationJiahua Dong, Xudong Wang, Wenqi Liang, Zongyan Han et al.AAAI 2026 · 1 citation
- Class-Incremental Grouping Network for Continual Audio-Visual LearningShentong Mo, Weiguo Pian, Yapeng TianICCV 2023 · 34 citations
- Continual Personalization for Diffusion ModelsYu-Chien Liao, Jr-Jen Chen, Chi-Pin Huang, Ci-Siang Lin et al.ICCV 2025 · 2 citations
