Continual Learning for Personalized Co-Speech Gesture Generation
Chaitanya Ahuja, Pratik Joshi, Ryo Ishii, Louis-Philippe Morency
摘要
Co-speech gestures are a key channel of human communication, making them important for personalized chat agents to generate. In the past, gesture generation models assumed that data for each speaker is available all at once, and in large amounts. However in practical scenarios, speaker data comes sequentially and in small amounts as the agent personalizes with more speakers, akin to a continual learning paradigm. While more recent works have shown progress in adapting to low-resource data, they catastrophically forget the gesture styles of initial speakers they were trained on. Also, prior generative continual learning works are not multimodal, making this space less studied. In this paper, we explore this new paradigm and propose C-DiffGAN: an approach that continually learns new speaker gesture styles with only a few minutes of per-speaker data, while retaining previously learnt styles. Inspired by prior continual learning works, C-DiffGAN encourages knowledge retention by 1) generating reminiscences of previous low-resource speaker data, then 2) crossmodally aligning to them to mitigate catastrophic forgetting. We quantitatively demonstrate improved performance and reduced forgetting over strong baselines through standard continual learning measures, reinforced by a qualitative user study that shows that our method produces more natural, style-preserving gestures. Code and videos can be found at https://chahuja.com/cdiffgan
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Mitigating Error Accumulation in Co-Speech Motion Generation via Global Rotation Diffusion and Multi-Level ConstraintsXiangyue Zhang, Jianfang Li, Jianqiang Ren, Jiaxu ZhangAAAI 2026 · 被引用 7 次
- SemGes: Semantics-Aware Co-Speech Gesture Generation Using Semantic Coherence and Relevance LearningLanmiao Liu, Esam Ghaleb, Asli Özyürek, Zerrin YumakICCV 2025 · 被引用 4 次
- Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned RepresentationsSangmin Lee, Bolin Lai, Fiona Ryan, Bikram Boote 等CVPR 2024
它引用的顶会 Paper5
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine 等NeurIPS 2020 · 被引用 2,345 次
- Image Generation From Small Datasets via Batch Statistics AdaptationAtsuhiro Noguchi, Tatsuya HaradaICCV 2019 · 被引用 211 次
- Low-Resource Adaptation for Personalized Co-Speech Gesture GenerationChaitanya Ahuja, Dong Won Lee, Louis-Philippe MorencyCVPR 2022 · 被引用 26 次
- MineGAN: Effective Knowledge Transfer From GANs to Target Domains With Few ImagesYaxing Wang, Abel Gonzalez-Garcia, David Berga, Luis Herranz 等CVPR 2020
相关 Paper
- GAN Memory with No ForgettingYulai Cong, Miaoyun Zhao, Jianqiao Li, Sijia Wang 等NeurIPS 2020 · 被引用 156 次
- Taming Diffusion Models for Audio-Driven Co-Speech Gesture GenerationLingting Zhu, Xian Liu, Xuanyu Liu, Rui Qian 等CVPR 2023
- Bring Your Dreams to Life: Continual Text-to-Video CustomizationJiahua Dong, Xudong Wang, Wenqi Liang, Zongyan Han 等AAAI 2026 · 被引用 1 次
- Class-Incremental Grouping Network for Continual Audio-Visual LearningShentong Mo, Weiguo Pian, Yapeng TianICCV 2023 · 被引用 34 次
- Continual Personalization for Diffusion ModelsYu-Chien Liao, Jr-Jen Chen, Chi-Pin Huang, Ci-Siang Lin 等ICCV 2025 · 被引用 2 次
