FADA: Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG Distillation
Tianyun Zhong, Chao Liang, Jianwen Jiang, Gaojie Lin, Jiaqi Yang, Zhou Zhao
Abstract
Diffusion-based audio-driven talking avatar methods have recently gained attention for their high-fidelity, vivid, and expressive results. However, their slow inference speed limits practical applications. Despite the development of various distillation techniques for diffusion models, we found that naive diffusion distillation methods do not yield satisfactory results. Distilled models exhibit reduced robustness with open-set input images and a decreased correlation between audio and video compared to teacher models, undermining the advantages of diffusion models. To address this, we propose FADA (Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG Distillation). We first designed a mixed-supervised loss to leverage data of varying quality and enhance the overall model capability as well as robustness. Additionally, we propose a multi-CFG distillation with learnable tokens to utilize the correlation between audio and reference image conditions, reducing the threefold inference runs caused by multi-CFG with acceptable quality degradation. Extensive experiments across multiple datasets show that FADA generates vivid videos comparable to recent diffusion model-based methods while achieving an NFE speedup of 4.17-12.5 times. Demos are available at our webpage https://fadavatar.github.io .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 56f2811f-1a85-4a00-af3c-e0cab2f32fabCited by top-tier papers2
- Instilling an Active Mind in Avatars via Cognitive SimulationJianwen Jiang, Weihong Zeng, Zerong Zheng, Jiaqi Yang et al.ICLR 2026 · 26 citations
- OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation ModelsGaojie Lin, Jianwen Jiang, Jiaqi Yang, Zerong Zheng et al.ICCV 2025 · 11 citations
Builds on29
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisYudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu et al.ICCV 2021 · 510 citations
Related papers
- Adaptive Video Distillation: Mitigating Oversaturation and Temporal Collapse in Few-Step GenerationYuyang You, Yongzhi Li, Jiahui Li, Yadong Mu et al.CVPR 2026 · 7 citations
- SwiftVideo: A Unified Framework for Few-Step Video Generation Through Trajectory-Distribution AlignmentYanxiao Sun, Jiafu Wu, Yun Cao, Chengming Xu et al.AAAI 2026 · 6 citations
- Simple and Fast Distillation of Diffusion ModelsZhenyu Zhou, Defang Chen, Can Wang, Chun Chen et al.NeurIPS 2024 · 44 citations
- FVGen: Accelerating Novel-View Synthesis with Adversarial Video Diffusion DistillationWenbin Teng, Gonglin Chen, Haiwei Chen, Yajie ZhaoICCV 2025
- Distillation of Discrete Diffusion through Dimensional CorrelationsSatoshi Hayakawa, Yuhta Takida, Masaaki Imaizumi, Hiromi Wakaki et al.ICML 2025
