Unsupervised Audiovisual Synthesis via Exemplar Autoencoders
Kangle Deng, Aayush Bansal, Deva Ramanan
Abstract
We present an unsupervised approach that converts the input speech of any individual into audiovisual streams of potentially-infinitely many output speakers. Our approach builds on simple autoencoders that project out-of-sample data onto the distribution of the training set. We use Exemplar Autoencoders to learn the voice, stylistic prosody, and visual appearance of a specific target exemplar speech. In contrast to existing methods, the proposed approach can be easily extended to an arbitrarily large number of speakers and styles using only 3 minutes of target audio-video data, without requiring any training data for the input speaker. To do so, we learn audiovisual bottleneck representations that capture the structured linguistic content of speech. We outperform prior approaches on both audio and video synthesis, and provide extensive qualitative analysis on our project page -- https://www.cs.cmu.edu/ exemplar-ae/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 726441b1-4772-4371-af12-502479b42e66Cited by top-tier papers2
- Global Prosody Style Transfer Without Text TranscriptionsKaizhi Qian, Yang Zhang, Shiyu Chang, Jinjun Xiong et al.ICML 2021 · 25 citations
- Moûsai: Efficient Text-to-Music Diffusion ModelsFlavio Schneider, Ojasv Kamal, Zhijing Jin, Bernhard SchölkopfACL 2024
Builds on5
- FaceForensics++: Learning to Detect Manipulated Facial ImagesAndreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess et al.ICCV 2019 · 2,966 citations
- EPIC-Fusion: Audio-Visual Temporal Binding for Egocentric Action RecognitionEvangelos Kazakos, Arsha Nagrani, Andrew Zisserman, Dima DamenICCV 2019 · 395 citations
- CNN-Generated Images Are Surprisingly Easy to Spot... for NowSheng-Yu Wang, Oliver Wang, Richard Zhang, Andrew Owens et al.CVPR 2020
- PULSE: Self-Supervised Photo Upsampling via Latent Space Exploration of Generative ModelsSachit Menon, Alexandru Damian, Shijia Hu, Nikhil Ravi et al.CVPR 2020
- CIAGAN: Conditional Identity Anonymization Generative Adversarial NetworksMaxim Maximov, Ismail Elezi, Laura Leal-TaixéCVPR 2020
Related papers
- Improving Zero-Shot Voice Style Transfer via Disentangled Representation LearningSiyang Yuan, Pengyu Cheng, Ruiyi Zhang, Weituo Hao et al.ICLR 2021 · 64 citations
- VisualVoice: Audio-Visual Speech Separation With Cross-Modal ConsistencyRuohan Gao, Kristen GraumanCVPR 2021
- Lip-to-Speech Synthesis for Arbitrary Speakers in the WildSindhu B. Hegde, K. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri et al.ACM MM 2022 · 15 citations
- Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised DisentanglementXueyao Zhang, Xiaohui Zhang, Kainan Peng, Zhenyu Tang et al.ICLR 2025
- ELF: Encoding Speaker-Specific Latent Speech Feature for Speech SynthesisJungil Kong, Junmo Lee, Jeongmin Kim, Beomjeong Kim et al.ICML 2024 · 3 citations
