Unsupervised Audiovisual Synthesis via Exemplar Autoencoders
Kangle Deng, Aayush Bansal, Deva Ramanan
摘要
We present an unsupervised approach that converts the input speech of any individual into audiovisual streams of potentially-infinitely many output speakers. Our approach builds on simple autoencoders that project out-of-sample data onto the distribution of the training set. We use Exemplar Autoencoders to learn the voice, stylistic prosody, and visual appearance of a specific target exemplar speech. In contrast to existing methods, the proposed approach can be easily extended to an arbitrarily large number of speakers and styles using only 3 minutes of target audio-video data, without requiring any training data for the input speaker. To do so, we learn audiovisual bottleneck representations that capture the structured linguistic content of speech. We outperform prior approaches on both audio and video synthesis, and provide extensive qualitative analysis on our project page -- https://www.cs.cmu.edu/ exemplar-ae/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Global Prosody Style Transfer Without Text TranscriptionsKaizhi Qian, Yang Zhang, Shiyu Chang, Jinjun Xiong 等ICML 2021 · 被引用 25 次
- Moûsai: Efficient Text-to-Music Diffusion ModelsFlavio Schneider, Ojasv Kamal, Zhijing Jin, Bernhard SchölkopfACL 2024
它引用的顶会 Paper5
- FaceForensics++: Learning to Detect Manipulated Facial ImagesAndreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess 等ICCV 2019 · 被引用 2,966 次
- EPIC-Fusion: Audio-Visual Temporal Binding for Egocentric Action RecognitionEvangelos Kazakos, Arsha Nagrani, Andrew Zisserman, Dima DamenICCV 2019 · 被引用 395 次
- CNN-Generated Images Are Surprisingly Easy to Spot... for NowSheng-Yu Wang, Oliver Wang, Richard Zhang, Andrew Owens 等CVPR 2020
- PULSE: Self-Supervised Photo Upsampling via Latent Space Exploration of Generative ModelsSachit Menon, Alexandru Damian, Shijia Hu, Nikhil Ravi 等CVPR 2020
- CIAGAN: Conditional Identity Anonymization Generative Adversarial NetworksMaxim Maximov, Ismail Elezi, Laura Leal-TaixéCVPR 2020
相关 Paper
- Improving Zero-Shot Voice Style Transfer via Disentangled Representation LearningSiyang Yuan, Pengyu Cheng, Ruiyi Zhang, Weituo Hao 等ICLR 2021 · 被引用 64 次
- VisualVoice: Audio-Visual Speech Separation With Cross-Modal ConsistencyRuohan Gao, Kristen GraumanCVPR 2021
- Lip-to-Speech Synthesis for Arbitrary Speakers in the WildSindhu B. Hegde, K. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri 等ACM MM 2022 · 被引用 15 次
- Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised DisentanglementXueyao Zhang, Xiaohui Zhang, Kainan Peng, Zhenyu Tang 等ICLR 2025
- ELF: Encoding Speaker-Specific Latent Speech Feature for Speech SynthesisJungil Kong, Junmo Lee, Jeongmin Kim, Beomjeong Kim 等ICML 2024 · 被引用 3 次
