Style Equalization: Unsupervised Learning of Controllable Generative Sequence Models
Jen-Hao Rick Chang, Ashish Shrivastava, Hema Koppula, Xiaoshuai Zhang, Oncel Tuzel
摘要
Controllable generative sequence models with the capability to extract and replicate the style of specific examples enable many applications, including narrating audiobooks in different voices, auto-completing and auto-correcting written handwriting, and generating missing training samples for downstream recognition tasks. However, under an unsupervised-style setting, typical training algorithms for controllable sequence generative models suffer from the training-inference mismatch, where the same sample is used as content and style input during training but unpaired samples are given during inference. In this paper, we tackle the training-inference mismatch encountered during unsupervised learning of controllable generative sequence models. The proposed method is simple yet effective, where we use a style transformation module to transfer target style information into an unrelated style input. This method enables training using unpaired content and style samples and thereby mitigate the training-inference mismatch. We apply style equalization to text-to-speech and text-to-handwriting synthesis on three datasets. We conduct thorough evaluation, including both quantitative and qualitative user studies. Our results show that by mitigating the training-inference mismatch with the proposed style equalization, we achieve style replication scores comparable to real data in our user studies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Alias-Free Generative Adversarial NetworksTero Karras, Miika Aittala, Samuli Laine, Erik Härkönen 等NeurIPS 2021 · 被引用 2,126 次
- Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment SearchJaehyeon Kim, Sungwon Kim, Jungil Kong, Sungroh YoonNeurIPS 2020 · 被引用 663 次
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin 等ICLR 2021 · 被引用 513 次
- How much Position Information Do Convolutional Neural Networks Encode?Md. Amirul Islam, Sen Jia, Neil D. B. BruceICLR 2020 · 被引用 392 次
- Unsupervised Speech Decomposition via Triple Information BottleneckKaizhi Qian, Yang Zhang, Shiyu Chang, Mark Hasegawa-Johnson 等ICML 2020 · 被引用 210 次
相关 Paper
- Exploring Contextual Word-level Style Relevance for Unsupervised Style TransferChulun Zhou, Liangyu Chen, Jiachen Liu, Xinyan Xiao 等ACL 2020 · 被引用 34 次
- Autoregressive Stylized Motion Synthesis With Generative FlowYu-Hui Wen, Zhipeng Yang, Hongbo Fu, Lin Gao 等CVPR 2021
- Improving Zero-Shot Voice Style Transfer via Disentangled Representation LearningSiyang Yuan, Pengyu Cheng, Ruiyi Zhang, Weituo Hao 等ICLR 2021 · 被引用 64 次
- Plug and Play Autoencoders for Conditional Text GenerationFlorian Mai, Nikolaos Pappas, Ivan Montero, Noah A. Smith 等EMNLP 2020 · 被引用 24 次
- Collaborative Learning of Bidirectional Decoders for Unsupervised Text Style TransferYun Ma, Yangbin Chen, Xudong Mao, Qing LiEMNLP 2021 · 被引用 6 次
