Educating Text Autoencoders: Latent Representation Guidance via Denoising
Tianxiao Shen, Jonas Mueller, Regina Barzilay, Tommi S. Jaakkola
Abstract
Generative autoencoders offer a promising approach for controllable text generation by leveraging their latent sentence representations. However, current models struggle to maintain coherent latent spaces required to perform meaningful text manipulations via latent vector operations. Specifically, we demonstrate by example that neural encoders do not necessarily map similar sentences to nearby latent vectors. A theoretical explanation for this phenomenon establishes that highcapacity autoencoders can learn an arbitrary mapping between sequences and associated latent representations. To remedy this issue, we augment adversarial autoencoders with a denoising objective where original sentences are reconstructed from perturbed versions (referred to as DAAE). We prove that this simple modification guides the latent space geometry of the resulting model by encouraging the encoder to map similar texts to similar latent representations. In empirical comparisons with various types of autoencoders, our model provides the best trade-off between generation quality and reconstruction capacity. Moreover, the improved geometry of the DAAE latent space enables zero-shot text style transfer via simple latent vector arithmetic. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 752f2b86-9fee-4587-b0c0-bd57a82b11afCited by top-tier papers8
- Data Augmentation for Cross-Domain Named Entity RecognitionShuguang Chen, Gustavo Aguilar, Leonardo Neves, Thamar SolorioEMNLP 2021 · 39 citations
- Plug and Play Autoencoders for Conditional Text GenerationFlorian Mai, Nikolaos Pappas, Ivan Montero, Noah A. Smith et al.EMNLP 2020 · 24 citations
- Composable Text Controls in Latent Space with ODEsGuangyi Liu, Zeyu Feng, Yuan Gao, Zichao Yang et al.EMNLP 2023 · 12 citations
- Unified Generation, Reconstruction, and Representation: Generalized Diffusion with Adaptive Latent Encoding-DecodingGuangyi Liu, Yu Wang, Zeyu Feng, Qiyu Wu et al.ICML 2024 · 9 citations
- SALSA: Semantically-Aware Latent Space AutoencoderKathryn E. Kirchoff, Travis Maxfield, Alexander Tropsha, Shawn M. GomezAAAI 2024 · 3 citations
Related papers
- On Variational Learning of Controllable Representations for Text without SupervisionPeng Xu, Jackie Chi Kit Cheung, Yanshuai CaoICML 2020 · 69 citations
- So Different Yet So Alike! Constrained Unsupervised Text Style TransferAbhinav Ramesh Kashyap, Devamanyu Hazarika, Min-Yen Kan, Roger Zimmermann et al.ACL 2022
- Revision in Continuous Space: Unsupervised Text Style Transfer without Adversarial LearningDayiheng Liu, Jie Fu, Yidan Zhang, Chris Pal et al.AAAI 2020 · 53 citations
- Adapting Language Models for Non-Parallel Author-Stylized RewritingBakhtiyar Syed, Gaurav Verma, Balaji Vasan Srinivasan, Anandhavelu Natarajan et al.AAAI 2020 · 53 citations
- Adversarial Latent AutoencodersStanislav Pidhorskyi, Donald A. Adjeroh, Gianfranco DorettoCVPR 2020
