A Novel Estimator of Mutual Information for Learning to Disentangle Textual Representations
Pierre Colombo, Pablo Piantanida, Chloé Clavel
Abstract
Learning disentangled representations of textual data is essential for many natural language tasks such as fair classification, style transfer and sentence generation, among others. The existent dominant approaches in the context of text data either rely on training an adversary (discriminator) that aims at making attribute values difficult to be inferred from the latent code or rely on minimising variational bounds of the mutual information between latent code and the value attribute. However, the available methods suffer of the impossibility to provide a fine-grained control of the degree (or force) of disentanglement. In contrast to adversarial methods, which are remarkably simple, although the adversary seems to be performing perfectly well during the training phase, after it is completed a fair amount of information about the undesired attribute still remains. This paper introduces a novel variational upper bound to the mutual information between an attribute and the latent code of an encoder. Our bound aims at controlling the approximation error via the Renyi's divergence, leading to both better disentangled representations and in particular, a precise control of the desirable degree of disentanglement than state-of-the-art methods proposed for textual data. Furthermore, it does not suffer from the degeneracy of other losses in multi-class scenarios. We show the superiority of this method on fair classification and on textual style transfer tasks. Additionally, we provide new insights illustrating various trade-offs in style transfer when attempting to learn disentangled representations and quality of the generated sentence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e667790-61dc-45a7-a016-122b705c0d83Cited by top-tier papers10
- A Differential Entropy Estimator for Training Neural NetworksGeorg Pichler, Pierre Jean A. Colombo, Malik Boudiaf, Günther Koliander et al.ICML 2022 · 28 citations
- Beyond Mahalanobis Distance for Textual OOD DetectionPierre Colombo, Eduardo Dadalto Câmara Gomes, Guillaume Staerman, Nathan Noiry et al.NeurIPS 2022 · 24 citations
- DARE: Disentanglement-Augmented Rationale ExtractionLinan Yue, Qi Liu, Yichao Du, Yanqing An et al.NeurIPS 2022 · 24 citations
- Learning Disentangled Representations of Negation and UncertaintyJake Vasilakes, Chrysoula Zerva, Makoto Miwa, Sophia AnaniadouACL 2022 · 22 citations
- What are the best Systems? New Perspectives on NLP BenchmarkingPierre Colombo, Nathan Noiry, Ekhine Irurozki, Stéphan ClémençonNeurIPS 2022 · 20 citations
Builds on4
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- CLUB: A Contrastive Log-ratio Upper Bound of Mutual InformationPengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu et al.ICML 2020 · 512 citations
- Guiding Attention in Sequence-to-Sequence Models for Dialogue Act PredictionPierre Colombo, Emile Chapuis, Matteo Manica, Emmanuel Vignon et al.AAAI 2020 · 69 citations
- Heavy-tailed Representations, Text Polarity Classification & Data AugmentationHamid Jalalzai, Pierre Colombo, Chloé Clavel, Éric Gaussier et al.NeurIPS 2020 · 33 citations
Related papers
- Improving Disentangled Text Representation Learning with Information-Theoretic GuidancePengyu Cheng, Martin Renqiang Min, Dinghan Shen, Christopher Malon et al.ACL 2020 · 66 citations
- Learning Disentangled Textual Representations via Statistical Measures of SimilarityPierre Colombo, Guillaume Staerman, Nathan Noiry, Pablo PiantanidaACL 2022
- Revision in Continuous Space: Unsupervised Text Style Transfer without Adversarial LearningDayiheng Liu, Jie Fu, Yidan Zhang, Chris Pal et al.AAAI 2020 · 53 citations
- Multi-type Disentanglement without Adversarial TrainingLei Sha, Thomas LukasiewiczAAAI 2021 · 13 citations
- Adversarial Disentanglement with Grouped ObservationsJózsef NémethAAAI 2020 · 8 citations
