A Novel Estimator of Mutual Information for Learning to Disentangle Textual Representations
Pierre Colombo, Pablo Piantanida, Chloé Clavel
摘要
Learning disentangled representations of textual data is essential for many natural language tasks such as fair classification, style transfer and sentence generation, among others. The existent dominant approaches in the context of text data either rely on training an adversary (discriminator) that aims at making attribute values difficult to be inferred from the latent code or rely on minimising variational bounds of the mutual information between latent code and the value attribute. However, the available methods suffer of the impossibility to provide a fine-grained control of the degree (or force) of disentanglement. In contrast to adversarial methods, which are remarkably simple, although the adversary seems to be performing perfectly well during the training phase, after it is completed a fair amount of information about the undesired attribute still remains. This paper introduces a novel variational upper bound to the mutual information between an attribute and the latent code of an encoder. Our bound aims at controlling the approximation error via the Renyi's divergence, leading to both better disentangled representations and in particular, a precise control of the desirable degree of disentanglement than state-of-the-art methods proposed for textual data. Furthermore, it does not suffer from the degeneracy of other losses in multi-class scenarios. We show the superiority of this method on fair classification and on textual style transfer tasks. Additionally, we provide new insights illustrating various trade-offs in style transfer when attempting to learn disentangled representations and quality of the generated sentence.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- A Differential Entropy Estimator for Training Neural NetworksGeorg Pichler, Pierre Jean A. Colombo, Malik Boudiaf, Günther Koliander 等ICML 2022 · 被引用 28 次
- Beyond Mahalanobis Distance for Textual OOD DetectionPierre Colombo, Eduardo Dadalto Câmara Gomes, Guillaume Staerman, Nathan Noiry 等NeurIPS 2022 · 被引用 24 次
- DARE: Disentanglement-Augmented Rationale ExtractionLinan Yue, Qi Liu, Yichao Du, Yanqing An 等NeurIPS 2022 · 被引用 24 次
- Learning Disentangled Representations of Negation and UncertaintyJake Vasilakes, Chrysoula Zerva, Makoto Miwa, Sophia AnaniadouACL 2022 · 被引用 22 次
- What are the best Systems? New Perspectives on NLP BenchmarkingPierre Colombo, Nathan Noiry, Ekhine Irurozki, Stéphan ClémençonNeurIPS 2022 · 被引用 20 次
它引用的顶会 Paper4
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- CLUB: A Contrastive Log-ratio Upper Bound of Mutual InformationPengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu 等ICML 2020 · 被引用 512 次
- Guiding Attention in Sequence-to-Sequence Models for Dialogue Act PredictionPierre Colombo, Emile Chapuis, Matteo Manica, Emmanuel Vignon 等AAAI 2020 · 被引用 69 次
- Heavy-tailed Representations, Text Polarity Classification & Data AugmentationHamid Jalalzai, Pierre Colombo, Chloé Clavel, Éric Gaussier 等NeurIPS 2020 · 被引用 33 次
相关 Paper
- Improving Disentangled Text Representation Learning with Information-Theoretic GuidancePengyu Cheng, Martin Renqiang Min, Dinghan Shen, Christopher Malon 等ACL 2020 · 被引用 66 次
- Learning Disentangled Textual Representations via Statistical Measures of SimilarityPierre Colombo, Guillaume Staerman, Nathan Noiry, Pablo PiantanidaACL 2022
- Revision in Continuous Space: Unsupervised Text Style Transfer without Adversarial LearningDayiheng Liu, Jie Fu, Yidan Zhang, Chris Pal 等AAAI 2020 · 被引用 53 次
- Multi-type Disentanglement without Adversarial TrainingLei Sha, Thomas LukasiewiczAAAI 2021 · 被引用 13 次
- Adversarial Disentanglement with Grouped ObservationsJózsef NémethAAAI 2020 · 被引用 8 次
