A Differential Entropy Estimator for Training Neural Networks
Georg Pichler, Pierre Jean A. Colombo, Malik Boudiaf, Günther Koliander, Pablo Piantanida
摘要
Mutual Information (MI) has been widely used as a loss regularizer for training neural networks. This has been particularly effective when learn disentangled or compressed representations of high dimensional data. However, differential entropy (DE), another fundamental measure of information, has not found widespread use in neural network training. Although DE offers a potentially wider range of applications than MI, off-the-shelf DE estimators are either non differentiable, computationally intractable or fail to adapt to changes in the underlying distribution. These drawbacks prevent them from being used as regularizers in neural networks training. To address shortcomings in previously proposed estimators for DE, here we introduce KNIFE, a fully parameterized, differentiable kernel-based estimator of DE. The flexibility of our approach also allows us to construct KNIFE-based estimators for conditional (on either discrete or continuous variables) DE, as well as MI. We empirically validate our method on high-dimensional synthetic data and further apply it to guide the training of neural networks for real-world tasks. Our experiments on a large variety of tasks, including visual domain adaptation, textual fair classification, and textual fine-tuning demonstrate the effectiveness of KNIFE-based estimation. Code can be found at https://github.com/g-pichler/knife.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Beyond Mahalanobis Distance for Textual OOD DetectionPierre Colombo, Eduardo Dadalto Câmara Gomes, Guillaume Staerman, Nathan Noiry 等NeurIPS 2022 · 被引用 24 次
- Robust Concept Erasure via Kernelized Rate-Distortion MaximizationSomnath Basu Roy Chowdhury, Nicholas Monath, Kumar Avinava Dubey, Amr Ahmed 等NeurIPS 2023 · 被引用 12 次
- Sourcerer: Sample-based Maximum Entropy Source Distribution EstimationJulius Vetter, Guy Moss, Cornelius Schröder, Richard Gao 等NeurIPS 2024 · 被引用 11 次
- Transductive Learning for Textual Few-Shot Classification in API-based Embedding ModelsPierre Colombo, Victor Pellegrain, Malik Boudiaf, Myriam Tami 等EMNLP 2023 · 被引用 7 次
- REMEDI: Corrective Transformations for Improved Neural Entropy EstimationViktor Nilsson, Anirban Samaddar, Sandeep Madireddy, Pierre NyquistICML 2024 · 被引用 3 次
它引用的顶会 Paper8
- Invariant Information Clustering for Unsupervised Image Classification and SegmentationXu Ji, Andrea Vedaldi, João F. HenriquesICCV 2019 · 被引用 956 次
- On Mutual Information Maximization for Representation LearningMichael Tschannen, Josip Djolonga, Paul K. Rubenstein, Sylvain Gelly 等ICLR 2020 · 被引用 559 次
- CLUB: A Contrastive Log-ratio Upper Bound of Mutual InformationPengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu 等ICML 2020 · 被引用 512 次
- Understanding the Limitations of Variational Mutual Information EstimatorsJiaming Song, Stefano ErmonICLR 2020 · 被引用 243 次
- Variational Information Bottleneck for Effective Low-Resource Fine-TuningRabeeh Karimi Mahabadi, Yonatan Belinkov, James HendersonICLR 2021 · 被引用 88 次
相关 Paper
- Scalable Infomin LearningYanzhi Chen, Weihao Sun, Yingzhen Li, Adrian WellerNeurIPS 2022 · 被引用 10 次
- Diffeomorphic Information Neural EstimationBao Duong, Thin NguyenAAAI 2023 · 被引用 10 次
- Learning Disentangled Textual Representations via Statistical Measures of SimilarityPierre Colombo, Guillaume Staerman, Nathan Noiry, Pablo PiantanidaACL 2022
- Connecting Jensen-Shannon and Kullback-Leibler Divergences: A New Bound for Representation LearningReuben Dorent, Polina Golland, William (Sandy) WellsNeurIPS 2025 · 被引用 7 次
- Beyond Normal: On the Evaluation of Mutual Information EstimatorsPawel Czyz, Frederic Grabowski, Julia E. Vogt, Niko Beerenwinkel 等NeurIPS 2023 · 被引用 72 次
