Representation Learning with Conditional Information Flow Maximization
Dou Hu, Lingwei Wei, Wei Zhou, Songlin Hu
Abstract
This paper proposes an information-theoretic representation learning framework, named conditional information flow maximization, to extract noise-invariant sufficient representations for the input data and target task. It promotes the learned representations have good feature uniformity and sufficient predictive ability, which can enhance the generalization of pre-trained language models (PLMs) for the target task. Firstly, an information flow maximization principle is proposed to learn more sufficient representations for the input and target by simultaneously maximizing both inputrepresentation and representation-label mutual information. Unlike the information bottleneck, we handle the input-representation information in an opposite way to avoid the overcompression issue of latent representations. Besides, to mitigate the negative effect of potential redundant features from the input, we design a conditional information minimization principle to eliminate negative redundant features while preserve noise-invariant features. Experiments on 13 language understanding benchmarks demonstrate that our method effectively improves the performance of PLMs for classification and regression. Extensive experiments show that the learned representations are more sufficient, robust and transferable.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 449f7453-797a-4ae4-b11e-8fc860335aadCited by top-tier papers2
- An Information-theoretic Multi-task Representation Learning Framework for Natural Language UnderstandingDou Hu, Lingwei Wei, Wei Zhou, Songlin HuAAAI 2025 · 3 citations
- Multi-Task Representation Alignment on Language Understanding: A Mutual Information PerspectiveDou Hu, Lingwei Wei, Hongjiang Xiao, Songlin Hu et al.ACL 2026
Builds on13
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
- What Makes for Good Views for Contrastive Learning?Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan et al.NeurIPS 2020 · 1,631 citations
- On Mutual Information Maximization for Representation LearningMichael Tschannen, Josip Djolonga, Paul K. Rubenstein, Sylvain Gelly et al.ICLR 2020 · 559 citations
Related papers
- Information Retention via Learning Supplemental FeaturesZhipeng Xie, Yahe LiICLR 2024 · 1 citation
- Learning is Forgetting; LLM Training As Lossy CompressionHenry Conklin, Tom Hosking, Yi Chern Tan, Jonathan D. Cohen et al.ICLR 2026 · 6 citations
- Text Representation Distillation via Information Bottleneck PrincipleYanzhao Zhang, Dingkun Long, Zehan Li, Pengjun XieEMNLP 2023 · 4 citations
- InfoBERT: Improving Robustness of Language Models from An Information Theoretic PerspectiveBoxin Wang, Shuohang Wang, Yu Cheng, Zhe Gan et al.ICLR 2021 · 132 citations
- Variational Information Bottleneck for Effective Low-Resource Fine-TuningRabeeh Karimi Mahabadi, Yonatan Belinkov, James HendersonICLR 2021 · 88 citations
