FlowNIB: An Information Bottleneck Analysis of Bidirectional vs. Unidirectional Language Models
Md Kowsher, Nusrat Jahan Prottasha, Shiyun Xu, Shetu Mohanto, Niloofar Yousefi, Ozlem O. Garibay, Chen Chen
Abstract
Bidirectional language models have better context understanding and perform better than unidirectional models on natural language understanding tasks, yet the theoretical reasons behind this advantage remain unclear. In this work, we investigate this disparity through the lens of the Information Bottleneck (IB) principle, which formalizes a trade-off between compressing input information and preserving task-relevant content. We propose FlowNIB, a dynamic and scalable method for estimating mutual information during training that addresses key limitations of classical IB approaches, including computational intractability and fixed trade-off schedules. Theoretically, we show that bidirectional models retain more mutual information and exhibit higher effective dimensionality than unidirectional models. To support this, we present a generalized framework for measuring representational complexity and prove that bidirectional representations are strictly more informative under mild conditions. We further validate our findings through extensive experiments across multiple models and tasks using FlowNIB, revealing how information is encoded and compressed throughout training. Together, our work provides a principled explanation for the effectiveness of bidirectional architectures and introduces a practical tool for analyzing information flow in deep language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 44b7d50e-07b2-453d-8551-c8262d349492Builds on7
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and InferenceBenjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller et al.ACL 2025 · 552 citations
- MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use CasesZechun Liu, Changsheng Zhao, Forrest N. Iandola, Chen Lai et al.ICML 2024 · 227 citations
- Predicting Through Generation: Why Generation Is Better for PredictionMd. Kowsher, Nusrat Jahan Prottasha, Prakash Bhat, Chun-Nam Yu et al.ACL 2025
Related papers
- Information Bottleneck Analysis of Deep Neural Networks via Lossy CompressionIvan Butakov, Aleksander Tolmachev, Sofia Malanchuk, Anna Neopryatnaya et al.ICLR 2024 · 20 citations
- Information Bottleneck: Exact Analysis of (Quantized) Neural NetworksStephan Sloth Lorenzen, Christian Igel, Mads NielsenICLR 2022 · 24 citations
- Cauchy-Schwarz Divergence Information Bottleneck for RegressionShujian Yu, Xi Yu, Sigurd Løkse, Robert Jenssen et al.ICLR 2024 · 16 citations
- Learning is Forgetting; LLM Training As Lossy CompressionHenry Conklin, Tom Hosking, Yi Chern Tan, Jonathan D. Cohen et al.ICLR 2026 · 6 citations
- Representation Learning with Conditional Information Flow MaximizationDou Hu, Lingwei Wei, Wei Zhou, Songlin HuACL 2024
