A Mutual Information Maximization Perspective of Language Representation Learning
Lingpeng Kong, Cyprien de Masson d'Autume, Lei Yu, Wang Ling, Zihang Dai, Dani Yogatama
Abstract
We show state-of-the-art word representation learning methods maximize an objective function that is a lower bound on the mutual information between different parts of a word sequence (i.e., a sentence). Our formulation provides an alternative perspective that unifies classical word embedding models (e.g., Skip-gram) and modern contextual embeddings (e.g., BERT, XLNet). In addition to enhancing our theoretical understanding of these methods, our derivation leads to a principled framework that can be used to construct new self-supervised tasks. We provide an example by drawing inspirations from related methods based on mutual information maximization that have been successful in computer vision, and introduce a simple self-supervised objective that maximizes the mutual information between a global sentence representation and n-grams in the sentence. Our analysis offers a holistic view of representation learning methods to transfer knowledge and translate progress across multiple domains (e.g., natural language processing, computer vision, audio processing).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cee84b41-39fb-4dd3-8ca5-cd76c97832c8Cited by top-tier papers61
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- What Makes for Good Views for Contrastive Learning?Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan et al.NeurIPS 2020 · 1,631 citations
- Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language ProcessingZihang Dai, Guokun Lai, Yiming Yang, Quoc LeNeurIPS 2020 · 273 citations
- Coreferential Reasoning Learning for Language RepresentationDeming Ye, Yankai Lin, Jiaju Du, Zhenghao Liu et al.EMNLP 2020 · 164 citations
- The Effect of Natural Distribution Shift on Question Answering ModelsJohn Miller, Karl Krauth, Benjamin Recht, Ludwig SchmidtICML 2020 · 158 citations
Builds on1
Related papers
- Self-Guided Contrastive Learning for BERT Sentence RepresentationsTaeuk Kim, Kang Min Yoo, Sang-goo LeeACL 2021
- An Unsupervised Sentence Embedding Method by Mutual Information MaximizationYan Zhang, Ruidan He, Zuozhu Liu, Kwan Hui Lim et al.EMNLP 2020 · 126 citations
- SLM: Learning a Discourse Language Representation with Sentence UnshufflingHaejun Lee, Drew A. Hudson, Kangwook Lee, Christopher D. ManningEMNLP 2020 · 2 citations
- Scalable Attentive Sentence Pair Modeling via Distilled Sentence EmbeddingOren Barkan, Noam Razin, Itzik Malkiel, Ori Katz et al.AAAI 2020 · 37 citations
- Self-Supervised Representation Learning From Multi-Domain DataZeyu Feng, Chang Xu, Dacheng TaoICCV 2019 · 46 citations
