Comparing Text Representations: A Theory-Driven Approach
Gregory Yauney, David Mimno
摘要
Much of the progress in contemporary NLP has come from learning representations, such as masked language model (MLM) contextual embeddings, that turn challenging problems into simple classification tasks. But how do we quantify and explain this effect? We adapt general tools from computational learning theory to fit the specific characteristics of text datasets and present a method to evaluate the compatibility between representations and tasks. Even though many tasks can be easily solved with simple bag-of-words (BOW) representations, BOW does poorly on hard natural language inference tasks. For one such task we find that BOW cannot distinguish between real and randomized labelings, while pre-trained MLM representations show 72x greater distinction between real and random labelings than BOW. This method provides a calibrated, quantitative measure of the difficulty of a classification-based NLP task, enabling comparisons between representations without requiring empirical evaluations that may be sensitive to initializations and hyperparameters. The method provides a fresh perspective on the patterns in a dataset and the alignment of those patterns with specific labels.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal 等ACL 2020 · 被引用 602 次
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang 等EMNLP 2020 · 被引用 538 次
- Emerging Cross-lingual Structure in Pretrained Language ModelsAlexis Conneau, Shijie Wu, Haoran Li, Luke Zettlemoyer 等ACL 2020 · 被引用 210 次
- A Mathematical Exploration of Why Language Models Help Solve Downstream TasksNikunj Saunshi, Sadhika Malladi, Sanjeev AroraICLR 2021 · 被引用 93 次
相关 Paper
- Efficient Pre-training of Masked Language Model via Concept-based Curriculum MaskingMingyu Lee, Jun-Hyung Park, Junho Kim, Kang-Min Kim 等EMNLP 2022 · 被引用 8 次
- Information-Theoretic Probing with Minimum Description LengthElena Voita, Ivan TitovEMNLP 2020 · 被引用 34 次
- Probing as Quantifying Inductive BiasAlexander Immer, Lucas Torroba Hennigen, Vincent Fortuin, Ryan CotterellACL 2022
- Clotho: Measuring Task-Specific Pre-Generation Test Adequacy for LLM InputsJuyeon Yoon, Somin Kim, Robert Feldt, Shin YooFSE 2026 · 被引用 2 次
- Understanding the Transferability of Representations via Task-RelatednessAkshay Mehra, Yunbei Zhang, Jihun HammNeurIPS 2024 · 被引用 13 次
