Fast, Effective, and Self-Supervised: Transforming Masked Language Models into Universal Lexical and Sentence Encoders
Fangyu Liu, Ivan Vulic, Anna Korhonen, Nigel Collier
摘要
Previous work has indicated that pretrained Masked Language Models (MLMs) are not effective as universal lexical and sentence encoders off-the-shelf, i.e., without further taskspecific fine-tuning on NLI, sentence similarity, or paraphrasing tasks using annotated task data. In this work, we demonstrate that it is possible to turn MLMs into effective lexical and sentence encoders even without any additional data, relying simply on self-supervision. We propose an extremely simple, fast, and effective contrastive learning technique, termed Mirror-BERT, which converts MLMs (e.g., BERT and RoBERTa) into such encoders in 20-30 seconds with no access to additional external knowledge. Mirror-BERT relies on identical and slightly modified string pairs as positive (i.e., synonymous) fine-tuning examples, and aims to maximise their similarity during "identity fine-tuning". We report huge gains over off-the-shelf MLMs with Mirror-BERT both in lexical-level and in sentencelevel tasks, across different domains and different languages. Notably, in sentence similarity (STS) and question-answer entailment (QNLI) tasks, our self-supervised Mirror-BERT model even matches the performance of the Sentence-BERT models from prior work which rely on annotated task data. Finally, we delve deeper into the inner workings of MLMs, and suggest some evidence on why this simple Mirror-BERT fine-tuning approach can yield effective universal lexical and sentence encoders.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- A Contrastive Framework for Neural Text GenerationYixuan Su, Tian Lan, Yan Wang, Dani Yogatama 等NeurIPS 2022 · 被引用 349 次
- Trans-Encoder: Unsupervised sentence-pair modelling through self- and mutual-distillationsFangyu Liu, Yunlong Jiao, Jordan Massiah, Emine Yilmaz 等ICLR 2022 · 被引用 36 次
- Ranking-Enhanced Unsupervised Sentence Representation LearningYeon Seonwoo, Guoyin Wang, Changmin Seo, Sajal Choudhary 等ACL 2023 · 被引用 12 次
- Punctuation-level Attack: Single-shot and Single Punctuation Can Fool Text ModelsWenqiang Wang, Chongyang Du, Tao Wang, Kaihao Zhang 等NeurIPS 2023 · 被引用 11 次
- miCSE: Mutual Information Contrastive Learning for Low-shot Sentence EmbeddingsTassilo Klein, Moin NabiACL 2023 · 被引用 10 次
它引用的顶会 Paper17
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
- AugMix: A Simple Data Processing Method to Improve Robustness and UncertaintyDan Hendrycks, Norman Mu, Ekin Dogus Cubuk, Barret Zoph 等ICLR 2020 · 被引用 1,572 次
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang 等EMNLP 2020 · 被引用 538 次
- Poly-encoders: Architectures and Pre-training Strategies for Fast and Accurate Multi-sentence ScoringSamuel Humeau, Kurt Shuster, Marie-Anne Lachaux, Jason WestonICLR 2020 · 被引用 316 次
相关 Paper
- Self-Guided Contrastive Learning for BERT Sentence RepresentationsTaeuk Kim, Kang Min Yoo, Sang-goo LeeACL 2021
- Improving Word Translation via Two-Stage Contrastive LearningYaoyiran Li, Fangyu Liu, Nigel Collier, Anna Korhonen 等ACL 2022 · 被引用 32 次
- ConSERT: A Contrastive Framework for Self-Supervised Sentence Representation TransferYuanmeng Yan, Rumei Li, Sirui Wang, Fuzheng Zhang 等ACL 2021
- QuASE: Question-Answer Driven Sentence EncodingHangfeng He, Qiang Ning, Dan RothACL 2020 · 被引用 32 次
- OssCSE: Overcoming Surface Structure Bias in Contrastive Learning for Unsupervised Sentence EmbeddingZhan Shi, Guoyin Wang, Ke Bai, Jiwei Li 等EMNLP 2023 · 被引用 3 次
