Headless Language Models: Learning without Predicting with Contrastive Weight Tying
Nathan Godey, Éric Villemonte de la Clergerie, Benoît Sagot
摘要
Self-supervised pre-training of language models usually consists in predicting probability distributions over extensive token vocabularies. In this study, we propose an innovative method that shifts away from probability prediction and instead focuses on reconstructing input embeddings in a contrastive fashion via Constrastive Weight Tying (CWT). We apply this approach to pretrain Headless Language Models in both monolingual and multilingual contexts. Our method offers practical advantages, substantially reducing training computational requirements by up to 20 times, while simultaneously enhancing downstream performance and data efficiency. We observe a significant +1.6 GLUE score increase and a notable +2.7 LAMBADA accuracy improvement compared to classical LMs within similar compute budgets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- In-Context Reinforcement Learning for Variable Action SpacesViacheslav Sinii, Alexander Nikulin, Vladislav Kurenkov, Ilya Zisman 等ICML 2024 · 被引用 26 次
- Zipfian WhiteningSho Yokoi, Han Bao, Hiroto Kurita, Hidetoshi ShimodairaNeurIPS 2024 · 被引用 3 次
它引用的顶会 Paper14
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 被引用 394 次
相关 Paper
- COCO-LM: Correcting and Contrasting Text Sequences for Language Model PretrainingYu Meng, Chenyan Xiong, Payal Bajaj, Saurabh Tiwary 等NeurIPS 2021 · 被引用 231 次
- SLlama: Parameter-Efficient Language Model Architecture for Enhanced Linguistic Competence Under Strict Data ConstraintsVictor Adelakun Omolaoye, Babajide Alamu Owoyele, Gerard de MeloEMNLP 2025
- Accelerating Training of Transformer-Based Language Models with Progressive Layer DroppingMinjia Zhang, Yuxiong HeNeurIPS 2020 · 被引用 126 次
- Rethinking Embedding Coupling in Pre-trained Language ModelsHyung Won Chung, Thibault Févry, Henry Tsai, Melvin Johnson 等ICLR 2021 · 被引用 11 次
- Cross-View Language Modeling: Towards Unified Cross-Lingual Cross-Modal Pre-trainingYan Zeng, Wangchunshu Zhou, Ao Luo, Ziming Cheng 等ACL 2023 · 被引用 18 次
