Supervised Contrastive Learning for Pre-trained Language Model Fine-tuning
Beliz Gunel, Jingfei Du, Alexis Conneau, Veselin Stoyanov
Abstract
State-of-the-art natural language understanding classification models follow two-stages: pre-training a large language model on an auxiliary task, and then fine-tuning the model on a task-specific labeled dataset using cross-entropy loss. Cross-entropy loss has several shortcomings that can lead to sub-optimal generalization and instability. Driven by the intuition that good generalization requires capturing the similarity between examples in one class and contrasting them with examples in other classes, we propose a supervised contrastive learning (SCL) objective for the fine-tuning stage. Combined with cross-entropy, the SCL loss we propose obtains improvements over a strong RoBERTa-Large baseline on multiple datasets of the GLUE benchmark in both the high-data and low-data regimes, and it does not require any specialized architecture, data augmentation of any kind, memory banks, or additional unsupervised data. We also demonstrate that the new objective leads to models that are more robust to different levels of noise in the training data, and can generalize better to related tasks with limited labeled task data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2fedb17d-4a38-4ebf-a317-3176cbd23acdCited by top-tier papers53
- Intent Contrastive Learning for Sequential RecommendationYongjun Chen, Zhiwei Liu, Jia Li, Julian J. McAuley et al.WWW 2022 · 429 citations
- Contrast and Generation Make BART a Good Dialogue Emotion RecognizerShimin Li, Hang Yan, Xipeng QiuAAAI 2022 · 121 citations
- Improved Text Classification via Contrastive Adversarial TrainingLin Pan, Chung-Wei Hang, Avirup Sil, Saloni PotdarAAAI 2022 · 115 citations
- Expanding Small-Scale Datasets with Guided ImaginationYifan Zhang, Daquan Zhou, Bryan Hooi, Kai Wang et al.NeurIPS 2023 · 84 citations
- Supervised Adversarial Contrastive Learning for Emotion Recognition in ConversationsDou Hu, Yinan Bao, Lingwei Wei, Wei Zhou et al.ACL 2023 · 64 citations
Builds on14
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
Related papers
- CoDA: Contrast-enhanced and Diversity-promoting Data Augmentation for Natural Language UnderstandingYanru Qu, Dinghan Shen, Yelong Shen, Sandra Sajeev et al.ICLR 2021 · 77 citations
- Contrastive Classification and Representation Learning with Probabilistic InterpretationRahaf Aljundi, Yash Patel, Milan Sulc, Nikolay Chumerin et al.AAAI 2023 · 10 citations
- OssCSE: Overcoming Surface Structure Bias in Contrastive Learning for Unsupervised Sentence EmbeddingZhan Shi, Guoyin Wang, Ke Bai, Jiwei Li et al.EMNLP 2023 · 3 citations
- Rethinking Denoised Auto-Encoding in Language Pre-TrainingFuli Luo, Pengcheng Yang, Shicheng Li, Xuancheng Ren et al.EMNLP 2021 · 8 citations
- COCO-LM: Correcting and Contrasting Text Sequences for Language Model PretrainingYu Meng, Chenyan Xiong, Payal Bajaj, Saurabh Tiwary et al.NeurIPS 2021 · 231 citations
