Contrastive Learning of Sentence Embeddings from Scratch
Junlei Zhang, Zhenzhong Lan, Junxian He
Abstract
Contrastive learning has been the dominant approach to train state-of-the-art sentence embeddings. Previous studies have typically learned sentence embeddings either through the use of human-annotated natural language inference (NLI) data or via large-scale unlabeled sentences in an unsupervised manner. However, even in the case of 1;unlabeled data, their acquisition presents challenges in certain domains due to various reasons. To address these issues, we present SynCSE, a contrastive learning framework that trains sentence embeddings with synthesized data. Specifically, we explore utilizing large language models to synthesize the required data samples for contrastive learning, including (1) producing positive and negative annotations given unlabeled sentences (SynCSE-partial), and (2) generating sentences along with their corresponding annotations from scratch (SynCSE-scratch). Experimental results on sentence similarity and reranking tasks indicate that both SynCSE-partial and SynCSE-scratch greatly outperform unsupervised baselines, and SynCSE-partial even achieves comparable performance to the supervised models in most settings. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f871d0f7-97e1-43be-b1f7-454e39f8e726Cited by top-tier papers7
- Subgraph-Aware Training of Language Models for Knowledge Graph Completion Using Structure-Aware Contrastive LearningYoumin Ko, Hyemin Yang, Taeuk Kim, Hyunjoon KimWWW 2025 · 10 citations
- Meta-Task Prompting Elicits Embeddings from Large Language ModelsYibin Lei, Di Wu, Tianyi Zhou, Tao Shen et al.ACL 2024 · 6 citations
- Optimizing Code Retrieval: High-Quality and Scalable Dataset Annotation through Large Language ModelsRui Li, Qi Liu, Liyang He, Zheng Zhang et al.EMNLP 2024 · 4 citations
- Template-assisted Contrastive Learning of Task-oriented Dialogue Sentence EmbeddingsMinsik Oh, Jiwei Li, Guoyin WangACL 2026 · 2 citations
- Contrastive Perplexity for Controlled Generation: An Application in Detoxifying Large Language ModelsTassilo Klein, Moin NabiACL 2025
Builds on12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- MIND: A Large-scale Dataset for News RecommendationFangzhao Wu, Ying Qiao, Jiun-Hung Chen, Chuhan Wu et al.ACL 2020 · 454 citations
- Debiased Contrastive Learning of Unsupervised Sentence RepresentationsKun Zhou, Beichen Zhang, Wayne Xin Zhao, Ji-Rong WenACL 2022 · 128 citations
- ZeroGen: Efficient Zero-shot Learning via Dataset GenerationJiacheng Ye, Jiahui Gao, Qintong Li, Hang Xu et al.EMNLP 2022 · 96 citations
Related papers
- Generate, Discriminate and Contrast: A Semi-Supervised Sentence Representation Learning FrameworkYiming Chen, Yan Zhang, Bin Wang, Zuozhu Liu et al.EMNLP 2022 · 9 citations
- Narrowing the Gap between Supervised and Unsupervised Sentence Representation Learning with Large Language ModelMingxin Li, Richong Zhang, Zhijie Nie, Yongyi MaoAAAI 2024 · 1 citation
- DeCLUTR: Deep Contrastive Learning for Unsupervised Textual RepresentationsJohn M. Giorgi, Osvald Nitski, Bo Wang, Gary D. BaderACL 2021
- English Contrastive Learning Can Learn Universal Cross-lingual Sentence EmbeddingsYau-Shian Wang, Ashley Wu, Graham NeubigEMNLP 2022 · 18 citations
- Unsupervised Sentence Representation via Contrastive Learning with Mixing NegativesYanzhao Zhang, Richong Zhang, Samuel Mensah, Xudong Liu et al.AAAI 2022 · 71 citations
