Contrastive Learning of Sentence Embeddings from Scratch
Junlei Zhang, Zhenzhong Lan, Junxian He
摘要
Contrastive learning has been the dominant approach to train state-of-the-art sentence embeddings. Previous studies have typically learned sentence embeddings either through the use of human-annotated natural language inference (NLI) data or via large-scale unlabeled sentences in an unsupervised manner. However, even in the case of 1;unlabeled data, their acquisition presents challenges in certain domains due to various reasons. To address these issues, we present SynCSE, a contrastive learning framework that trains sentence embeddings with synthesized data. Specifically, we explore utilizing large language models to synthesize the required data samples for contrastive learning, including (1) producing positive and negative annotations given unlabeled sentences (SynCSE-partial), and (2) generating sentences along with their corresponding annotations from scratch (SynCSE-scratch). Experimental results on sentence similarity and reranking tasks indicate that both SynCSE-partial and SynCSE-scratch greatly outperform unsupervised baselines, and SynCSE-partial even achieves comparable performance to the supervised models in most settings. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Subgraph-Aware Training of Language Models for Knowledge Graph Completion Using Structure-Aware Contrastive LearningYoumin Ko, Hyemin Yang, Taeuk Kim, Hyunjoon KimWWW 2025 · 被引用 10 次
- Meta-Task Prompting Elicits Embeddings from Large Language ModelsYibin Lei, Di Wu, Tianyi Zhou, Tao Shen 等ACL 2024 · 被引用 6 次
- Optimizing Code Retrieval: High-Quality and Scalable Dataset Annotation through Large Language ModelsRui Li, Qi Liu, Liyang He, Zheng Zhang 等EMNLP 2024 · 被引用 4 次
- Template-assisted Contrastive Learning of Task-oriented Dialogue Sentence EmbeddingsMinsik Oh, Jiwei Li, Guoyin WangACL 2026 · 被引用 2 次
- Contrastive Perplexity for Controlled Generation: An Application in Detoxifying Large Language ModelsTassilo Klein, Moin NabiACL 2025
它引用的顶会 Paper12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- MIND: A Large-scale Dataset for News RecommendationFangzhao Wu, Ying Qiao, Jiun-Hung Chen, Chuhan Wu 等ACL 2020 · 被引用 454 次
- Debiased Contrastive Learning of Unsupervised Sentence RepresentationsKun Zhou, Beichen Zhang, Wayne Xin Zhao, Ji-Rong WenACL 2022 · 被引用 128 次
- ZeroGen: Efficient Zero-shot Learning via Dataset GenerationJiacheng Ye, Jiahui Gao, Qintong Li, Hang Xu 等EMNLP 2022 · 被引用 96 次
相关 Paper
- Generate, Discriminate and Contrast: A Semi-Supervised Sentence Representation Learning FrameworkYiming Chen, Yan Zhang, Bin Wang, Zuozhu Liu 等EMNLP 2022 · 被引用 9 次
- Narrowing the Gap between Supervised and Unsupervised Sentence Representation Learning with Large Language ModelMingxin Li, Richong Zhang, Zhijie Nie, Yongyi MaoAAAI 2024 · 被引用 1 次
- DeCLUTR: Deep Contrastive Learning for Unsupervised Textual RepresentationsJohn M. Giorgi, Osvald Nitski, Bo Wang, Gary D. BaderACL 2021
- English Contrastive Learning Can Learn Universal Cross-lingual Sentence EmbeddingsYau-Shian Wang, Ashley Wu, Graham NeubigEMNLP 2022 · 被引用 18 次
- Unsupervised Sentence Representation via Contrastive Learning with Mixing NegativesYanzhao Zhang, Richong Zhang, Samuel Mensah, Xudong Liu 等AAAI 2022 · 被引用 71 次
