PAUSE: Positive and Annealed Unlabeled Sentence Embedding
Lele Cao, Emil Larsson, Vilhelm von Ehrenheim, Dhiana Deva Cavalcanti Rocha, Anna Martin, Sonja Horn
摘要
Sentence embedding refers to a set of effective and versatile techniques for converting raw text into numerical vector representations that can be used in a wide range of natural language processing (NLP) applications. The majority of these techniques are either supervised or unsupervised. Compared to the unsupervised methods, the supervised ones make less assumptions about optimization objectives and usually achieve better results. However, the training requires a large amount of labeled sentence pairs, which is not available in many industrial scenarios. To that end, we propose a generic and end-to-end approach -PAUSE (Positive and Annealed Unlabeled Sentence Embedding), capable of learning high-quality sentence embeddings from a partially labeled dataset. We experimentally show that PAUSE achieves, and sometimes surpasses, state-ofthe-art results using only a small fraction of labeled sentence pairs on various benchmark tasks. When applied to a real industrial use case where labeled samples are scarce, PAUSE encourages us to extend our dataset without the burden of extensive manual annotation work.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Non-Linguistic Supervision for Contrastive Learning of Sentence EmbeddingsYiren Jian, Chongyang Gao, Soroush VosoughiNeurIPS 2022 · 被引用 20 次
- A Unified Positive-Unlabeled Learning Framework for Document-Level Relation Extraction with Different Levels of LabelingYe Wang, Xinxin Liu, Wenxin Hu, Tao ZhangEMNLP 2022 · 被引用 18 次
- A Positive-Unlabeled Metric Learning Framework for Document-Level Relation Extraction with Incomplete LabelingYe Wang, Huazheng Pan, Tao Zhang, Wen Wu 等AAAI 2024 · 被引用 11 次
它引用的顶会 Paper8
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- An Unsupervised Sentence Embedding Method by Mutual Information MaximizationYan Zhang, Ruidan He, Zuozhu Liu, Kwan Hui Lim 等EMNLP 2020 · 被引用 126 次
- Self-PU: Self Boosted and Calibrated Positive-Unlabeled TrainingXuxi Chen, Wuyang Chen, Tianlong Chen, Ye Yuan 等ICML 2020 · 被引用 100 次
- Semantic Re-tuning with Contrastive TensionFredrik Carlsson, Amaru Cuba Gyllensten, Evangelia Gogoulou, Erik Ylipää Hellqvist 等ICLR 2021 · 被引用 86 次
- Language-agnostic BERT Sentence EmbeddingFangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan 等ACL 2022
相关 Paper
- Generate, Discriminate and Contrast: A Semi-Supervised Sentence Representation Learning FrameworkYiming Chen, Yan Zhang, Bin Wang, Zuozhu Liu 等EMNLP 2022 · 被引用 9 次
- DeCLUTR: Deep Contrastive Learning for Unsupervised Textual RepresentationsJohn M. Giorgi, Osvald Nitski, Bo Wang, Gary D. BaderACL 2021
- Contrastive Learning of Sentence Embeddings from ScratchJunlei Zhang, Zhenzhong Lan, Junxian HeEMNLP 2023 · 被引用 9 次
- Ranking-Enhanced Unsupervised Sentence Representation LearningYeon Seonwoo, Guoyin Wang, Changmin Seo, Sajal Choudhary 等ACL 2023 · 被引用 12 次
- Scalable Attentive Sentence Pair Modeling via Distilled Sentence EmbeddingOren Barkan, Noam Razin, Itzik Malkiel, Ori Katz 等AAAI 2020 · 被引用 37 次
