Phrase-BERT: Improved Phrase Embeddings from BERT with an Application to Corpus Exploration
Shufan Wang, Laure Thompson, Mohit Iyyer
摘要
Phrase representations derived from BERT often do not exhibit complex phrasal compositionality, as the model relies instead on lexical similarity to determine semantic relatedness. In this paper, we propose a contrastive fine-tuning objective that enables BERT to produce more powerful phrase embeddings. Our approach (Phrase-BERT) relies on a dataset of diverse phrasal paraphrases, which is automatically generated using a paraphrase generation model, as well as a large-scale dataset of phrases in context mined from the Books3 corpus. Phrase-BERT outperforms baselines across a variety of phrase-level similarity tasks, while also demonstrating increased lexical diversity between nearest neighbors in the vector space. Finally, as a case study, we show that Phrase-BERT embeddings can be easily integrated with a simple autoencoder to build a phrase-based neural topic model that interprets topics as mixtures of words and phrases by performing a nearest neighbor search in the embedding space. Crowdsourced evaluations demonstrate that this phrase-based topic model produces more coherent and meaningful topics than baseline word and phrase-level topic models, further validating the utility of Phrase-BERT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- #InsTag: Instruction Tagging for Analyzing Supervised Fine-tuning of Large Language ModelsKeming Lu, Hongyi Yuan, Zheng Yuan, Runji Lin 等ICLR 2024 · 被引用 119 次
- UCTopic: Unsupervised Contrastive Learning for Phrase Representations and Topic MiningJiacheng Li, Jingbo Shang, Julian J. McAuleyACL 2022 · 被引用 68 次
- Improving Word Translation via Two-Stage Contrastive LearningYaoyiran Li, Fangyu Liu, Nigel Collier, Anna Korhonen 等ACL 2022 · 被引用 32 次
- Drowzee: Metamorphic Testing for Fact-Conflicting Hallucination Detection in Large Language ModelsNingke Li, Yuekang Li, Yi Liu, Ling Shi 等OOPSLA 2024 · 被引用 26 次
- Efficient semantic uncertainty quantification in language models via diversity-steered samplingJi Won Park, Kyunghyun ChoNeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper7
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang 等EMNLP 2020 · 被引用 538 次
- Interpreting Pretrained Contextualized Representations via Reductions to Static EmbeddingsRishi Bommasani, Kelly Davis, Claire CardieACL 2020 · 被引用 137 次
- Assessing Phrasal Representation and Composition in TransformersLang Yu, Allyson EttingerEMNLP 2020 · 被引用 60 次
- Reformulating Unsupervised Style Transfer as Paraphrase GenerationKalpesh Krishna, John Wieting, Mohit IyyerEMNLP 2020 · 被引用 9 次
相关 Paper
- Sentence Representation Learning with Generative Objective rather than Contrastive ObjectiveBohong Wu, Hai ZhaoEMNLP 2022 · 被引用 3 次
- Topic Modeling as Multi-Objective Contrastive OptimizationThong Thanh Nguyen, Xiaobao Wu, Xinshuai Dong, Cong-Duy T. Nguyen 等ICLR 2024 · 被引用 13 次
- Self-Guided Contrastive Learning for BERT Sentence RepresentationsTaeuk Kim, Kang Min Yoo, Sang-goo LeeACL 2021
- Composition-contrastive Learning for Sentence EmbeddingsSachin Chanchani, Ruihong HuangACL 2023 · 被引用 11 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
