A Bilingual Generative Transformer for Semantic Sentence Embedding
John Wieting, Graham Neubig, Taylor Berg-Kirkpatrick
Abstract
Semantic sentence embedding models encode natural language sentences into vectors, such that closeness in embedding space indicates closeness in the semantics between the sentences. Bilingual data offers a useful signal for learning such embeddings: properties shared by both sentences in a translation pair are likely semantic, while divergent properties are likely stylistic or language-specific. We propose a deep latent variable model that attempts to perform source separation on parallel sentences, isolating what they have in common in a latent semantic vector, and explaining what is left over with language-specific latent vectors. Our proposed approach differs from past work on semantic sentence encoding in two ways. First, by using a variational probabilistic framework, we introduce priors that encourage source separation, and can use our model's posterior to predict sentence embeddings for monolingual data at test time. Second, we use high-capacity transformers as both data generating distributions and inference networkscontrasting with most past work on sentence embeddings. In experiments, our approach substantially outperforms the state-of-the-art on a standard suite of unsupervised semantic similarity evaluations. Further, we demonstrate that our approach yields the largest gains on more difficult subsets of these evaluations where simple word overlap is not a good indicator of similarity. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e67fecac-15d4-443e-b80f-6fb4a64cf0d9Cited by top-tier papers6
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- ConvFiT: Conversational Fine-Tuning of Pretrained Language ModelsIvan Vulic, Pei-Hao Su, Samuel Coope, Daniela Gerz et al.EMNLP 2021 · 30 citations
- Language-agnostic Representation from Multilingual Sentence Encoders for Cross-lingual Similarity EstimationNattapong Tiyajamorn, Tomoyuki Kajiwara, Yuki Arase, Makoto OnizukaEMNLP 2021 · 16 citations
- Do We Run How We Say We Run? Formalization and Practice of Governance in OSS CommunitiesMahasweta Chakraborti, Curtis Atkisson, Stefan Stanciulescu, Vladimir Filkov et al.CHI 2024 · 6 citations
- Retrofitting Multilingual Sentence Embeddings with Abstract Meaning RepresentationDeng Cai, Xin Li, Jackie Chun-Sing Ho, Lidong Bing et al.EMNLP 2022 · 4 citations
Related papers
- Beyond Contrastive Learning: A Variational Generative Model for Multilingual RetrievalJohn Wieting, Jonathan H. Clark, William W. Cohen, Graham Neubig et al.ACL 2023 · 3 citations
- Language-agnostic BERT Sentence EmbeddingFangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan et al.ACL 2022
- Cross-lingual Sentence Embedding using Multi-Task LearningKoustava Goswami, Sourav Dutta, Haytham Assem, Theodorus Fransen et al.EMNLP 2021 · 9 citations
- Making Monolingual Sentence Embeddings Multilingual using Knowledge DistillationNils Reimers, Iryna GurevychEMNLP 2020 · 54 citations
- DeCLUTR: Deep Contrastive Learning for Unsupervised Textual RepresentationsJohn M. Giorgi, Osvald Nitski, Bo Wang, Gary D. BaderACL 2021
