Universal Sentence Representation Learning with Conditional Masked Language Model
Ziyi Yang, Yinfei Yang, Daniel Cer, Jax Law, Eric Darve
摘要
This paper presents a novel training method, Conditional Masked Language Modeling (CMLM), to effectively learn sentence representations on large scale unlabeled corpora. CMLM integrates sentence representation learning into MLM training by conditioning on the encoded vectors of adjacent sentences. Our English CMLM model achieves state-ofthe-art performance on SentEval (Conneau and Kiela, 2018), even outperforming models learned using supervised signals. As a fully unsupervised learning method, CMLM can be conveniently extended to a broad range of languages and domains. We find that a multilingual CMLM model co-trained with bitext retrieval (BR) and natural language inference (NLI) tasks outperforms the previous state-of-the-art multilingual models by a large margin, e.g. 10% improvement upon baseline models on cross-lingual semantic search. We explore the same language bias of the learned representations, and propose a simple, post-training and model agnostic approach to remove the language identifying information from the representation while still retaining sentence semantics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Learning to Generalize to More: Continuous Semantic Augmentation for Neural Machine TranslationXiangpeng Wei, Heng Yu, Yue Hu, Rongxiang Weng 等ACL 2022 · 被引用 26 次
- Autonomous LLM-Enhanced Adversarial Attack for Text-to-MotionHonglei Miao, Fan Ma, Ruijie Quan, Kun Zhan 等AAAI 2025 · 被引用 11 次
- DenoSent: A Denoising Objective for Self-Supervised Sentence Representation LearningXinghao Wang, Junliang He, Pengyu Wang, Yunhua Zhou 等AAAI 2024 · 被引用 11 次
- Global Selection of Contrastive Batches via Optimization on Sample PermutationsVin Sachidananda, Ziyi Yang, Chenguang ZhuICML 2023 · 被引用 6 次
- Dual-Alignment Pre-training for Cross-lingual Sentence EmbeddingZiheng Li, Shaohan Huang, Zihan Zhang, Zhi-Hong Deng 等ACL 2023 · 被引用 5 次
它引用的顶会 Paper7
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu 等ICLR 2020 · 被引用 1,170 次
- ETC: Encoding Long and Structured Inputs in TransformersJoshua Ainslie, Santiago Ontañón, Chris Alberti, Vaclav Cvicek 等EMNLP 2020 · 被引用 268 次
- Making Monolingual Sentence Embeddings Multilingual using Knowledge DistillationNils Reimers, Iryna GurevychEMNLP 2020 · 被引用 54 次
相关 Paper
- Language-agnostic BERT Sentence EmbeddingFangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan 等ACL 2022
- Bilingual alignment transfers to multilingual alignment for unsupervised parallel text miningChih-chan Tien, Shane Steinert-ThrelkeldACL 2022 · 被引用 10 次
- Modeling Sequential Sentence Relation to Improve Cross-lingual Dense RetrievalShunyu Zhang, Yaobo Liang, Ming Gong, Daxin Jiang 等ICLR 2023 · 被引用 1 次
- Discovering Low-rank Subspaces for Language-agnostic Multilingual RepresentationsZhihui Xie, Handong Zhao, Tong Yu, Shuai LiEMNLP 2022 · 被引用 3 次
- Distilling Knowledge Learned in BERT for Text GenerationYen-Chun Chen, Zhe Gan, Yu Cheng, Jingzhou Liu 等ACL 2020 · 被引用 116 次
