Unsupervised Domain Clusters in Pretrained Language Models
Roee Aharoni, Yoav Goldberg
Abstract
The notion of "in-domain data" in NLP is often over-simplistic and vague, as textual data varies in many nuanced linguistic aspects such as topic, style or level of formality. In addition, domain labels are many times unavailable, making it challenging to build domainspecific systems. We show that massive pretrained language models implicitly learn sentence representations that cluster by domains without supervision -suggesting a simple datadriven definition of domains in textual data. We harness this property and propose domain data selection methods based on such models, which require only a small set of in-domain monolingual data. We evaluate our data selection methods for neural machine translation across five diverse domains, where they outperform an established approach as measured by both BLEU and by precision and recall of sentence selection with respect to an oracle.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3be44db1-104c-421d-90cf-bf833d159e4fCited by top-tier papers51
- MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment AnalysisDevamanyu Hazarika, Roger Zimmermann, Soujanya PoriaACM MM 2020 · 1,037 citations
- Nearest Neighbor Machine TranslationUrvashi Khandelwal, Angela Fan, Dan Jurafsky, Luke Zettlemoyer et al.ICLR 2021 · 323 citations
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo et al.ACL 2020 · 93 citations
- Goal Driven Discovery of Distributional Differences via Language DescriptionsRuiqi Zhong, Peter Zhang, Steve Li, Jinwoo Ahn et al.NeurIPS 2023 · 81 citations
- Neuro-Symbolic Language Modeling with Automaton-augmented RetrievalUri Alon, Frank F. Xu, Junxian He, Sudipta Sengupta et al.ICML 2022 · 79 citations
Builds on3
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Emerging Cross-lingual Structure in Pretrained Language ModelsAlexis Conneau, Shijie Wu, Haoran Li, Luke Zettlemoyer et al.ACL 2020 · 210 citations
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo et al.ACL 2020 · 93 citations
Related papers
- Generalised Unsupervised Domain Adaptation of Neural Machine Translation with Cross-Lingual Data SelectionThuy-Trang Vu, Xuanli He, Dinh Q. Phung, Gholamreza HaffariEMNLP 2021 · 2 citations
- The Inductive Bias of In-Context Learning: Rethinking Pretraining Example DesignYoav Levine, Noam Wies, Daniel Jannai, Dan Navon et al.ICLR 2022 · 43 citations
- Entity Extraction in Low Resource Domains with Selective Pre-training of Large Language ModelsAniruddha Mahapatra, Sharmila Reddy Nangi, Aparna Garimella, Anandhavelu NatarajanEMNLP 2022 · 4 citations
- LAMDAS: LLM as an Implicit Classifier for Domain-specific Data SelectionJian Wu, Hang Yu, Bingchang Liu, Wenjie Yang et al.AAAI 2026 · 1 citation
- Enhancing Neural Machine Translation Through Target Language Data: A kNN-LM Approach for Domain AdaptationAbudurexiti Reheman, Hongyu Liu, Junhao Ruan, Abudukeyumu Abudula et al.ACL 2025
