Genre as Weak Supervision for Cross-lingual Dependency Parsing
Max Müller-Eberstein, Rob van der Goot, Barbara Plank
Abstract
Recent work has shown that monolingual masked language models learn to represent data-driven notions of language variation which can be used for domain-targeted training data selection. Dataset genre labels are already frequently available, yet remain largely unexplored in cross-lingual setups. We harness this genre metadata as a weak supervision signal for targeted data selection in zeroshot dependency parsing. Specifically, we project treebank-level genre information to the finer-grained sentence level, with the goal to amplify information implicitly stored in unsupervised contextualized representations. We demonstrate that genre is recoverable from multilingual contextual embeddings and that it provides an effective signal for training data selection in cross-lingual, zero-shot scenarios. For 12 low-resource language treebanks, six of which are test-only, our genre-specific methods significantly outperform competitive baselines as well as recent embedding-based methods for data selection. Moreover, genre-based data selection provides new state-of-the-art results for three of these target languages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aee4fcd1-4f3d-4646-a7c8-96598decbd3dBuilds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- From Zero to Hero: On the Limitations of Zero-Shot Language Transfer with Multilingual TransformersAnne Lauscher, Vinit Ravishankar, Ivan Vulic, Goran GlavasEMNLP 2020 · 235 citations
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo et al.ACL 2020 · 93 citations
- Unsupervised Domain Clusters in Pretrained Language ModelsRoee Aharoni, Yoav GoldbergACL 2020 · 13 citations
Related papers
- Language Embeddings for Typology and Cross-lingual Transfer LearningDian Yu, Taiqi He, Kenji SagaeACL 2021
- Finding Universal Grammatical Relations in Multilingual BERTEthan A. Chi, John Hewitt, Christopher D. ManningACL 2020 · 7 citations
- Revisiting Tri-training of Dependency ParsersJoachim Wagner, Jennifer FosterEMNLP 2021
- How do languages influence each other? Studying cross-lingual data sharing during LM fine-tuningRochelle Choenni, Dan Garrette, Ekaterina ShutovaEMNLP 2023 · 2 citations
- Unsupervised Interlingual Semantic Representations from Sentence Embeddings for Zero-Shot Cross-Lingual TransferChanny Hong, Jaeyeon Lee, Jungkwon LeeAAAI 2020 · 1 citation
