A Second Wave of UD Hebrew Treebanking and Cross-Domain Parsing
Amir Zeldes, Nick Howell, Noam Ordan, Yifat Ben Moshe
Abstract
Foundational Hebrew NLP tasks such as segmentation, tagging and parsing, have relied to date on various versions of the Hebrew Treebank (HTB, Sima'an et al. 2001). However, the data in HTB, a single-source newswire corpus, is now over 30 years old, and does not cover many aspects of contemporary Hebrew on the web. This paper presents a new, freely available UD treebank of Hebrew stratified from a range of topics selected from Hebrew Wikipedia. In addition to introducing the corpus and evaluating the quality of its annotations, we deploy automatic validation tools based on grew (Guillaume, 2021), and conduct the first cross domain parsing experiments in Hebrew. We obtain new state-of-the-art (SOTA) results on UD NLP tasks, using a combination of the latest language modelling and some incremental improvements to existing transformer based approaches. We also release a new version of the UD HTB matching annotation scheme updates from our new corpus.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b1b30aaf-0ed2-458c-83dc-fabe9e856bf3Builds on3
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Predictive Biases in Natural Language Processing Models: A Conceptual Framework and OverviewDeven Shah, H. Andrew Schwartz, Dirk HovyACL 2020 · 93 citations
- AlephBERT: Language Model Pre-training and Evaluation from Sub-Word to Sentence LevelAmit Seker, Elron Bandel, Dan Bareket, Idan Brusilovsky et al.ACL 2022 · 53 citations
Related papers
- The UD-NewsCrawl Treebank: Reflections and Challenges from a Large-scale Tagalog Syntactic Annotation ProjectAngelina Aspra Aquino, Lester James Validad Miranda, Elsie Marie T. OrACL 2025
- GDTB: Genre Diverse Data for English Shallow Discourse Parsing across Modalities, Text Types, and DomainsYang Janet Liu, Tatsuya Aoyama, Wesley Scivetti, Yilun Zhu et al.EMNLP 2024
- Revisiting Supertagging for faster HPSG parsingOlga Zamaraeva, Carlos Gómez-RodríguezEMNLP 2024
- "Wikily" Supervised Neural Translation Tailored to Cross-Lingual TasksMohammad Sadegh Rasooli, Chris Callison-Burch, Derry Tanti WijayaEMNLP 2021 · 3 citations
- Domain-Adapted Dependency Parsing for Cross-Domain Named Entity RecognitionChenxiao Dou, Xianghui Sun, Yaoshu Wang, Yunjie Ji et al.AAAI 2023 · 8 citations
