AlephBERT: Language Model Pre-training and Evaluation from Sub-Word to Sentence Level
Amit Seker, Elron Bandel, Dan Bareket, Idan Brusilovsky, Refael Shaked Greenfeld, Reut Tsarfaty
Abstract
Large Pre-trained Language Models (PLMs) have become ubiquitous in the development of language understanding technology and lie at the heart of many artificial intelligence advances. While advances reported for English using PLMs are unprecedented, reported advances using PLMs for Hebrew are few and far between. The problem is twofold. First, so far, Hebrew resources for training large language models are not of the same magnitude as their English counterparts. Second, most benchmarks available to evaluate progress in Hebrew NLP require morphological boundaries which are not available in the output of PLMs. In this work we remedy both aspects. We present AlephBERT, a large PLM for Modern Hebrew, trained on larger vocabulary and a larger dataset than any Hebrew PLM before. Moreover, we introduce a novel neural architecture that recovers the morphological segments encoded in contextualized embedding vectors. Based on this new morphological component we offer an evaluation suite consisting of multiple tasks and benchmarks that cover sentencelevel, word-level and sub-word level analyses. On all tasks, AlephBERT obtains state-of-theart results beyond contemporary Hebrew stateof-the-art models. We make our AlephBERT model, the morphological extraction component, and the Hebrew evaluation suite publicly available, for future investigations and evaluations of Hebrew PLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- SSPAttack: A Simple and Sweet Paradigm for Black-Box Hard-Label Textual Adversarial AttackHan Liu, Zhi Xu, Xiaotong Zhang, Xiaoming Xu et al.AAAI 2023 · 31 citations
- Structural Priming Demonstrates Abstract Grammatical Representations in Multilingual Language ModelsJames A. Michaelov, Catherine Arnett, Tyler A. Chang, Ben BergenEMNLP 2023 · 6 citations
- A Second Wave of UD Hebrew Treebanking and Cross-Domain ParsingAmir Zeldes, Nick Howell, Noam Ordan, Yifat Ben MosheEMNLP 2022 · 6 citations
- Beyond Word Boundaries: A Hebrew Coreference Benchmark and an Evaluation Protocol for Morphologically Complex TextRefael Shaked Greenfeld, Reut TsarfatyACL 2026
- The Enemy from Within: A Study of Political Delegitimization Discourse in Israeli Political SpeechNaama Rivlin-Angert, Guy Mor-LanEMNLP 2025
Builds on4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource LanguagesPedro Javier Ortiz Suárez, Laurent Romary, Benoît SagotACL 2020 · 72 citations
- KLEJ: Comprehensive Benchmark for Polish Language UnderstandingPiotr Rybak, Robert Mroczkowski, Janusz Tracz, Ireneusz GawlikACL 2020 · 4 citations
- From SPMRL to NMRL: What Did We Learn (and Unlearn) in a Decade of Parsing Morphologically-Rich Languages (MRLs)?Reut Tsarfaty, Dan Bareket, Stav Klein, Amit SekerACL 2020 · 2 citations
Related papers
- ARBERT & MARBERT: Deep Bidirectional Transformers for ArabicMuhammad Abdul-Mageed, AbdelRahim A. Elmadany, El Moatez Billah NagoudiACL 2021
- Revisiting Pre-trained Language Models and their Evaluation for Arabic Natural Language ProcessingAbbas Ghaddar, Yimeng Wu, Sunyam Bagga, Ahmad Rashid et al.EMNLP 2022 · 17 citations
- Superbizarre Is Not Superb: Derivational Morphology Improves BERT's Interpretation of Complex WordsValentin Hofmann, Janet B. Pierrehumbert, Hinrich SchützeACL 2021
- DagoBERT: Generating Derivational Morphology with a Pretrained Language ModelValentin Hofmann, Janet B. Pierrehumbert, Hinrich SchützeEMNLP 2020 · 1 citation
- BERT-like Models for Slavic Morpheme SegmentationDmitry Morozov, Lizaveta Astapenka, Anna V. Glazkova, Timur Garipov et al.ACL 2025 · 1 citation
