ALLaM: Large Language Models for Arabic and English
M. Saiful Bari, Yazeed Alnumay, Norah A. Alzahrani, Nouf M. Alotaibi, Hisham Abdullah Alyahya, Sultan Alrashed, Faisal Abdulrahman Mirza, Shaykhah Z. Alsubaie, Hassan A. Alahmed, Ghadah Alabduljabbar, Raghad Alkhathran, Yousef Almushayqih
Abstract
We present ALLaM: Arabic Large Language Model, a series of large language models to support the ecosystem of Arabic Language Technologies (ALT). ALLaM is carefully trained considering the values of language alignment and knowledge transfer at scale. Our autoregressive decoder-only architecture models demonstrate how second-language acquisition via vocabulary expansion and pretraining on a mixture of Arabic and English text can steer a model towards a new language (Arabic) without any catastrophic forgetting in the original language (English). Furthermore, we highlight the effectiveness of using parallel/translated data to aid the process of knowledge alignment between languages. Finally, we show that extensive alignment with human preferences can significantly enhance the performance of a language model compared to models of a larger scale with lower quality alignment. ALLaM achieves state-of-the-art performance in various Arabic benchmarks, including MMLU Arabic, ACVA, and Arabic Exams. Our aligned models improve both in Arabic and English from their base aligned models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 85d3ae01-adf3-4960-bab3-c2d6d2dc26ecCited by top-tier papers8
- Lost in the Mix: Evaluating LLM Understanding of Code-Switched TextAmr Mohamed, Yang Zhang, Michalis Vazirgiannis, Guokan ShangACL 2026 · 9 citations
- MedAraBench: Large-scale Arabic Medical Question Answering Dataset and BenchmarkMouath Abu Daoud, Leen Kharouf, Omar El Hajj, Dana El Samad et al.ICLR 2026 · 4 citations
- Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMsWafa Al Ghallabi, Ritesh Thawkar, Sara Ghaboura, Ketan Pravin More et al.EMNLP 2025 · 1 citation
- OntologyRAG-Q: Resource Development and Benchmarking for Retrieval-Augmented Question Answering in Qur'anic TafsirSadam Al-Azani, Maad Alowaifeer, Alhanoof Alhunief, Ahmed AbdelaliEMNLP 2025 · 1 citation
- SAHM: A Benchmark for Arabic Financial and Shari'ah-Compliant ReasoningRania Elbadry, Sarfraz Ahmad, Ahmed Heakl, Dani Bouch et al.ACL 2026 · 1 citation
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
Related papers
- Alignment at Pre-training! Towards Native Alignment for Arabic LLMsJuhao Liang, Zhenyang Cai, Jianqing Zhu, Huang Huang et al.NeurIPS 2024 · 19 citations
- RLHF Can Speak Many Languages: Unlocking Multilingual Preference Optimization for LLMsJohn Dang, Arash Ahmadian, Kelly Marchisio, Julia Kreutzer et al.EMNLP 2024 · 4 citations
- LifeAlign: Lifelong Alignment for Large Language Models with Memory-Augmented Focalized Preference OptimizationJunsong Li, Jie Zhou, Bihao Zhan, Yutao Yang et al.AAAI 2026 · 3 citations
- Revisiting Pre-trained Language Models and their Evaluation for Arabic Natural Language ProcessingAbbas Ghaddar, Yimeng Wu, Sunyam Bagga, Ahmad Rashid et al.EMNLP 2022 · 17 citations
- ARBERT & MARBERT: Deep Bidirectional Transformers for ArabicMuhammad Abdul-Mageed, AbdelRahim A. Elmadany, El Moatez Billah NagoudiACL 2021
