Alignment at Pre-training! Towards Native Alignment for Arabic LLMs
Juhao Liang, Zhenyang Cai, Jianqing Zhu, Huang Huang, Kewei Zong, Bang An, Mosen Alharthi, Juncai He, Lian Zhang, Haizhou Li, Benyou Wang, Jinchao Xu
Abstract
The alignment of large language models (LLMs) is critical for developing effective and safe language models. Traditional approaches focus on aligning models during the instruction tuning or reinforcement learning stages, referred to in this paper as post alignment'. We argue that alignment during the pre-training phase, which we term native alignment', warrants investigation. Native alignment aims to prevent unaligned content from the beginning, rather than relying on post-hoc processing. This approach leverages extensively aligned pre-training data to enhance the effectiveness and usability of pre-trained models. Our study specifically explores the application of native alignment in the context of Arabic LLMs. We conduct comprehensive experiments and ablation studies to evaluate the impact of native alignment on model performance and alignment stability. Additionally, we release open-source Arabic LLMs that demonstrate state-of-the-art performance on various benchmarks, providing significant benefits to the Arabic LLM community.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eee08f2f-ac52-4716-8005-3daaa31eacefCited by top-tier papers2
- AraEval: An Arabic Multi-Task Evaluation Suite for Large Language ModelsAlhanoof Althnian, Norah A. Alzahrani, Shaykhah Z. Alsubaie, Eman Albilali et al.EMNLP 2025
- The Stackelberg Speaker: Optimizing Persuasive Communication in Social Deduction GamesZheng Zhang, Deheng Ye, Peilin Zhao, Hao WangACL 2026
Builds on10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Llemma: An Open Language Model for MathematicsZhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos et al.ICLR 2024 · 433 citations
Related papers
- ALLaM: Large Language Models for Arabic and EnglishM. Saiful Bari, Yazeed Alnumay, Norah A. Alzahrani, Nouf M. Alotaibi et al.ICLR 2025 · 4 citations
- Getting More from Less: Large Language Models are Good Spontaneous Multilingual LearnersShimao Zhang, Changjiang Gao, Wenhao Zhu, Jiajun Chen et al.EMNLP 2024 · 1 citation
- Gaining Wisdom from Setbacks: Aligning Large Language Models via Mistake AnalysisKai Chen, Chunwei Wang, Kuo Yang, Jianhua Han et al.ICLR 2024 · 47 citations
- Safety Pretraining: Toward the Next Generation of Safe AIPratyush Maini, Sachin Goyal, Dylan Sam, Alexander Robey et al.NeurIPS 2025 · 50 citations
- PreAlign: Boosting Cross-Lingual Transfer by Early Establishment of Multilingual AlignmentJiahuan Li, Shujian Huang, Aarron Ching, Xinyu Dai et al.EMNLP 2024 · 5 citations
