SLlama: Parameter-Efficient Language Model Architecture for Enhanced Linguistic Competence Under Strict Data Constraints
Victor Adelakun Omolaoye, Babajide Alamu Owoyele, Gerard de Melo
Abstract
Scaling data and model size has driven recent advances in language modeling, but this strategy falters under scenarios with strict data constraints, as in the BabyLM Challenge. However, insights from training compute-optimal large language models highlight that smaller models trained on more data outperform larger counterparts trained inadequately, emphasizing the need for compact architectures. Furthermore, while embedding weight tying is a common parameter-reduction technique, we find that it significantly diminishes linguistic competence in compact models. In response, we explore alternative architectural strategies that preserve the parameter-efficiency of tied models without sacrificing the representational benefits of untied embeddings. Consequently, we introduce SLlama, a Llama-3 architecture variant that incorporates targeted modifications-Repeated Reduced Hidden Size and Projection (RRHP), Permutated Weight Attention (PWA), Shared Projection Multi-Layer Perceptron (SPMLP), and Layer Weight Sharing-to compress Transformer components. Without relying on distillation, SLlama achieves a 31.72% improvement in linguistic knowledge acquisition over the Baby Llama baseline, with a comparable GLUE score and significantly lower parameter count. These results demonstrate that welldesigned, compact models can rival larger ones under strict data constraints.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3c1c6bfa-41e6-4648-a2b9-a4c02e7af1adBuilds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang et al.ACL 2022 · 844 citations
Related papers
- Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRASangmin Bae, Adam Fisch, Hrayr Harutyunyan, Ziwei Ji et al.ICLR 2025
- Sheared LLaMA: Accelerating Language Model Pre-training via Structured PruningMengzhou Xia, Tianyu Gao, Zhiyuan Zeng, Danqi ChenICLR 2024 · 453 citations
- Headless Language Models: Learning without Predicting with Contrastive Weight TyingNathan Godey, Éric Villemonte de la Clergerie, Benoît SagotICLR 2024 · 5 citations
- Effective Distillation to Hybrid xLSTM ArchitecturesLukas Hauzenberger, Niklas Schmidinger, Thomas Schmied, Anamaria-Roberta Hartl et al.ICML 2026 · 3 citations
- BOLT: Fewer Tokens but More Performance Retention for Efficient Vision-Language Models InferenceJiahua Bao, Siyao Cheng, Jiaxing Du, Changjiang He et al.ACM MM 2025
