Lune

EMNLP2025Top-tier venue

Training compute-optimal transformer encoder models

Megi Dervishi, Alexandre Allauzen, Gabriel Synnaeve, Yann LeCun

2025Year
1Citations
1Top-tier citations

Abstract

Transformer encoders are critical for a wide range of Natural Language Processing (NLP) tasks, yet their compute-efficiency remains poorly understood. We present the first comprehensive empirical investigation of compute-optimal pretraining for encoder transformers using the Masked Language Modeling (MLM) objective. Across hundreds of carefully controlled runs we vary model size, data size, batch size, learning rate, and masking ratio, with increasing compute budget. The compute-optimal data-to-model ratio of Transformer encoder models is 10 to 100 times larger than the ratio of auto-regressive models. Using these recipes, we train OptiBERT, a family of compute-optimal BERT-style models that matches or surpasses leading baselines-including ModernBERT and NeoBERT-on GLUE and MTEB while training with dramatically less FLOPS.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 29c8365b-3512-4c20-bc6f-2df4eaea41b0

Cited by top-tier papers1

Ask how each one uses it

Builds on21

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines