Separating Representation from Reconstruction Enables Scalable Text Encoders
Megi Dervishi, Mathurin VIDEAU, Yann LeCun
Abstract
While decoders have rapidly scaled, encoders have remained largely unchanged since BERT. We revisit this disparity by frozen backbone evaluation via probing. Under this lens, the representations of BERT encoders become increasingly unexploitable by frozen probes, despite improved perplexity. The misalignment originates in BERT's flat design, which couples representation learning to the token reconstruction loss. We propose CrossBERT, a two-part architecture that separates the learning of high-quality encoded representations from the rigid grounding of token reconstruction. This design further enables high masking ratios (%) and gradient collection over all tokens via a Complementary Masking Strategy, respectively increasing throughput by to and sample efficiency by . Overall, CrossBERT demonstrates monotonic scaling and superior performance on MTEB(eng, v2) and frozen GLUE benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a62dcf77-3455-4c0d-8705-e47f593010edBuilds on9
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised LearningAdrien Bardes, Jean Ponce, Yann LeCunICLR 2022 · 1,226 citations
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and InferenceBenjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller et al.ACL 2025 · 552 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
Related papers
- Rethinking JEPA: Compute‑Efficient Video Self-Supervised Learning with Frozen TeachersXianhang Li, Chen Huang, Chun-Liang Li, Eran Malach et al.ICLR 2026 · 6 citations
- MosaicBERT: A Bidirectional Encoder Optimized for Fast PretrainingJacob P. Portes, Alexander Trott, Sam Havens, Daniel King et al.NeurIPS 2023 · 46 citations
- KV-Embedding: Training-free Text Embedding via Internal KV Re-routing in Decoder-only LLMsYixuan Tang, Yi YangACL 2026 · 1 citation
- Training compute-optimal transformer encoder modelsMegi Dervishi, Alexandre Allauzen, Gabriel Synnaeve, Yann LeCunEMNLP 2025 · 1 citation
- Drop your Decoder: Pre-training with Bag-of-Word Prediction for Dense Passage RetrievalGuangyuan Ma, Xing Wu, Zijia Lin, Songlin HuSIGIR 2024 · 5 citations
