Lune

ICML2026Top-tier venue

Separating Representation from Reconstruction Enables Scalable Text Encoders

Megi Dervishi, Mathurin VIDEAU, Yann LeCun

2026Year

Abstract

While decoders have rapidly scaled, encoders have remained largely unchanged since BERT. We revisit this disparity by frozen backbone evaluation via probing. Under this lens, the representations of BERT encoders become increasingly unexploitable by frozen probes, despite improved perplexity. The misalignment originates in BERT's flat design, which couples representation learning to the token reconstruction loss. We propose CrossBERT, a two-part architecture that separates the learning of high-quality encoded representations from the rigid grounding of token reconstruction. This design further enables high masking ratios (≥50\geq 50%) and gradient collection over all tokens via a Complementary Masking Strategy, respectively increasing throughput by 1.51.5 to 2×2\times and sample efficiency by 2×2\times. Overall, CrossBERT demonstrates monotonic scaling and superior performance on MTEB(eng, v2) and frozen GLUE benchmarks.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext a62dcf77-3455-4c0d-8705-e47f593010ed

Builds on9

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines