NextLevelBERT: Masked Language Modeling with Higher-Level Representations for Long Documents
Tamara Czinczoll, Christoph Hönes, Maximilian Schall, Gerard de Melo
Abstract
While (large) language models have significantly improved over the last years, they still struggle to sensibly process long sequences found, e.g., in books, due to the quadratic scaling of the underlying attention mechanism. To address this, we propose NextLevelBERT, a Masked Language Model operating not on tokens, but on higher-level semantic representations in the form of text embeddings. We pretrain NextLevelBERT to predict the vector representation of entire masked text chunks and evaluate the effectiveness of the resulting document vectors on three types of tasks: 1) Semantic Textual Similarity via zero-shot document embeddings, 2) Long document classification, 3) Multiple-choice question answering. We find that next-level Masked Language Modeling is an effective technique to tackle long-document use cases and can outperform much larger embedding models as long as the required level of detail of semantic information is not too fine. Our models and code are publicly available online.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3d139f5b-b26e-481b-b649-24cd1436ec39Cited by top-tier papers1
Ask how each one uses itBuilds on12
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
Related papers
- Hierarchical Token Prepending: Enhancing Information Flow in Decoder-based LLM EmbeddingsXueying Ding, Xingyue Huang, Mingxuan Ju, Liam Collins et al.ACL 2026 · 3 citations
- LinkBERT: Pretraining Language Models with Document LinksMichihiro Yasunaga, Jure Leskovec, Percy LiangACL 2022 · 463 citations
- Dynamic Chunking and Selection for Reading Comprehension of Ultra-Long Context in Large Language ModelsBoheng Sheng, Jiacheng Yao, Meicong Zhang, Guoxiu HeACL 2025 · 9 citations
- SenseBERT: Driving Some Sense into BERTYoav Levine, Barak Lenz, Or Dagan, Ori Ram et al.ACL 2020 · 27 citations
- Modeling Sequential Sentence Relation to Improve Cross-lingual Dense RetrievalShunyu Zhang, Yaobo Liang, Ming Gong, Daxin Jiang et al.ICLR 2023 · 1 citation
