Lune

ICML2026Top-tier venue

L3L^3: Large Lookup Layers

Albert Tseng, Chris De Sa

2026Year
2Citations
2Top-tier citations

Abstract

Modern sparse language models typically achieve sparsity through Mixture-of-Experts (MoE) layers, which dynamically route tokens to dense MLP "experts." However, dynamic hard routing has a number of drawbacks, such as potentially poor hardware efficiency and needing auxiliary losses for stable training. In contrast, the tokenizer embedding table, which is natively sparse, largely avoids these issues by selecting a single embedding per token at the cost of not having contextual information. In this work, we introduce the Large Lookup Layer (L 3 ), which generalizes embedding tables to model decoder layers as a means of further scaling sparsity. L 3 layers use static token-based routing to aggregate a set of learned embeddings per token in a contextdependent way, allowing the model to efficiently balance memory and compute by caching information in embeddings. L 3 has two main components: (1) a systems-friendly architecture that allows for fast training and CPU-offloaded inference with no overhead, and (2) an information-theoretic embedding allocation algorithm that effectively balances speed and quality. We empirically test L 3 by training transformers with up to 2.6B active parameters and find that L 3 strongly outperforms both dense models and iso-sparse MoEs in both language modeling and downstream tasks.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 73ea276e-9569-4a65-b32b-cf373cba3c0c

Cited by top-tier papers2

Ask how each one uses it

Builds on12

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines