Scaling Embedding Layers in Language Models
Da Yu, Edith Cohen, Badih Ghazi, Yangsibo Huang, Pritish Kamath, Ravi Kumar, Daogao Liu, Chiyuan Zhang
摘要
We propose (calable, ontextualized, ffloaded, -gram mbedding), a new method for extending input embedding layers to enhance language model performance. To avoid increased decoding costs, retains the original vocabulary while introducing embeddings for a set of frequent n-grams. These embeddings provide contextualized representation for each input token and are learned with a separate model during training. After training, embeddings are precomputed and stored in off-accelerator memory; during inference, querying them has minimal impact on latency due to the low complexity of embedding lookups. enables two new scaling strategies: increasing the number of n-gram embeddings and scaling the model used to learn them, both while maintaining fixed accelerator usage during inference (in terms of FLOPS and memory). We show that scaling both aspects enables a model with 1B accelerator-resident parameters to outperform a 1.9B-parameter baseline across diverse corpora, while using only about half the FLOPS and accelerator memory during inference.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language ModelsXin Cheng, Wangding Zeng, Damai Dai, Qinyu Chen 等ACL 2026 · 被引用 57 次
- Exploiting Vocabulary Frequency Imbalance in Language Model Pre-trainingWoojin Chung, Jeonghoon KimNeurIPS 2025 · 被引用 6 次
- Proxy Compression for Language ModelingLin Zheng, Li Xinyu, Qian Liu, Xiachong Feng 等ICML 2026 · 被引用 3 次
- : Large Lookup LayersAlbert Tseng, Chris De SaICML 2026 · 被引用 2 次
- Inner-layer Token Self-modulation as Another Scaling Axis for LLMsYebin Yang, Huaijin Wu, Jingtao Han, Yu Wang 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper17
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- MEGABYTE: Predicting Million-byte Sequences with Multiscale TransformersLili Yu, Daniel Simig, Colin Flaherty, Armen Aghajanyan 等NeurIPS 2023 · 被引用 197 次
- Scaling Laws with Vocabulary: Larger Models Deserve Larger VocabulariesChaofan Tao, Qian Liu, Longxu Dou, Niklas Muennighoff 等NeurIPS 2024 · 被引用 135 次
- Byte Latent Transformer: Patches Scale Better Than TokensArtidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodríguez, John Nguyen 等ACL 2025 · 被引用 116 次
相关 Paper
- Parallel Scaling Law for Language ModelsMouxiang Chen, Binyuan Hui, Zeyu Cui, Jiaxi Yang 等NeurIPS 2025 · 被引用 33 次
- Scaling Attention Beyond GPUs for LLM InferenceWeishu Deng, Yujie Yang, Peiran Du, Lingfeng Xiang 等HPDC 2026
- Context-level Language Modeling by Learning Predictive Context Embeddingsbeiya dai, Yuliang Liu, Yunchong Song, Daozheng Xue 等ICML 2026 · 被引用 5 次
- SpecExec: Massively Parallel Speculative Decoding For Interactive LLM Inference on Consumer DevicesRuslan Svirschevski, Avner May, Zhuoming Chen, Beidi Chen 等NeurIPS 2024 · 被引用 70 次
- LLM in a flash: Efficient Large Language Model Inference with Limited MemoryKeivan Alizadeh, Iman Mirzadeh, Dmitry Belenko, S. Khatamifard 等ACL 2024 · 被引用 73 次
