Nugget: Neural Agglomerative Embeddings of Text
Guanghui Qin, Benjamin Van Durme
Abstract
Embedding text sequences is a widespread requirement in modern language understanding. Existing approaches focus largely on constant-size representations. This is problematic, as the amount of information contained in text often varies with the length of the input. We propose a solution called Nugget, which encodes language into a representation based on a dynamically selected subset of input tokens. These nuggets are learned through tasks like autoencoding and machine translation, and intuitively segment language into meaningful units. We demonstrate Nugget outperforms related approaches in tasks involving semantic comparison. Finally, we illustrate these compact units allow for expanding the contextual window of a language model (LM), suggesting new future LMs that can condition on significantly larger amounts of content.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e2e0860-12cf-4368-ad84-29209024b8dcCited by top-tier papers11
- In-context Autoencoder for Context Compression in a Large Language ModelTao Ge, Jing Hu, Lei Wang, Xun Wang et al.ICLR 2024 · 158 citations
- Efficient Large Multi-modal Models via Visual Context CompressionJieneng Chen, Luoxin Ye, Ju He, Zhaoyang Wang et al.NeurIPS 2024 · 49 citations
- Multi-Vector Index Compression in Any ModalityHanxiang Qin, Alexander Martin, Rohan Jha, Chunsheng Zuo et al.SIGIR 2026 · 11 citations
- Optimizing Retrieval-augmented Reader Models via Token EliminationMoshe Berchansky, Peter Izsak, Avi Caciularu, Ido Dagan et al.EMNLP 2023 · 4 citations
- By Tying Embeddings You Are Assuming the Distributional HypothesisFrancesco Bertolotti, Walter CazzolaICML 2024 · 4 citations
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang et al.NeurIPS 2022 · 1,546 citations
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier et al.ICLR 2020 · 833 citations
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang et al.EMNLP 2020 · 538 citations
Related papers
- Efficient Transformers with Dynamic Token PoolingPiotr Nawrot, Jan Chorowski, Adrian Lancucki, Edoardo Maria PontiACL 2023 · 14 citations
- Pre-training Universal Language RepresentationYian Li, Hai ZhaoACL 2021
- Dataset Decomposition: Faster LLM Training with Variable Sequence Length CurriculumHadi Pouransari, Chun-Liang Li, Jen-Hao Rick Chang, Pavan Kumar Anasosalu Vasu et al.NeurIPS 2024 · 40 citations
- Copy is All You NeedTian Lan, Deng Cai, Yan Wang, Heyan Huang et al.ICLR 2023
- Cramming 1568 Tokens into a Single Vector and Back Again: Exploring the Limits of Embedding Space CapacityYuri Kuratov, Mikhail Arkhipov, Aydar Bulatov, Mikhail BurtsevACL 2025
