Frozen LLMs are Native Decoders for High-Norm Semantic Vectors
Yunsheng Zeng, Yongmei Tan
Abstract
Large language models (LLMs) are designed for discrete tokens, yet they operate in a continuous embedding space. Recent context compression methods exploit this property by encoding text into dense vectors for frozen LLM decoding. However, a key question remains unanswered: how does a frozen LLM interpret continuous vectors that encode complex semantics? We investigate this through controlled reconstruction experiments. Our analysis reveals a critical geometric property: successful compression encoders learn to produce vectors with L2 norms two orders of magnitude higher than standard embeddings. Norm-scaling interventions provide strong evidence that this highnorm regime is an enabling factor for frozen-LLM decoding, while leaving open whether the effect arises from attention dominance, lownorm suppression, or both. Based on this finding, we propose a landmark-based compression framework for long contexts. Our encoder uses bidirectional attention over landmark tokens, which captures global dependencies and avoids semantic fragmentation from segment-based methods. Experiments on text reconstruction and four QA benchmarks provide evidence for our method. At 4x compression, our method is strongest on SQuAD and AdversarialQA and remains competitive on average.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5c9a4dfc-d723-4e1c-a778-6011e776f49dBuilds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- In-context Autoencoder for Context Compression in a Large Language ModelTao Ge, Jing Hu, Lei Wang, Xun Wang et al.ICLR 2024 · 158 citations
- xRAG: Extreme Context Compression for Retrieval-augmented Generation with One TokenXin Cheng, Xun Wang, Xingxing Zhang, Tao Ge et al.NeurIPS 2024 · 156 citations
- The Power of Scale for Parameter-Efficient Prompt TuningBrian Lester, Rami Al-Rfou, Noah ConstantEMNLP 2021 · 94 citations
Related papers
- Latent-Condensed Transformer for Efficient Long Context ModelingZeng You, Yaofo Chen, Qiuwu Chen, Ying Sun et al.ACL 2026
- GMSA: Enhancing Context Compression via Group Merging and Layer Semantic AlignmentJiwei Tang, Zhicheng Zhang, Shunlong Wu, Jingheng Ye et al.ACL 2026 · 24 citations
- A Simple and Effective L_2 Norm-Based Strategy for KV Cache CompressionAlessio Devoto, Yu Zhao, Simone Scardapane, Pasquale MinerviniEMNLP 2024 · 3 citations
- EAKV: An Entropy-Driven Adaptive KV Compression Framework for Long Video UnderstandingHengrui Hu, Jingyu Li, Juntao Liang, Guanyu Chen et al.ICML 2026
- Autoencoding-Free Context Compression for LLMs via Contextual Semantic AnchorsXin Liu, Runsong Zhao, Pengcheng Huang, Xinyu Liu et al.ICLR 2026 · 16 citations
