LLM-based Embeddings: Attention Values Encode Sentence Semantics Better Than Hidden States
Yeqin Zhang, Yunfei Wang, Jiaxuan Chen, Ke Qin, Yizheng Zhao, Cam-Tu Nguyen
Abstract
Sentence representations are foundational to many Natural Language Processing (NLP) applications. While recent methods leverage Large Language Models (LLMs) to derive sentence representations, most rely on final-layer hidden states, which are optimized for next-token prediction and thus often fail to capture global, sentence-level semantics. This paper introduces a novel perspective, demonstrating that attention value vectors capture sentence semantics more effectively than hidden states. We propose Value Aggregation (VA), a simple method that pools token values across multiple layers and token indices. In a training-free setting, VA outperforms other LLM-based embeddings, even matches or surpasses the ensemble-based MetaEOL. Furthermore, we demonstrate that when paired with suitable prompts, the layer attention outputs can be interpreted as aligned weighted value vectors. Specifically, the attention scores of the last token function as the weights, while the output projection matrix () aligns these weighted value vectors with the common space of the LLM residual stream. This refined method, termed Aligned Weighted VA (AlignedWVA), achieves state-of-the-art performance among training-free LLM-based embeddings, outperforming the high-cost MetaEOL by a substantial margin. Finally, we highlight the potential of obtaining strong LLM embedding models through fine-tuning Value Aggregation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2a6822fe-a775-423e-932b-4a897e34230cBuilds on15
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same CoinEnrique Queipo-de-Llano, Alvaro Arroyo, Federico Barbero, Xiaowen Dong et al.ICLR 2026 · 56 citations
- Label Words are Anchors: An Information Flow Perspective for Understanding In-Context LearningLean Wang, Lei Li, Damai Dai, Deli Chen et al.EMNLP 2023 · 26 citations
- MMTEB: Massive Multilingual Text Embedding BenchmarkKenneth C. Enevoldsen, Isaac Chung, Imene Kerboua, Márton Kardos et al.ICLR 2025 · 10 citations
- Meta-Task Prompting Elicits Embeddings from Large Language ModelsYibin Lei, Di Wu, Tianyi Zhou, Tao Shen et al.ACL 2024 · 6 citations
Related papers
- KV-Embedding: Training-free Text Embedding via Internal KV Re-routing in Decoder-only LLMsYixuan Tang, Yi YangACL 2026 · 1 citation
- Token Prepending: A Training-Free Approach for Eliciting Better Sentence Embeddings from LLMsYuchen Fu, Zifeng Cheng, Zhiwei Jiang, Zhonghui Wang et al.ACL 2025
- Your Mixture-of-Experts LLM Is Secretly an Embedding Model for FreeZiyue Li, Tianyi ZhouICLR 2025
- Ensemble Learning for Heterogeneous Large Language Models with Deep Parallel CollaborationYichong Huang, Xiaocheng Feng, Baohang Li, Yang Xiang et al.NeurIPS 2024 · 94 citations
- Towards Improved Sentence Representations using Token GraphsKrishna Sri Ipsit Mantri, Carola-Bibiane Schönlieb, Zorah Lähner, Moshe EliasofICLR 2026 · 1 citation
