ESE: Espresso Sentence Embeddings
Xianming Li, Zongxi Li, Jing Li, Haoran Xie, Qing Li
Abstract
High-quality sentence embeddings are fundamental in many natural language processing (NLP) tasks, such as semantic textual similarity (STS) and retrievalaugmented generation (RAG). However, most existing methods leverage fixedlength sentence embeddings from full-layer language models, which lack the scalability to accommodate the diverse available resources across various applications. Viewing this gap, we propose a novel sentence embedding model Espresso Sentence Embeddings (ESE) with two learning processes. First, the learn-to-express process encodes more salient representations to shallow layers. Second, the learn-to-compress process compacts essential features into the initial dimensions using Principal Component Analysis (PCA). This way, ESE can scale model depth via the former process and embedding size via the latter. Extensive experiments on STS and RAG suggest that ESE can effectively produce high-quality sentence embeddings with less model depth and embedding size, enhancing embedding inference efficiency. The code is available at https://github.com/SeanLee97/AnglE/blob/main/README_ESE.md .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7626bb6b-ccc7-40b2-b844-704d36da371bCited by top-tier papers6
- Meta-Task Prompting Elicits Embeddings from Large Language ModelsYibin Lei, Di Wu, Tianyi Zhou, Tao Shen et al.ACL 2024 · 6 citations
- Making Large Language Models Efficient Dense RetrieversYibin Lei, Shwai He, Ang Li, Andrew YatesACL 2026 · 2 citations
- TALAS: Teacher-Anchored Layer Alignment with Adaptive Sharpness-Aware Minimization for Embedding DistillationQuoc Phong Dao, Hoang Son Nguyen, Pham Khanh Chi, Linh Ngo Van et al.ACL 2026
- ML-Embed: Inclusive and Efficient Embeddings for a Multilingual WorldZiyin Zhang, Zihan Liao, Hang Yu, Peng Di et al.ICML 2026
- CondAmbigQA: A Benchmark and Dataset for Conditional Ambiguous Question AnsweringZongxi Li, Yang Li, Haoran Xie, S. Joe QinEMNLP 2025
Builds on8
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Matryoshka Representation LearningAditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford et al.NeurIPS 2022 · 364 citations
- Semantic Re-tuning with Contrastive TensionFredrik Carlsson, Amaru Cuba Gyllensten, Evangelia Gogoulou, Erik Ylipää Hellqvist et al.ICLR 2021 · 86 citations
- AoE: Angle-optimized Embeddings for Semantic Textual SimilarityXianming Li, Jing LiACL 2024 · 22 citations
Related papers
- Static Word Embeddings for Sentence Semantic RepresentationTakashi Wada, Yuki Hirakawa, Ryotaro Shimizu, Takahiro Kawashima et al.EMNLP 2025 · 1 citation
- An Unsupervised Sentence Embedding Method by Mutual Information MaximizationYan Zhang, Ruidan He, Zuozhu Liu, Kwan Hui Lim et al.EMNLP 2020 · 126 citations
- A Contrastive Framework for Learning Sentence Representations from Pairwise and Triple-wise Perspective in Angular SpaceYuhao Zhang, Hongji Zhu, Yongliang Wang, Nan Xu et al.ACL 2022 · 94 citations
- 3R: Enhancing Sentence Representation Learning via Redundant Representation ReductionLongxuan Ma, Xiao Wu, Yuxin Huang, Shengxiang Gao et al.EMNLP 2025
- Generate, Discriminate and Contrast: A Semi-Supervised Sentence Representation Learning FrameworkYiming Chen, Yan Zhang, Bin Wang, Zuozhu Liu et al.EMNLP 2022 · 9 citations
