ESE: Espresso Sentence Embeddings
Xianming Li, Zongxi Li, Jing Li, Haoran Xie, Qing Li
摘要
High-quality sentence embeddings are fundamental in many natural language processing (NLP) tasks, such as semantic textual similarity (STS) and retrievalaugmented generation (RAG). However, most existing methods leverage fixedlength sentence embeddings from full-layer language models, which lack the scalability to accommodate the diverse available resources across various applications. Viewing this gap, we propose a novel sentence embedding model Espresso Sentence Embeddings (ESE) with two learning processes. First, the learn-to-express process encodes more salient representations to shallow layers. Second, the learn-to-compress process compacts essential features into the initial dimensions using Principal Component Analysis (PCA). This way, ESE can scale model depth via the former process and embedding size via the latter. Extensive experiments on STS and RAG suggest that ESE can effectively produce high-quality sentence embeddings with less model depth and embedding size, enhancing embedding inference efficiency. The code is available at https://github.com/SeanLee97/AnglE/blob/main/README_ESE.md .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Meta-Task Prompting Elicits Embeddings from Large Language ModelsYibin Lei, Di Wu, Tianyi Zhou, Tao Shen 等ACL 2024 · 被引用 6 次
- Making Large Language Models Efficient Dense RetrieversYibin Lei, Shwai He, Ang Li, Andrew YatesACL 2026 · 被引用 2 次
- TALAS: Teacher-Anchored Layer Alignment with Adaptive Sharpness-Aware Minimization for Embedding DistillationQuoc Phong Dao, Hoang Son Nguyen, Pham Khanh Chi, Linh Ngo Van 等ACL 2026
- ML-Embed: Inclusive and Efficient Embeddings for a Multilingual WorldZiyin Zhang, Zihan Liao, Hang Yu, Peng Di 等ICML 2026
- CondAmbigQA: A Benchmark and Dataset for Conditional Ambiguous Question AnsweringZongxi Li, Yang Li, Haoran Xie, S. Joe QinEMNLP 2025
它引用的顶会 Paper8
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Matryoshka Representation LearningAditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford 等NeurIPS 2022 · 被引用 364 次
- Semantic Re-tuning with Contrastive TensionFredrik Carlsson, Amaru Cuba Gyllensten, Evangelia Gogoulou, Erik Ylipää Hellqvist 等ICLR 2021 · 被引用 86 次
- AoE: Angle-optimized Embeddings for Semantic Textual SimilarityXianming Li, Jing LiACL 2024 · 被引用 22 次
相关 Paper
- Static Word Embeddings for Sentence Semantic RepresentationTakashi Wada, Yuki Hirakawa, Ryotaro Shimizu, Takahiro Kawashima 等EMNLP 2025 · 被引用 1 次
- An Unsupervised Sentence Embedding Method by Mutual Information MaximizationYan Zhang, Ruidan He, Zuozhu Liu, Kwan Hui Lim 等EMNLP 2020 · 被引用 126 次
- A Contrastive Framework for Learning Sentence Representations from Pairwise and Triple-wise Perspective in Angular SpaceYuhao Zhang, Hongji Zhu, Yongliang Wang, Nan Xu 等ACL 2022 · 被引用 94 次
- 3R: Enhancing Sentence Representation Learning via Redundant Representation ReductionLongxuan Ma, Xiao Wu, Yuxin Huang, Shengxiang Gao 等EMNLP 2025
- Generate, Discriminate and Contrast: A Semi-Supervised Sentence Representation Learning FrameworkYiming Chen, Yan Zhang, Bin Wang, Zuozhu Liu 等EMNLP 2022 · 被引用 9 次
