Effective post-training embedding compression via temperature control in contrastive training
Georgiana Dinu, Corey D. Barrett, Yi Xiang, Miguel Romero Calvo, Anna Currey, Xing Niu
Abstract
Fixed-size learned representations (dense representations, or embeddings) are widely used in many machine learning applications across language, vision or speech modalities. This paper investigates the role of the temperature parameter in contrastive training for text embeddings. We shed light on the impact this parameter has on the intrinsic dimensionality of the embedding spaces obtained, and show that lower intrinsic dimensionality is further correlated with effective compression of embeddings. We still observe a trade-off between absolute performance and effective compression and we propose temperature aggregation methods which reduce embedding size by an order of magnitude with minimal impact on quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bce7e6db-8157-48a8-aa2d-0dbc3564550fCited by top-tier papers2
- Closing the Modality Gap Aligns Group-Wise SemanticsEleonora Grassucci, Giordano Cicchetti, Emanuele Frasca, Aurelio Uncini et al.ICLR 2026 · 5 citations
- Bridging Domain Expertise and Generalization for Performance EstimationShuxuan Li, Zhilin Zhao, Quyu Kong, Wei-Shi ZhengCVPR 2026 · 1 citation
Builds on13
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
- Matryoshka Representation LearningAditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford et al.NeurIPS 2022 · 364 citations
- Can contrastive learning avoid shortcut solutions?Joshua Robinson, Li Sun, Ke Yu, Kayhan Batmanghelich et al.NeurIPS 2021 · 185 citations
Related papers
- Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variablesYu Gui, Cong Ma, Zongming MaNeurIPS 2025 · 9 citations
- Repurposing Language Models into Embedding Models: Finding the Compute-Optimal RecipeAlbert Q. Jiang, Alicja Ziarko, Bartosz Piotrowski, Wenda Li et al.NeurIPS 2024 · 4 citations
- Length-Induced Embedding Collapse in PLM-based ModelsYuqi Zhou, Sunhao Dai, Zhanshuo Cao, Xiao Zhang et al.ACL 2025 · 8 citations
- Scaling Laws For Dense RetrievalYan Fang, Jingtao Zhan, Qingyao Ai, Jiaxin Mao et al.SIGIR 2024 · 26 citations
- Following the Autoregressive Nature of LLM Embeddings via Compression and AlignmentJingcheng Deng, Zhongtao Jiang, Liang Pang, Zihao Wei et al.EMNLP 2025
