Following the Autoregressive Nature of LLM Embeddings via Compression and Alignment
Jingcheng Deng, Zhongtao Jiang, Liang Pang, Zihao Wei, Liwei Chen, Kun Xu, Yang Song, Huawei Shen, Xueqi Cheng
Abstract
A new trend uses LLMs as dense text encoders via contrastive learning. However, since LLM embeddings predict the probability distribution of the next token, they are inherently generative and distributive, conflicting with contrastive learning, which requires embeddings to capture full-text semantics and align via cosine similarity. This discrepancy hinders the full utilization of LLMs' pre-training capabilities, resulting in inefficient learning. In response to this issue, we propose AutoRegEmbed, a new contrastive learning method built on embedding conditional probability distributions, which integrates two core tasks: information compression and conditional distribution alignment. The information compression task encodes text into the embedding space, ensuring that the embedding vectors capture global semantics. The conditional distribution alignment task focuses on aligning text embeddings with positive samples embeddings by leveraging the conditional distribution of embeddings while simultaneously reducing the likelihood of generating negative samples from text embeddings, thereby achieving embedding alignment and uniformity. Experimental results demonstrate that our method significantly outperforms traditional contrastive learning approaches and achieves performance comparable to state-ofthe-art models when using the same amount of data. Our code is available at https:// github.com/TrustedLLM/AutoRegEmbed
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 69fe20a2-957e-4e86-bceb-3b1db8a8b48fCited by top-tier papers3
- RLKD: Distilling LLMs' Reasoning via Reinforcement LearningShicheng Xu, Liang Pang, Yunchang Zhu, Jia Gu et al.AAAI 2026 · 2 citations
- Learning to Compress: Unlocking the Potential of Large Language Models for Text RepresentationYeqin Zhang, Yizheng Zhao, Chen Hu, Binxing Jiao et al.AAAI 2026 · 2 citations
- TermGPT: Multi-Level Contrastive Fine-Tuning for Terminology Adaptation in Legal and Financial DomainsYidan Sun, Mengying Zhu, Feiyue Chen, Yangyang Wu et al.AAAI 2026 · 1 citation
Builds on30
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
Related papers
- Learning to Look at the Other Side: A Semantic Probing Study of Word Embeddings in LLMs with Enabled Bidirectional AttentionZhaoxin Feng, Jianfei Ma, Emmanuele Chersoni, Xiaojing Zhao et al.ACL 2025
- Focusing Condition: Inference-Time Self-Contrastive Steering Elicits Better Conditional Text Embeddings in LLMsZifeng Cheng, Lingyun Qian, Zhiwei Jiang, Cong Wang et al.ACL 2026
- Compressing then Matching: An Efficient Pre-training Paradigm for Multimodal EmbeddingDa Li, Yuxiao Luo, Keping Bi, Jiafeng Guo et al.ACL 2026 · 3 citations
- DaRec: A Disentangled Alignment Framework for Large Language Model and Recommender SystemXihong Yang, Heming Jing, Zixing Zhang, Jindong Wang et al.ICDE 2025 · 2 citations
- SemCSE: Semantic Contrastive Sentence Embeddings Using LLM-Generated Summaries For Scientific AbstractsMarc Felix Brinner, Sina ZarrießEMNLP 2025
