Conan-Embedding-v2: Training an LLM from Scratch for Text Embeddings
Shiyu Li, Yang Tang, Ruijie Liu, Shi-Zhe Chen, Xi Chen
Abstract
Large language models (LLMs) have recently demonstrated excellent performance in text embedding tasks. Previous work usually use LoRA to fine-tune existing LLMs, which are limited by the data and training gap between LLMs and embedding models. In this work, we introduce Conan-embedding-v2, a new 1.4Bparameter LLM trained from scratch and finetuned as a text embedder. First, we add news data and multilingual pairs for LLM pretraining to bridge the data gap. Based on this, we propose a cross-lingual retrieval dataset that enables the LLM to better integrate embeddings across different languages. Second, whereas LLMs use a causal mask with token-level loss, embedding models use a bidirectional mask with sentence-level loss. This training gap makes full fine-tuning less effective than LoRA. We introduce a soft-masking mechanism to gradually transition between these two types of masks, enabling the model to learn more comprehensive representations. Based on this, we propose a dynamic hard negative mining method that exposes the model to more difficult negative examples throughout the training process. Being intuitive and effective, with only approximately 1.4B parameters, Conanembedding-v2 achieves SOTA performance on both the Massive Text Embedding Benchmark (MTEB) and Chinese MTEB (May 19, 2025).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5236b0fd-1ffd-4aff-8729-c1f8fd188617Cited by top-tier papers2
- ReSeek: A Self-Correcting Framework for Search Agents with Instructive RewardsShiyu Li, Yifan Wang, Peiming Li, Zheng Wei et al.ICML 2026 · 8 citations
- Beyond Surface Features: Advancing Medical Vision-Language Alignment via Dynamic Evidence-Guided Preference OptimizationZixuan Huang, Zhihong Zhu, Xiaolong Liu, Yanchao Hao et al.ACL 2026
Builds on7
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- Matryoshka Representation LearningAditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford et al.NeurIPS 2022 · 364 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
Related papers
- KaLM-Embedding-V2: Superior Training Techniques and Data Inspire A Versatile Embedding ModelXinping Zhao, Xinshuo Hu, Zifei Shan, Shouzheng Huang et al.ICLR 2026 · 47 citations
- Improving Text Embeddings with Large Language ModelsLiang Wang, Nan Yang, Xiaolong Huang, Linjun Yang et al.ACL 2024
- Training LLMs to be Better Text Embedders through Bidirectional ReconstructionChang Su, Dengliang Shi, Siyuan Huang, Jintao Du et al.EMNLP 2025
- MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal EmbeddingsHaonan Chen, Hong Liu, Yuping Luo, Liang Wang et al.ACL 2026 · 20 citations
- Negative Matters: Multi-Granularity Hard-Negative Synthesis and Anchor-Token-Aware Pooling for Enhanced Text EmbeddingsTengyu Pan, Zhichao Duan, Zhenyu Li, Bowen Dong et al.ACL 2025 · 3 citations
