KaLM-Embedding-V2: Superior Training Techniques and Data Inspire A Versatile Embedding Model
Xinping Zhao, Xinshuo Hu, Zifei Shan, Shouzheng Huang, Yao Zhou, Xin Zhang, Zetian Sun, Zhenyu Liu, Dongfang Li, Xinyuan Wei, Youcheng Pan, Yang Xiang
摘要
Recent advancements in Large Language Models (LLMs)-based text embedding models primarily focus on data scaling or synthesis, yet limited exploration of training techniques and data quality, thereby constraining performance. In this work, we propose KaLM-Embedding-V2 from the Lychee-KaLM team, a series of versatile and compact embedding models, systematically incentivizing advanced embedding capability in LLMs by superior training techniques and high-quality data. For model architecture, we implement the models on a 0.5B compact size with simple mean-pooling to produce fixed-length embeddings and remove the causal attention mask to enable fully bidirectional representation learning. For training techniques, we propose a progressive multi-stage training pipeline: pre-training on weakly supervised large-scale datasets, fine-tuning with supervised high-quality datasets, and contrastive distillation with fine-grained soft signals, integrated with focal-style reweighting and online hard-negative mixing to emphasize difficult samples and enrich hard negatives, respectively. For training data, we curate over 20 categories for pre-training and 100 categories for fine-tuning and contrastive distillation, to improve both performance and generalization, leveraging task-specific instructions, hard-negative mining, and example-based multi-class labeling to ensure high quality. Combining these techniques, our KaLM-Embedding-V2 series achieves state-of-the-art performance on the Massive Text Embedding Benchmark, outperforming models of comparable size and rivaling models 3-26x larger, setting a new standard for versatile and compact embedding models under 1B parameters. The code, data, and models are available at https://kalm-embedding.github.io/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Balancing the Blend: An Experimental Analysis of Trade-offs in Hybrid SearchMengzhao Wang, Boyu Tan, Yunjun Gao, Hai Jin 等VLDB 2026 · 被引用 8 次
- Structured Episodic Event MemoryZhengxuan Lu, Dongfang Li, Yukun Shi, Beilun Wang 等ACL 2026 · 被引用 1 次
- CONE: Embeddings for Complex Numerical Data Preserving Unit and Variable SemanticsGyanendra Shrestha, Anna Pyayt, Michael N. GubanovSIGMOD 2026
- No More K-means: Single-Stage Sparse Coding for Efficient Multi-Vector RetrievalLixuan Guo, Yifei Wang, Tiansheng Wen, Aosong Feng 等ICML 2026
- ML-Embed: Inclusive and Efficient Embeddings for a Multilingual WorldZiyin Zhang, Zihan Liao, Hang Yu, Peng Di 等ICML 2026
它引用的顶会 Paper16
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Matryoshka Representation LearningAditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford 等NeurIPS 2022 · 被引用 364 次
- Crosslingual Generalization through Multitask FinetuningNiklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts 等ACL 2023 · 被引用 319 次
- Large Dual Encoders Are Generalizable RetrieversJianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai 等EMNLP 2022 · 被引用 145 次
- Large Language Models for Data Annotation and Synthesis: A SurveyZhen Tan, Dawei Li, Song Wang, Alimohammad Beigi 等EMNLP 2024 · 被引用 119 次
相关 Paper
- Conan-Embedding-v2: Training an LLM from Scratch for Text EmbeddingsShiyu Li, Yang Tang, Ruijie Liu, Shi-Zhe Chen 等EMNLP 2025
- Improving Text Embeddings with Large Language ModelsLiang Wang, Nan Yang, Xiaolong Huang, Linjun Yang 等ACL 2024
- Training LLMs to be Better Text Embedders through Bidirectional ReconstructionChang Su, Dengliang Shi, Siyuan Huang, Jintao Du 等EMNLP 2025
- NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding ModelsChankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman 等ICLR 2025
- Negative Matters: Multi-Granularity Hard-Negative Synthesis and Anchor-Token-Aware Pooling for Enhanced Text EmbeddingsTengyu Pan, Zhichao Duan, Zhenyu Li, Bowen Dong 等ACL 2025 · 被引用 3 次
