When is an Embedding Model More Promising than Another?
Maxime Darrin, Philippe Formont, Ismail Ben Ayed, Jackie CK Cheung, Pablo Piantanida
摘要
Embedders play a central role in machine learning, projecting any object into numerical representations that can, in turn, be leveraged to perform various downstream tasks. The evaluation of embedding models typically depends on domain-specific empirical approaches utilizing downstream tasks, primarily because of the lack of a standardized framework for comparison. However, acquiring adequately large and representative datasets for conducting these assessments is not always viable and can prove to be prohibitively expensive and time-consuming. In this paper, we present a unified approach to evaluate embedders. First, we establish theoretical foundations for comparing embedding models, drawing upon the concepts of sufficiency and informativeness. We then leverage these concepts to devise a tractable comparison criterion (information sufficiency), leading to a task-agnostic and self-supervised ranking procedure. We demonstrate experimentally that our approach aligns closely with the capability of embedding models to facilitate various downstream tasks in both natural language processing and molecular biology. This effectively offers practitioners a valuable tool for prioritizing model trials.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Statistical Deficiency for Task Inclusion EstimationLoïc Fosse, Frédéric Béchet, Benoît Favre, Géraldine Damnati 等ACL 2025
- Towards an Explainable Comparison and Alignment of Feature EmbeddingsMohammad Jalali, Bahar Dibaei Nia, Farzan FarniaICML 2025
它引用的顶会 Paper28
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen 等NeurIPS 2020 · 被引用 3,042 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Self-Supervised Graph Transformer on Large-Scale Molecular DataYu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie 等NeurIPS 2020 · 被引用 1,113 次
相关 Paper
- Just Rank: Rethinking Evaluation with Word and Sentence SimilaritiesBin Wang, C.-C. Jay Kuo, Haizhou LiACL 2022 · 被引用 33 次
- Learning Task-Agnostic Representations through Multi-Teacher DistillationPhilippe Formont, Maxime Darrin, Banafsheh Karimian, Eric Granger 等NeurIPS 2025 · 被引用 6 次
- Interpretable Text Embeddings and Text Similarity Explanation: A SurveyJuri Opitz, Lucas Möller, Andrianos Michail, Sebastian Padó 等EMNLP 2025 · 被引用 3 次
- Observatory: Characterizing Embeddings of Relational TablesTianji Cong, Madelon Hulsebos, Zhenjie Sun, Paul Groth 等VLDB 2024 · 被引用 19 次
- EmbedLLM: Learning Compact Representations of Large Language ModelsRichard Zhuang, Tianhao Wu, Zhaojin Wen, Andrew Li 等ICLR 2025
