Evaluating Transferability in Retrieval Tasks: An Approach Using MMD and Kernel Methods
Mengyu Dai, Amir Hossein Raffiee, Aashish Jain, Joshua Correa
摘要
Retrieval tasks play central roles in real-world machine learning systems such as search engine, recommender system, and retrieval-augmented generation (RAG). Achieving decent performance in these tasks often requires fine-tuning various pretrained models on specific datasets and selecting the best candidate, a process that can be both time and resource consuming. To tackle the problem, we introduce a novel and efficient method, called RetMMD, that leverages Maximum Mean Discrepancy (MMD) and kernel methods to assess the transferability of pretrained models in retrieval tasks. RetMMD is calculated on pretrained model and target dataset without any fine-tuning involved. Specifically, given some query, we quantify the distribution discrepancy between relevant and irrelevant document embeddings, by estimating the similarities within their mappings in the finetuned embedding space through kernel method. This discrepancy is averaged over multiple queries, taking into account the distribution characteristics of the target dataset. Experiments suggest that the proposed metric calculated on pretrained models closely aligns with retrieval performance post fine-tuning. The observation holds across a variety of datasets, including image, text, and multi-modal domains, indicating the potential of using MMD and kernel methods for transfer learning evaluation in retrieval scenarios. In addition, we also design a way of evaluating dataset transferability for retrieval tasks, with experimental results demonstrating the effectiveness of the proposed approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Unified Transferability Metrics for Time Series Foundation ModelsWeiyang Zhang, Xinyang Chen, Xiucheng Li, Kehai Chen 等NeurIPS 2025 · 被引用 5 次
- Real-Time Multimodal Fingertip Contact Detection via Depth and Motion Fusion for Vision-Based Human–Computer InteractionMukhiddin Toshpulatov, Wookey Lee, Suan Lee, Geehyuk LeeCVPR 2026
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
相关 Paper
- MMD Guidance: Training-Free Distribution Adaptation for Diffusion Models via Maximum Mean Discrepancy GuidanceMatina Mahdizadeh Sani, Nima Jamali, Mohammad Jalali, Farzan FarniaICML 2026 · 被引用 4 次
- Fast and Accurate Transferability Measurement by Evaluating Intra-class Feature VarianceHuiwen Xu, U KangICCV 2023 · 被引用 12 次
- Misalign, Contrast then Distill: Rethinking Misalignments in Language-Image PretrainingBumsoo Kim, Yeonsik Jo, Jinhyung Kim, Seung-Hwan KimICCV 2023 · 被引用 11 次
- Foundation Model is Efficient Multimodal Multitask Model SelectorFanqing Meng, Wenqi Shao, Zhanglin Peng, Chonghe Jiang 等NeurIPS 2023 · 被引用 26 次
- Why is a Bird's Caption a Good Demonstration? Towards Effective Multimodal In-Context Learning without Dedicated DataJunlin Fang, Wenya Wang, Lingli Zhang, Fengmao LvACM MM 2025 · 被引用 3 次
