Improving Personalized Search with Regularized Low-Rank Parameter Updates
Fiona Ryan, Josef Sivic, Fabian Caba Heilbron, Judy Hoffman, James M. Rehg, Bryan C. Russell
Abstract
Personalized vision-language retrieval seeks to recognize new concepts (e.g., "my dog Fido") from only a few examples. This task is challenging because it requires not only learning a new concept from a few images, but also integrating the personal and general knowledge together to recognize the concept in different contexts. In this paper, we show how to effectively adapt the internal representation of a vision-language dual encoder model for personalized vision-language retrieval. We find that regularized low-rank adaption of a small set of parameters in the language encoder's final layer serves as a highly effective alternative to textual inversion for recognizing the personal concept while preserving general knowledge. Additionally, we explore strategies for combining parameters of multiple learned personal concepts, finding that parameter addition is effective. To evaluate how well general knowledge is preserved in a finetuned representation, we introduce a metric that measures image retrieval accuracy based on captions generated by a vision language model (VLM). Our approach achieves state-of-the-art accuracy on two benchmarks for personalized image retrieval with natural language queries -DeepFashion2 and ConCon-Chi -outperforming the prior art by 4% -22% on personal retrievals.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7315cf97-caff-47b1-91ab-e510a8ed033eCited by top-tier papers2
- CLAY: Conditional Visual Similarity Modulation in Vision-Language Embedding SpaceSohwi Lim, Lee Hyoseok, Jungjoon Park, Tae-Hyun OhCVPR 2026 · 2 citations
- Personalized Image Descriptions from Attention SequencesRuoyu Xue, Hieu Le, Jingyi Xu, Sounak Mondal et al.CVPR 2026 · 2 citations
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- Multimodal Few-Shot Learning with Frozen Language ModelsMaria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami et al.NeurIPS 2021 · 1,020 citations
- Prompt-aligned Gradient for Prompt TuningBeier Zhu, Yulei Niu, Yucheng Han, Yue Wu et al.ICCV 2023 · 475 citations
Related papers
- ConCon-Chi: Concept-Context Chimera Benchmark for Personalized Vision-Language TasksAndrea Rosasco, Stefano Berti, Giulia Pasquale, Damiano Malafronte et al.CVPR 2024 · 1 citation
- Meta-Personalizing Vision-Language Models to Find Named Instances in VideoChun-Hsiao Yeh, Bryan C. Russell, Josef Sivic, Fabian Caba Heilbron et al.CVPR 2023
- Stop Mixing Things Up! BISCUIT Teaches Vision-Language Models to Learn New Concepts from Images on the SpotJiahua Bao, Siyao Cheng, Jiaxing Du, Yuhang Jia et al.AAAI 2026
- Ego: Embedding-Guided Personalization of Vision-Language ModelsSoroush Seifi, Simon Gardier, Vaggelis Dorovatas, Daniel Olmeda Reino et al.CVPR 2026
- Seeing Through Words: Controlling Visual Retrieval Quality with Language ModelsJianglin Lu, Simon Jenni, Kushal Kafle, Jing Shi et al.ICLR 2026 · 3 citations
