Personalized Parameter-Efficient Fine-Tuning of Foundation Models for Multimodal Recommendation
Sunwoo Kim, Hyunjin Hwang, Kijung Shin
Abstract
In recent years, substantial research has integrated multimodal item metadata into recommender systems, often by using pre-trained multimodal foundation models to encode such data. Since these models are not originally trained for recommendation tasks, recent works efficiently adapt them via parameter-efficient fine-tuning (PEFT). However, even with PEFT, item embeddings from multimodal foundation models remain user-blind: item embeddings are not conditioned on user interests, despite the fact that users with diverse interests attend to different item aspects. To address this limitation, we propose PerPEFT, a personalized PEFT strategy for multimodal recommendation. Specifically, PerPEFT groups users by interest and assigns a distinct PEFT module to each group, enabling each module to capture the fine-grained item aspects most predictive of that group's purchase decisions. We further introduce a specialized training technique that strengthens this user-group conditioning. Notably, PerPEFT is PEFT-agnostic and can be paired with any PEFT method applicable to multimodal foundation models. Through extensive experiments, we show that (1) PerPEFT outperforms the strongest baseline by up to 15.3% (NDCG@20) and (2) delivers consistent gains across diverse PEFT variants. It is noteworthy that, even with personalization, PEFT remains lightweight, adding only 1.3% of the parameter count of the foundation model. We provide our code and datasets at https://github.com/kswoo97/PerPEFT . CCS Concepts • Information systems → Recommender systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 47ebfcf2-b6e9-4957-a844-1194088152d0Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li et al.SIGIR 2020 · 4,448 citations
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context LearningHaokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta et al.NeurIPS 2022 · 1,483 citations
- Contrastive Learning with Hard Negative SamplesJoshua David Robinson, Ching-Yao Chuang, Suvrit Sra, Stefanie JegelkaICLR 2021 · 999 citations
Related papers
- IISAN: Efficiently Adapting Multimodal Representation for Sequential Recommendation with Decoupled PEFTJunchen Fu, Xuri Ge, Xin Xin, Alexandros Karatzoglou et al.SIGIR 2024 · 40 citations
- CoPL: Parameter-Efficient Collaborative Prompt Learning for Audio-Visual TasksYihan Zhao, Wei Xi, Yuhang Cui, Gairui Bai et al.ACM MM 2024 · 3 citations
- Parameter-Efficient Adaptation for MLLMs via Implicit Modality DecompositionMingfang Zhang, Yunhong Wang, Lu Wang, Jiaxin ChenCVPR 2026
- PEFT-BoA: Parameter-Efficient Fine-Tuning with Bag-of-Adapters for Multi-Modal Object Re-identificationHongchao Li, Guangxing Liu, Xixi Wang, Baihe Liang et al.AAAI 2026
- Personalized Pieces: Efficient Personalized Large Language Models through Collaborative EffortsZhaoxuan Tan, Zheyuan Liu, Meng JiangEMNLP 2024 · 11 citations
