IISAN: Efficiently Adapting Multimodal Representation for Sequential Recommendation with Decoupled PEFT
Junchen Fu, Xuri Ge, Xin Xin, Alexandros Karatzoglou, Ioannis Arapakis, Jie Wang, Joemon M. Jose
Abstract
Multimodal foundation models are transformative in sequential recommender systems, leveraging powerful representation learning capabilities. While Parameter-efficient Fine-tuning (PEFT) is commonly used to adapt foundation models for recommendation tasks, most research prioritizes parameter efficiency, often overlooking critical factors like GPU memory efficiency and training speed. Addressing this gap, our paper introduces IISAN (Intra-and Inter-modal Side Adapted Network for Multimodal Representation) 1 , a simple plug-and-play architecture using a Decoupled PEFT structure and exploiting both intra-and inter-modal adaptation.
IISAN matches the performance of full fine-tuning (FFT) and state-of-the-art PEFT. More importantly, it significantly reduces GPU memory usage -from 47GB to just 3GB for multimodal sequential recommendation tasks. Additionally, it accelerates training time per epoch from 443s to 22s compared to FFT. This is also a notable improvement over the Adapter and LoRA, which require 37-39 GB GPU memory and 350-380 seconds per epoch for training.
Furthermore, we propose a new composite efficiency metric, TPME (Training-time, Parameter, and GPU Memory Efficiency) to alleviate the prevalent misconception that "parameter efficiency represents overall efficiency". TPME provides more comprehensive insights into practical efficiency comparisons between different methods. Besides, we give an accessible efficiency analysis of all PEFT and FFT approaches, which demonstrate the superiority of IISAN. Code is available at https://github.com/GAIR-Lab/IISAN.
1 Same pronunciation as the name of Thailand's largest region "Isan".
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5fc795fb-3bef-4f26-a930-d362b58e6bfdCited by top-tier papers16
- Hierarchical Time-Aware Mixture of Experts for Multi-Modal Sequential RecommendationShengzhe Zhang, Liyi Chen, Dazhong Shen, Chao Wang et al.WWW 2025 · 29 citations
- Multi-Modal Multi-Behavior Sequential Recommendation with Conditional Diffusion-Based Feature DenoisingXiaoxi Cui, Weihai Lu, Yu Tong, Yiheng Li et al.SIGIR 2025 · 21 citations
- When Large Vision Language Models Meet Multimodal Sequential Recommendation: An Empirical StudyPeilin Zhou, Chao Liu, Jing Ren, Xinfeng Zhou et al.WWW 2025 · 21 citations
- Order-agnostic Identifier for Large Language Model-based Generative RecommendationXinyu Lin, Haihan Shi, Wenjie Wang, Fuli Feng et al.SIGIR 2025 · 15 citations
- Towards End-to-End Explainable Facial Action Unit Recognition via Vision-Language Joint LearningXuri Ge, Junchen Fu, Fuhai Chen, Shan An et al.ACM MM 2024 · 12 citations
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
Related papers
- Personalized Parameter-Efficient Fine-Tuning of Foundation Models for Multimodal RecommendationSunwoo Kim, Hyunjin Hwang, Kijung ShinWWW 2026 · 1 citation
- Customizing Language Models with Instance-wise LoRA for Sequential RecommendationXiaoyu Kong, Jiancan Wu, An Zhang, Leheng Sheng et al.NeurIPS 2024 · 66 citations
- CoLA: Cross-Modal Low-rank Adaptation for Multimodal Downstream TasksWish Suharitdamrong, Tony Alex, Muhammad Awais, Sara AtitoICML 2026
- Parameter-Free Fine-tuning via Redundancy Elimination for Vision Foundation ModelsJiahuan Long, Tingsong Jiang, Wen Yao, Yizhe Xiong et al.AAAI 2026
- Faster Parameter-Efficient Tuning with Token Redundancy ReductionKwonyoung Kim, Jungin Park, Jin Kim, Hyeongjun Kwon et al.CVPR 2025
