PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning
Yibo Lyu, Rui Shao, Gongwei Chen, Yijie Zhu, Weili Guan, Liqiang Nie
Abstract
As multimedia content expands, the demand for unified multimodal retrieval (UMR) in real-world applications increases. Recent work leverages multimodal large language models to tackle this task. However, the large number of parameters leads to high training resource demands and low inference efficiency. To address this issue, we propose the PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning, an efficient approach to enhancing the unified retrieval capabilities from both structure and learning perspectives: 1) From the perspective of model structure, to retain the most retrieval-relevant components within MLLMs, we analyze and propose Layer-Pruned Self-Distillation approach. It structurally prunes the model by preserving only the shallow layers, substantially reducing the parameters of MLLM. Moreover, we use self-distillation to mitigate the representational degradation caused by pruning. It reuses the feature from dropped deep layers as the teacher signal, where the supervised signal enables the retrieval embedding token to efficiently inherit effective representational capacity, resulting in a more compact model. 2) From the perspective of model learning, to mitigate representation degradation caused by rapid convergence during multimodal contrastive learning, we propose Modality-Adaptive Contrastive Learning Loss (MAC-Loss). It adaptively separates in-batch negative candidate samples into harder intramodality and simpler inter-modality groups based on each query's target modality. Assigning each group a temperature coefficient with different strategies enables each query to adaptively focus on challenging in-batch negatives, reducing the resource demands of multimodal contrastive learning. Experiments demonstrate that our approach achieves double efficiency, significantly reduces resource consumption while maintaining most of the performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 53399e8f-db28-4c2c-9830-449b5bbd643bCited by top-tier papers5
- SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic ManipulationWei Li, Renshan Zhang, Rui Shao, Zhijian Fang et al.AAAI 2026 · 13 citations
- HiconAgent: History Context-aware Policy Optimization for GUI AgentsXurui Zhou, Gongwei Chen, Yuquan Xie, Zaijing Li et al.CVPR 2026 · 11 citations
- ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic ManipulationWei Li, Jizhihui Liu, Yixing Li, Junwen Tong et al.CVPR 2026 · 8 citations
- H-GAR: A Hierarchical Interaction Framework via Goal-Driven Observation-Action Refinement for Robotic ManipulationYijie Zhu, Rui Shao, Ziyang Liu, Jie He et al.AAAI 2026 · 5 citations
- FreeRet: MLLMs as Training-Free RetrieversYuhan Zhu, Xiangyu Zeng, Chenting Wang, Xinhao Li et al.ICML 2026 · 5 citations
Builds on37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Unified Knowledge Maintenance Pruning and Progressive Recovery with Weight Recalling for Large Vision-Language ModelsZimeng Wu, Jiaxin Chen, Yunhong WangAAAI 2025 · 4 citations
- Compressing then Matching: An Efficient Pre-training Paradigm for Multimodal EmbeddingDa Li, Yuxiao Luo, Keping Bi, Jiafeng Guo et al.ACL 2026 · 3 citations
- Retrv-MoE: Scaling Unified Multimodal Retrieval with Sparse Mixture-of-ExpertsTongxu Lin, Jiayin XiaoKDD 2026
- C3CMR: Cross-Modality Cross-Instance Contrastive Learning for Cross-Media RetrievalJunsheng Wang, Tiantian Gong, Zhixiong Zeng, Changchang Sun et al.ACM MM 2022 · 12 citations
- LLaVA-MoD: Making LLaVA Tiny via MoE-Knowledge DistillationFangxun Shu, Yue Liao, Lei Zhang, Le Zhuo et al.ICLR 2025
