PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning
Yibo Lyu, Rui Shao, Gongwei Chen, Yijie Zhu, Weili Guan, Liqiang Nie
摘要
As multimedia content expands, the demand for unified multimodal retrieval (UMR) in real-world applications increases. Recent work leverages multimodal large language models to tackle this task. However, the large number of parameters leads to high training resource demands and low inference efficiency. To address this issue, we propose the PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning, an efficient approach to enhancing the unified retrieval capabilities from both structure and learning perspectives: 1) From the perspective of model structure, to retain the most retrieval-relevant components within MLLMs, we analyze and propose Layer-Pruned Self-Distillation approach. It structurally prunes the model by preserving only the shallow layers, substantially reducing the parameters of MLLM. Moreover, we use self-distillation to mitigate the representational degradation caused by pruning. It reuses the feature from dropped deep layers as the teacher signal, where the supervised signal enables the retrieval embedding token to efficiently inherit effective representational capacity, resulting in a more compact model. 2) From the perspective of model learning, to mitigate representation degradation caused by rapid convergence during multimodal contrastive learning, we propose Modality-Adaptive Contrastive Learning Loss (MAC-Loss). It adaptively separates in-batch negative candidate samples into harder intramodality and simpler inter-modality groups based on each query's target modality. Assigning each group a temperature coefficient with different strategies enables each query to adaptively focus on challenging in-batch negatives, reducing the resource demands of multimodal contrastive learning. Experiments demonstrate that our approach achieves double efficiency, significantly reduces resource consumption while maintaining most of the performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic ManipulationWei Li, Renshan Zhang, Rui Shao, Zhijian Fang 等AAAI 2026 · 被引用 13 次
- HiconAgent: History Context-aware Policy Optimization for GUI AgentsXurui Zhou, Gongwei Chen, Yuquan Xie, Zaijing Li 等CVPR 2026 · 被引用 11 次
- ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic ManipulationWei Li, Jizhihui Liu, Yixing Li, Junwen Tong 等CVPR 2026 · 被引用 8 次
- H-GAR: A Hierarchical Interaction Framework via Goal-Driven Observation-Action Refinement for Robotic ManipulationYijie Zhu, Rui Shao, Ziyang Liu, Jie He 等AAAI 2026 · 被引用 5 次
- FreeRet: MLLMs as Training-Free RetrieversYuhan Zhu, Xiangyu Zeng, Chenting Wang, Xinhao Li 等ICML 2026 · 被引用 5 次
它引用的顶会 Paper37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- Unified Knowledge Maintenance Pruning and Progressive Recovery with Weight Recalling for Large Vision-Language ModelsZimeng Wu, Jiaxin Chen, Yunhong WangAAAI 2025 · 被引用 4 次
- Compressing then Matching: An Efficient Pre-training Paradigm for Multimodal EmbeddingDa Li, Yuxiao Luo, Keping Bi, Jiafeng Guo 等ACL 2026 · 被引用 3 次
- Retrv-MoE: Scaling Unified Multimodal Retrieval with Sparse Mixture-of-ExpertsTongxu Lin, Jiayin XiaoKDD 2026
- C3CMR: Cross-Modality Cross-Instance Contrastive Learning for Cross-Media RetrievalJunsheng Wang, Tiantian Gong, Zhixiong Zeng, Changchang Sun 等ACM MM 2022 · 被引用 12 次
- LLaVA-MoD: Making LLaVA Tiny via MoE-Knowledge DistillationFangxun Shu, Yue Liao, Lei Zhang, Le Zhuo 等ICLR 2025
