RAP: Retrieval-Augmented Personalization for Multimodal Large Language Models
Haoran Hao, Jiaming Han, Changsheng Li, Yu-Feng Li, Xiangyu Yue
摘要
concept editing via updating the external database. To further improve generation quality and alignment with userspecific information, we design a pipeline for data collection and create a specialized dataset for personalized training of MLLMs. Based on the dataset, we train a series of MLLMs as personalized multimodal assistants. By pretraining on large-scale dataset, RAP-MLLMs can generalize to infinite visual concepts without additional finetuning. Our models demonstrate outstanding flexibility and generation quality across a variety of tasks, such as personalized image captioning, question answering and visual recognition. The code, data and models are available at https://hoar012.github.io/RAP-Project/ . Number of Image Data Requirements for Personalization Support Method Positive Negative Caption Description Question-Answer Recognition Real-time edit Text-only QA Fine-tuning n -Yes Yes No No ✗ ✓ MyVLM [2] n 150 Yes No Yes Yes ✗ ✗ Yo'LLaVA [32] n 200 No No Yes Yes ✗ ✓ RAP(Ours) 1 -No Yes No No ✓ ✓ through vision-language alignment brings powerful multimodal LLMs (MLLMs) [12, 15, 29, 33, 45, 51, 56] . MLLMs have shown significant improvement in various tasks, such as image description and question answering, highlighting their potential as humans' assistants. However, their lack of user-specific knowledge continues to limit their effectiveness as personalized assistants in daily life.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- PersonaVLM: Long-Term Personalized Multimodal LLMsChang Nie, Chaoyou Fu, Yifan Zhang, Haihua Yang 等CVPR 2026 · 被引用 11 次
- Unified Personalized Understanding, Generating and EditingYu Zhong, Tianwei Lin, Ruike Zhu, Yuqian Yuan 等CVPR 2026 · 被引用 7 次
- Personalized Image Descriptions from Attention SequencesRuoyu Xue, Hieu Le, Jingyi Xu, Sounak Mondal 等CVPR 2026 · 被引用 2 次
- Contextualized Visual Personalization in Vision-Language ModelsYeongtak Oh, Sangwon Yu, Junsung Park, Han Cheol Moon 等ICML 2026 · 被引用 1 次
- TAMEing Long Contexts in Personalization: Towards Training-Free and State-Aware MLLM Personalized AssistantRongpei Hong, Jian Lang, Ting Zhong, Yong Wang 等KDD 2026
它引用的顶会 Paper29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil 等ICLR 2024 · 被引用 1,798 次
相关 Paper
- A Comprehensive Overhaul of Multimodal Assistant with Small Language ModelsMinjie Zhu, Yichen Zhu, Ning Liu, Xin Liu 等AAAI 2025 · 被引用 30 次
- RePIC: Reinforced Post-Training for Personalizing Multi-Modal Language ModelsYeongtak Oh, Dohyun Chung, Juhyeon Shin, Sangha Park 等NeurIPS 2025 · 被引用 12 次
- Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth FusionJiuhai Chen, Jianwei Yang, Haiping Wu, Dianqi Li 等CVPR 2025
- Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level AlignmentPritam Sarkar, Sayna Ebrahimi, Ali Etemad, Ahmad Beirami 等ICLR 2025
- Training-Free Personalization via Retrieval and Reasoning on FingerprintsDeepayan Das, Davide Talon, Yiming Wang, Massimiliano Mancini 等ICCV 2025 · 被引用 1 次
