Lune

CVPR2025顶会

RAP: Retrieval-Augmented Personalization for Multimodal Large Language Models

Haoran Hao, Jiaming Han, Changsheng Li, Yu-Feng Li, Xiangyu Yue

2025年份
7顶会引用

摘要

concept editing via updating the external database. To further improve generation quality and alignment with userspecific information, we design a pipeline for data collection and create a specialized dataset for personalized training of MLLMs. Based on the dataset, we train a series of MLLMs as personalized multimodal assistants. By pretraining on large-scale dataset, RAP-MLLMs can generalize to infinite visual concepts without additional finetuning. Our models demonstrate outstanding flexibility and generation quality across a variety of tasks, such as personalized image captioning, question answering and visual recognition. The code, data and models are available at https://hoar012.github.io/RAP-Project/ . Number of Image Data Requirements for Personalization Support Method Positive Negative Caption Description Question-Answer Recognition Real-time edit Text-only QA Fine-tuning n -Yes Yes No No ✗ ✓ MyVLM [2] n 150 Yes No Yes Yes ✗ ✗ Yo'LLaVA [32] n 200 No No Yes Yes ✗ ✓ RAP(Ours) 1 -No Yes No No ✓ ✓ through vision-language alignment brings powerful multimodal LLMs (MLLMs) [12, 15, 29, 33, 45, 51, 56] . MLLMs have shown significant improvement in various tasks, such as image description and question answering, highlighting their potential as humans' assistants. However, their lack of user-specific knowledge continues to limit their effectiveness as personalized assistants in daily life.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper7

问问它们各自怎么用它

它引用的顶会 Paper29

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖