Personalized Visual Instruction Tuning
Renjie Pi, Jianshu Zhang, Tianyang Han, Jipeng Zhang, Rui Pan, Tong Zhang
摘要
Recent advancements in multimodal large language models (MLLMs) have demonstrated significant progress; however, these models exhibit a notable limitation, which we refer to as "face blindness". Specifically, they can engage in general conversations but fail to conduct personalized dialogues targeting at specific individuals. This deficiency hinders the application of MLLMs in personalized settings, such as tailored visual assistants on mobile devices, or domestic robots that need to recognize members of the family. In this paper, we introduce Personalized Visual Instruction Tuning (PVIT), a novel data curation and training framework designed to enable MLLMs to identify target individuals within an image and engage in personalized and coherent dialogues. Our approach involves the development of a sophisticated pipeline that autonomously generates training data containing personalized conversations. This pipeline leverages the capabilities of various visual experts, image generation models, and (multi-modal) large language models. To evaluate the personalized potential of MLLMs, we present a benchmark called P-Bench, which encompasses various question types with different levels of difficulty. The experiments demonstrate a substantial personalized performance enhancement after fine-tuning with our curated dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept TokensRuichuan An, Sihan Yang, Renrui Zhang, Zijun Shen 等NeurIPS 2025 · 被引用 61 次
- Personalized Generation In Large Model Era: A SurveyYiyan Xu, Jinghao Zhang, Alireza Salemi, Xinting Hu 等ACL 2025 · 被引用 45 次
- Backdoor Cleaning without External Guidance in MLLM Fine-tuningXuankun Rong, Wenke Huang, Jian Liang, Jinhe Bi 等NeurIPS 2025 · 被引用 39 次
- RePIC: Reinforced Post-Training for Personalizing Multi-Modal Language ModelsYeongtak Oh, Dohyun Chung, Juhyeon Shin, Sangha Park 等NeurIPS 2025 · 被引用 12 次
- PersonaVLM: Long-Term Personalized Multimodal LLMsChang Nie, Chaoyou Fu, Yifan Zhang, Haihua Yang 等CVPR 2026 · 被引用 11 次
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
相关 Paper
- Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction TuningJi Hyeok Jung, Eun Tae Kim, Seo Yeon Kim, Joo Ho Lee 等CVPR 2025
- Learning to Instruct for Visual Instruction TuningZhihan Zhou, Feng Hong, Jiaan Luo, Yushi Ye 等NeurIPS 2025 · 被引用 8 次
- Corrupted but Not Broken: Understanding and Mitigating the Negative Impacts of Corrupted Data in Visual Instruction TuningYunhao Gou, Hansi Yang, Zhili Liu, Kai Chen 等EMNLP 2025
- PGT: Procedurally Generated Tasks for improving visual grounding in MLLMsRim Assouel, Amir Bar, Michal Drozdzal, Adriana Romero-SorianoICML 2026
- VIGC: Visual Instruction Generation and CorrectionBin Wang, Fan Wu, Xiao Han, Jiahui Peng 等AAAI 2024 · 被引用 95 次
