ICLR2025
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
Shuo Li, Tao Ji, Xiaoran Fan, Linsheng Lu, Leyi Yang, Yuming Yang, Zhiheng Xi, Rui Zheng, Yuran Wang, Xiaohui Zhao, Tao Gui, Qi Zhang, Xuanjing Huang
Abstract
Sycophancy, a common hallucination issue in large language models (LLMs), leads them to blindly agree with users, even when users' opinions are harmful. As LLMs expand into other modalities like vision-language models (VLMs), the saying "seeing is believing" raises the question: do VLMs still exhibit sycophancy when given images as evidence? This paper presents the first sycophancy evaluation benchmark for VLMs, named MM-SY, which covers ten diverse visual understanding tasks. We reveal that VLMs still sycophantically agree with users while ignoring visual facts, influenced by various factors like different tasks, user tones, model sizes, etc. To mitigate it, inspired by methods for reducing hallucination in LLMs, we investigate three methods: prompt-based, supervised fine-tuning, and direct preference optimization. We find that their ability to reduce sycophancy improves progressively. However, this mitigation has made the VLM more stubborn and less receptive to corrections. To balance the trade-off, we analyze the causes of sycophancy and explore a simple training-free approach, with experiments validating its effectiveness. 1
