Fine-Grained Verifiers: Preference Modeling as Next-token Prediction in Vision-Language Alignment
Chenhang Cui, An Zhang, Yiyang Zhou, Zhaorun Chen, Gelei Deng, Huaxiu Yao, Tat-Seng Chua
Abstract
The recent advancements in large language models (LLMs) and pre-trained vision models have accelerated the development of vision-language large models (VLLMs), enhancing the interaction between visual and linguistic modalities. Despite their notable success across various domains, VLLMs face challenges in modality alignment, which can lead to issues like hallucinations and unsafe content generation. Current alignment techniques often rely on coarse feedback and external datasets, limiting scalability and performance. In this paper, we propose FiSAO (Fine-Grained Self-Alignment Optimization), a novel self-alignment method that utilizes the model's own visual encoder as a fine-grained verifier to improve visionlanguage alignment without the need for additional data. By leveraging token-level feedback from the vision encoder, FiSAO significantly improves vision-language alignment, even surpassing traditional preference tuning methods that require additional data. Through both theoretical analysis and experimental validation, we demonstrate that FiSAO effectively addresses the misalignment problem in VLLMs, marking the first instance of token-level rewards being applied to such models. Our code is avaliable at https://github.com/gzcch/FISAO_ICLR .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 236efea9-8ea0-411f-94e2-180e3e0a834eCited by top-tier papers8
- Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMsYifan Shen, Yuanzhe Liu, Jingyuan Zhu, Xu Cao et al.NeurIPS 2025 · 41 citations
- MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language ModelsQiyan Zhao, Xiaofeng Zhang, Yiheng Li, Yun Xing et al.ACM MM 2025 · 11 citations
- From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical VisualizationHaonian Ji, Shi Qiu, Siyang Xin, Siwei Han et al.ICLR 2026 · 6 citations
- Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMsXudong Li, Mengdan Zhang, Peixian Chen, Xiawu Zheng et al.NeurIPS 2025 · 4 citations
- MJ-Video: Benchmarking and Rewarding Video Generation with Fine-Grained Video PreferenceHaibo Tong, Zhaoyang Wang, Zhaorun Chen, Haonian Ji et al.NeurIPS 2025 · 1 citation
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Calibrated Self-Rewarding Vision Language ModelsYiyang Zhou, Zhiyuan Fan, Dongjie Cheng, Sihan Yang et al.NeurIPS 2024 · 77 citations
- Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference OptimizationShuo Xing, Peiran Li, Yuping Wang, Ruizheng Bai et al.EMNLP 2025 · 2 citations
- Language Bias in LVLMs: From In-Depth Analysis to Simple and Effective MitigationYangneng Chen, Jing LiICML 2026
- From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language ModelsYuying Shang, Xinyi Zeng, Yutao Zhu, Xiao Yang et al.ACM MM 2025 · 5 citations
- HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language ModelsSongtao Jiang, Yan Zhang, Yeying Jin, Zhihang Tang et al.ACL 2025
