Structured Policy Optimization: Enhance Large Vision-Language Model via Self-Referenced Dialogue
Guohao Sun, Can Qin, Yihao Feng, Zeyuan Chen, Ran Xu, Sohail A. Dianat, Majid Rabbani, Raghuveer Rao, Zhiqiang Tao
Abstract
Preference optimization algorithms typically enhance LLM response quality by leveraging human feedback on multiple answers given a fixed instruction. However, these methods often lack capturing the dynamic nature of conversational exchanges. For large vision-language models (LVLMs), direct preference optimization (DPO) can over-emphasize linguistic nuances while overlooking visual context. To address this challenge, we introduce structured policy optimization (SPO) -a novel preference optimization method that simultaneously aligns preference instructions, responses, and dialogue interactions to improve multi-modal understanding and reasoning capabilities. The efficacy of SPO is attributed to one key design: treating the questioning and answering as a sequential action and binding them through a trajectory reward. This reward formulation better aligns with real-world dialogue studies and eliminates the need for fixed instructions. We evaluate our models on interleaved benchmarks, including image, multiimage, and video-based understanding and reasoning tasks. Experimental results show that the proposed SPO finetuning LVLM with multi-modal preference data can align with human preference more efficiently than DPO. The code is available at https://github.com/heliossun/ Structure-Policy-Optimization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d542965-b54f-4bcd-99fe-c1c7ce540a40Cited by top-tier papers1
Ask how each one uses itBuilds on19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- TPO: Aligning Large Language Models with Multi-branch & Multi-step Preference TreesWeibin Liao, Xu Chu, Yasha WangICLR 2025
- mDPO: Conditional Preference Optimization for Multimodal Large Language ModelsFei Wang, Wenxuan Zhou, James Y. Huang, Nan Xu et al.EMNLP 2024 · 11 citations
- Multi-step Visual Reasoning with Visual Tokens Scaling and VerificationTianyi Bai, Zengjie Hu, Fupeng Sun, Jiantao Qiu et al.NeurIPS 2025 · 22 citations
- Structured Preference Optimization for Vision-Language Long-Horizon Task PlanningXiwen Liang, Min Lin, Weiqi Ruan, Rongtao Xu et al.EMNLP 2025
- Spotlight on Token Perception for Multimodal Reinforcement LearningSiyuan Huang, Xiaoye Qu, Yafu Li, Yun Luo et al.ICLR 2026 · 45 citations
