Dynamics-Aware Preference Optimization for Vision-Language Models
jusheng zhang, Kaitong Cai, Jing Yang, Jian Wang, Keze Wang
Abstract
Preference-based finetuning of vision-language models (VLMs) is notoriously unstable, i.e., trivially wrong negatives inject uninformative gradients that distort optimization and degrade calibration. This work revisits this issue through the lens of learning dynamics and identifies a core pathology, i.e., the squeezing effect, where easy negatives retain large, misaligned gradients despite negligible loss. To address this, we propose Cooling-Weighted Direct Preference Optimization (CW-DPO), a two-stage framework that smooths and then stabilizes the alignment process. Stage 1 employs a constrained SFT phase with low-weight "gentle negatives" to regularize overconfident distributions and flatten the loss landscape. Stage 2 introduces a competence-aware cooling weight that adaptively scales negative gradients according to the model's average per-token log-probability, suppressing uninformative updates while emphasizing hard, on-policy contrasts. This dynamics-aware weighting effectively mitigates the squeezing effect and enables smoother convergence. Extensive and comprehensive results on the mainstream benchmarks, i.e., COCO, Flickr30k, NoCaps, MMMU, and MMBench1.1, our CW-DPO achieves stateof-the-art performance, e.g., +3.4 CIDEr over PPO and +2.4% absolute accuracy on MMMU, while improving calibration and halving convergence steps. This justifies that smoothing before cooling constitutes a simple yet general principle for robust VLM preference optimization. https://github.com/jushengzhang/Dynamics-Aware-Preference-Optimization
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Sharpness-Aware Minimization in Logit Space Efficiently Enhances Direct Preference OptimizationHaocheng Luo, Zehang Deng, Thanh-Toan Do, Mehrtash Harandi et al.ICLR 2026 · 1 citation
- VPO: Reasoning Preferences Optimization Based on V-Usable InformationZecheng Wang, Chunshan Li, Yupeng Zhang, Han Liu et al.NeurIPS 2025 · 1 citation
- Learning Dynamics of LLM FinetuningYi Ren, Danica J. SutherlandICLR 2025
- Optimal Transport-Based Token Weighting scheme for Enhanced Preference OptimizationMeng Li, Guangda Huzhang, Haibo Zhang, Xiting Wang et al.ACL 2025
- β-DPO: Direct Preference Optimization with Dynamic βJunkang Wu, Yuexiang Xie, Zhengyi Yang, Jiancan Wu et al.NeurIPS 2024 · 114 citations
