Lune

CVPR2026顶会

Dynamics-Aware Preference Optimization for Vision-Language Models

jusheng zhang, Kaitong Cai, Jing Yang, Jian Wang, Keze Wang

出版方
2026年份

摘要

Preference-based finetuning of vision-language models (VLMs) is notoriously unstable, i.e., trivially wrong negatives inject uninformative gradients that distort optimization and degrade calibration. This work revisits this issue through the lens of learning dynamics and identifies a core pathology, i.e., the squeezing effect, where easy negatives retain large, misaligned gradients despite negligible loss. To address this, we propose Cooling-Weighted Direct Preference Optimization (CW-DPO), a two-stage framework that smooths and then stabilizes the alignment process. Stage 1 employs a constrained SFT phase with low-weight "gentle negatives" to regularize overconfident distributions and flatten the loss landscape. Stage 2 introduces a competence-aware cooling weight that adaptively scales negative gradients according to the model's average per-token log-probability, suppressing uninformative updates while emphasizing hard, on-policy contrasts. This dynamics-aware weighting effectively mitigates the squeezing effect and enables smoother convergence. Extensive and comprehensive results on the mainstream benchmarks, i.e., COCO, Flickr30k, NoCaps, MMMU, and MMBench1.1, our CW-DPO achieves stateof-the-art performance, e.g., +3.4 CIDEr over PPO and +2.4% absolute accuracy on MMMU, while improving calibration and halving convergence steps. This justifies that smoothing before cooling constitutes a simple yet general principle for robust VLM preference optimization. https://github.com/jushengzhang/Dynamics-Aware-Preference-Optimization

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper23

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖