Escaping the Likelihood Trap: Geometric Diversity Optimization for Long-Form Image Captioning
Qingmei Tang, Shuai Hao, Rong Fu, Zirui Mo, Xiang Liu, Jiaxuan Lu, Wenyu Wang
摘要
The utility of Vision-Language Models (VLMs) in reasoning and auditing tasks hinges on their ability to exhaustively describe visual scenes. However, current models exhibit a pathology we term the Likelihood Trap: standard alignment objectives, specifically MLE and KLregularization, drive generation toward generic, high-probability templates, systematically suppressing fine-grained details. To overcome this, we introduce Geo-RL, a framework that shifts the optimization target from mode-seeking persample objectives toward diversity-oriented setlevel RL. Geo-RL reformulates caption generation as maximizing the volume of a parallelotope in semantic space. By leveraging Determinantal Point Processes (DPPs), we enforce orthogonality among sampled descriptions, ensuring that they span the image's full semantic support. Crucially, we derive a closed-form leaveone-out marginal reward, enabling stable policy optimization. Empirically, Geo-RL escapes the trap, achieving a significant improvement in semantic richness and detail coverage without compromising visual grounding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- Post-training Large Language Models for Diverse High-Quality ResponsesYilei Chen, Souradip Chakraborty, Lorenz Wolf, Ioannis Paschalidis 等ICLR 2026 · 被引用 20 次
- EagleVision: A Dual-Stage Framework with BEV-grounding-based Chain-of-Thought for Spatial IntelligenceJiaxu Wan, Xu Wang, Mengwei Xie, Hang Zhang 等CVPR 2026 · 被引用 3 次
- Recognition through Reasoning: Reinforcing Image Geo-localization with Large Vision-Language ModelsLing Li, Yao Zhou, Yuxuan Liang, Fugee Tsung 等NeurIPS 2025 · 被引用 30 次
- GeoDiT: A Diffusion-based Vision-Language Model for Geospatial UnderstandingJiaqi Liu, Ronghao Fu, Haoran Liu, Lang Sun 等CVPR 2026 · 被引用 2 次
- Perceptual Flow Network for Visually Grounded ReasoningYangfu Li, Yuning Gong, Hongjian Zhan, Teng Li 等ICML 2026
