Lune

ICML2026Top-tier venue

Escaping the Likelihood Trap: Geometric Diversity Optimization for Long-Form Image Captioning

Qingmei Tang, Shuai Hao, Rong Fu, Zirui Mo, Xiang Liu, Jiaxuan Lu, Wenyu Wang

2026Year

Abstract

The utility of Vision-Language Models (VLMs) in reasoning and auditing tasks hinges on their ability to exhaustively describe visual scenes. However, current models exhibit a pathology we term the Likelihood Trap: standard alignment objectives, specifically MLE and KLregularization, drive generation toward generic, high-probability templates, systematically suppressing fine-grained details. To overcome this, we introduce Geo-RL, a framework that shifts the optimization target from mode-seeking persample objectives toward diversity-oriented setlevel RL. Geo-RL reformulates caption generation as maximizing the volume of a parallelotope in semantic space. By leveraging Determinantal Point Processes (DPPs), we enforce orthogonality among sampled descriptions, ensuring that they span the image's full semantic support. Crucially, we derive a closed-form leaveone-out marginal reward, enabling stable policy optimization. Empirically, Geo-RL escapes the trap, achieving a significant improvement in semantic richness and detail coverage without compromising visual grounding.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext f9fb72a5-74ed-4900-9f45-f9d3e072aa1b

Builds on16

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines