Lune

NeurIPS2025Top-tier venue

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs

Xiyao Wang, Zhengyuan Yang, Chao Feng, Yuhang Zhou, Xiaoyu Liu, Yongyuan Liang, Ming Li, Ziyi Zang, Linjie Li, Chung-Ching Lin, Kevin Lin, Furong Huang, Lijuan Wang

2025Year
27Citations
6Top-tier citations

Abstract

Reinforcement learning (RL) has shown great effectiveness for fine-tuning large language models (LLMs) using tasks that are challenging yet easily verifiable, such as math reasoning or code generation. However, extending this success to visual perception in vision-language models (VLMs) has been impeded by the scarcity of vision-centric tasks that are simultaneously challenging and unambiguously verifiable. To this end, we introduce ViCrit (Visual Caption Hallucination Critic), an RL proxy task that trains VLMs to localize a subtle, synthetic visual hallucination injected into paragraphs of human-written image captions. Starting from a 200-word captions, we inject a single, subtle visual description error-altering a few words on objects, attributes, counts, or spatial relations-and task the model to pinpoint the corrupted span given the image and the modified caption. This formulation preserves the full perceptual difficulty while providing a binary, exactmatch reward that is easy to compute and unambiguous. Models trained with the ViCrit Task exhibit substantial gains across a variety of VL benchmarks. Crucially, the improvements transfer beyond natural-image training data to abstract image reasoning and visual math, showing promises of learning to perceive rather than barely memorizing seen objects. To facilitate evaluation, we further introduce ViCrit-Bench, a category-balanced diagnostic benchmark that systematically probes perception errors across diverse image domains and error types. Together, our results demonstrate that fine-grained hallucination criticism is an effective and generalizable objective for enhancing visual perception in VLMs.

The image showcases a social gathering of Caucasian individuals, both male and female, ranging from middle age to about 60, seated at multiple tables inside a room that appears to be a café or restaurant. The café's walls are a light brown to mustard yellow, adorned with an eclectic mix of picture frames and flags, including one particularly striking black flag with curved white stitching that reads both "true" and "false." There is a tall vertical window on the left side, offering a view of trees and parked cars outside. Hanging from the ceiling are two distinct light fixtures: a black wrought iron chandelier with six gold-colored bulbs, and a single glass pendant light with a black wire. Additionally, a lamp occupies the corner on the left side. Near this window, a woman dressed in black and wearing glasses is seated alone with an iPad on the table, a coffee cup beside her, and she is gazing out the window. Nearby, a group of four individuals, predominantly young men, are engaged in conversation and one is looking at his phone. To the right, there are smaller tables, where pairs of people, including some young women, are conversing. At one table in the lower right corner, a man with headphones and a blue jacket looks down, perhaps immersed in his own world. The atmosphere is lively, with a mix of discussions and some quiet moments of individual focus.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext ff52e408-5f3c-4cdc-8a28-726dbc97e5d2

Cited by top-tier papers6

Ask how each one uses it

Builds on30

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines