ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs
Xiyao Wang, Zhengyuan Yang, Chao Feng, Yuhang Zhou, Xiaoyu Liu, Yongyuan Liang, Ming Li, Ziyi Zang, Linjie Li, Chung-Ching Lin, Kevin Lin, Furong Huang, Lijuan Wang
Abstract
Reinforcement learning (RL) has shown great effectiveness for fine-tuning large language models (LLMs) using tasks that are challenging yet easily verifiable, such as math reasoning or code generation. However, extending this success to visual perception in vision-language models (VLMs) has been impeded by the scarcity of vision-centric tasks that are simultaneously challenging and unambiguously verifiable. To this end, we introduce ViCrit (Visual Caption Hallucination Critic), an RL proxy task that trains VLMs to localize a subtle, synthetic visual hallucination injected into paragraphs of human-written image captions. Starting from a 200-word captions, we inject a single, subtle visual description error-altering a few words on objects, attributes, counts, or spatial relations-and task the model to pinpoint the corrupted span given the image and the modified caption. This formulation preserves the full perceptual difficulty while providing a binary, exactmatch reward that is easy to compute and unambiguous. Models trained with the ViCrit Task exhibit substantial gains across a variety of VL benchmarks. Crucially, the improvements transfer beyond natural-image training data to abstract image reasoning and visual math, showing promises of learning to perceive rather than barely memorizing seen objects. To facilitate evaluation, we further introduce ViCrit-Bench, a category-balanced diagnostic benchmark that systematically probes perception errors across diverse image domains and error types. Together, our results demonstrate that fine-grained hallucination criticism is an effective and generalizable objective for enhancing visual perception in VLMs.
The image showcases a social gathering of Caucasian individuals, both male and female, ranging from middle age to about 60, seated at multiple tables inside a room that appears to be a café or restaurant. The café's walls are a light brown to mustard yellow, adorned with an eclectic mix of picture frames and flags, including one particularly striking black flag with curved white stitching that reads both "true" and "false." There is a tall vertical window on the left side, offering a view of trees and parked cars outside. Hanging from the ceiling are two distinct light fixtures: a black wrought iron chandelier with six gold-colored bulbs, and a single glass pendant light with a black wire. Additionally, a lamp occupies the corner on the left side. Near this window, a woman dressed in black and wearing glasses is seated alone with an iPad on the table, a coffee cup beside her, and she is gazing out the window. Nearby, a group of four individuals, predominantly young men, are engaged in conversation and one is looking at his phone. To the right, there are smaller tables, where pairs of people, including some young women, are conversing. At one table in the lower right corner, a man with headphones and a blue jacket looks down, perhaps immersed in his own world. The atmosphere is lively, with a mix of discussions and some quiet moments of individual focus.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ff52e408-5f3c-4cdc-8a28-726dbc97e5d2Cited by top-tier papers6
- Visual Jigsaw Post-Training Improves MLLMsPenghao Wu, Yushan Zhang, Haiwen Diao, Bo Li et al.ICLR 2026 · 25 citations
- Vision-Zero: Scalable VLM Self-Evolution via Multi-Agent Self-PlayQinsi Wang, Bo Liu, Tianyi Zhou, Jing Shi et al.ICLR 2026 · 24 citations
- Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware DecodingZhongxing Xu, Zhonghua Wang, Zhe Qian, Dachuan Shi et al.CVPR 2026 · 16 citations
- Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language ModelsYu Zeng, Wenxuan Huang, Shiting Huang, Xikun Bao et al.ICLR 2026 · 11 citations
- Med-Scout: Curing MLLMs' Geometric Blindness in Medical Perception via Geometry-Aware RL Post-TrainingAnglin Liu, Ruichao Chen, Yi Lu, Hongxia Xu et al.ICML 2026
Builds on30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
Related papers
- CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement LearningLong Xing, Xiaoyi Dong, Yuhang Zang, Yuhang Cao et al.ICLR 2026 · 37 citations
- On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMsRosie Zhao, Anshul Shah, Xiaoyu Zhu, Xinke Deng et al.ICML 2026 · 10 citations
- VES-RFT: Rewarding Visual Evidence Sensitivity to Mitigate Hallucinations in Large Vision–Language ModelsXUEGE HOU, Wenshuo Li, Yali Li, Han Shu et al.CVPR 2026
- VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual ReasoningXueqing Wu, Yuheng Ding, Bingxuan Li, Pan Lu et al.CVPR 2025
- VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward ModelsJiacheng Ruan, Wenzhen Yuan, Xiqi Gao, Ye Guo et al.ICCV 2025 · 22 citations
