ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs
Xiyao Wang, Zhengyuan Yang, Chao Feng, Yuhang Zhou, Xiaoyu Liu, Yongyuan Liang, Ming Li, Ziyi Zang, Linjie Li, Chung-Ching Lin, Kevin Lin, Furong Huang, Lijuan Wang
摘要
Reinforcement learning (RL) has shown great effectiveness for fine-tuning large language models (LLMs) using tasks that are challenging yet easily verifiable, such as math reasoning or code generation. However, extending this success to visual perception in vision-language models (VLMs) has been impeded by the scarcity of vision-centric tasks that are simultaneously challenging and unambiguously verifiable. To this end, we introduce ViCrit (Visual Caption Hallucination Critic), an RL proxy task that trains VLMs to localize a subtle, synthetic visual hallucination injected into paragraphs of human-written image captions. Starting from a 200-word captions, we inject a single, subtle visual description error-altering a few words on objects, attributes, counts, or spatial relations-and task the model to pinpoint the corrupted span given the image and the modified caption. This formulation preserves the full perceptual difficulty while providing a binary, exactmatch reward that is easy to compute and unambiguous. Models trained with the ViCrit Task exhibit substantial gains across a variety of VL benchmarks. Crucially, the improvements transfer beyond natural-image training data to abstract image reasoning and visual math, showing promises of learning to perceive rather than barely memorizing seen objects. To facilitate evaluation, we further introduce ViCrit-Bench, a category-balanced diagnostic benchmark that systematically probes perception errors across diverse image domains and error types. Together, our results demonstrate that fine-grained hallucination criticism is an effective and generalizable objective for enhancing visual perception in VLMs.
The image showcases a social gathering of Caucasian individuals, both male and female, ranging from middle age to about 60, seated at multiple tables inside a room that appears to be a café or restaurant. The café's walls are a light brown to mustard yellow, adorned with an eclectic mix of picture frames and flags, including one particularly striking black flag with curved white stitching that reads both "true" and "false." There is a tall vertical window on the left side, offering a view of trees and parked cars outside. Hanging from the ceiling are two distinct light fixtures: a black wrought iron chandelier with six gold-colored bulbs, and a single glass pendant light with a black wire. Additionally, a lamp occupies the corner on the left side. Near this window, a woman dressed in black and wearing glasses is seated alone with an iPad on the table, a coffee cup beside her, and she is gazing out the window. Nearby, a group of four individuals, predominantly young men, are engaged in conversation and one is looking at his phone. To the right, there are smaller tables, where pairs of people, including some young women, are conversing. At one table in the lower right corner, a man with headphones and a blue jacket looks down, perhaps immersed in his own world. The atmosphere is lively, with a mix of discussions and some quiet moments of individual focus.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Visual Jigsaw Post-Training Improves MLLMsPenghao Wu, Yushan Zhang, Haiwen Diao, Bo Li 等ICLR 2026 · 被引用 25 次
- Vision-Zero: Scalable VLM Self-Evolution via Multi-Agent Self-PlayQinsi Wang, Bo Liu, Tianyi Zhou, Jing Shi 等ICLR 2026 · 被引用 24 次
- Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware DecodingZhongxing Xu, Zhonghua Wang, Zhe Qian, Dachuan Shi 等CVPR 2026 · 被引用 16 次
- Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language ModelsYu Zeng, Wenxuan Huang, Shiting Huang, Xikun Bao 等ICLR 2026 · 被引用 11 次
- Med-Scout: Curing MLLMs' Geometric Blindness in Medical Perception via Geometry-Aware RL Post-TrainingAnglin Liu, Ruichao Chen, Yi Lu, Hongxia Xu 等ICML 2026
它引用的顶会 Paper30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
相关 Paper
- CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement LearningLong Xing, Xiaoyi Dong, Yuhang Zang, Yuhang Cao 等ICLR 2026 · 被引用 37 次
- On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMsRosie Zhao, Anshul Shah, Xiaoyu Zhu, Xinke Deng 等ICML 2026 · 被引用 10 次
- VES-RFT: Rewarding Visual Evidence Sensitivity to Mitigate Hallucinations in Large Vision–Language ModelsXUEGE HOU, Wenshuo Li, Yali Li, Han Shu 等CVPR 2026
- VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual ReasoningXueqing Wu, Yuheng Ding, Bingxuan Li, Pan Lu 等CVPR 2025
- VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward ModelsJiacheng Ruan, Wenzhen Yuan, Xiqi Gao, Ye Guo 等ICCV 2025 · 被引用 22 次
