OCR-Critic: Aligning Multimodal Large Language Models' Perception through Critical Feedback
Qiuna Tan, Runqi Qiao, Guanting Dong, Yifan Zhang, Minhui Wu, Jiapeng Wang, Miaoxuan Zhang, Yida Xu, Chong Sun, Chen Li, Honggang Zhang
摘要
Recent advancements in Large Multimodal Models have demonstrated impressive performance in various tasks. However, their capabilities in error detection and resolution for Optical Character Recognition (OCR) remain underexplored. To address this gap, we construct the first visual instruction tuning dataset specifically for detailed OCR error analysis. Building on this foundation, we develop a universal, plug-and-play OCR-Critic model that incorporates three novel dynamic alignment strategies. These strategies systematically mitigate LMMs' weaknesses in OCR tasks by providing coarse-to-fine error feedback. To comprehensively evaluate these capabilities, we introduce OCR-ERROR, a benchmark designed to assess LMMs' ability to detect and categorize OCR errors, covering two task types, diverse error categories, and 2,400 rigorously validated samples. Experimental results show that OCR-Critic effectively identifies fine-grained OCR errors across multiple domains. With the integration of our dynamic alignment strategies, the LMM further achieves substantial performance gains on four prominent benchmarks, demonstrating both versatility and effectiveness.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in LiteracyZhibo Yang, Jun Tang, Zhaohai Li, Pengfei Wang 等ICCV 2025 · 被引用 15 次
- LLaVA-Critic: Learning to Evaluate Multimodal ModelsTianyi Xiong, Xiyao Wang, Dong Guo, Qinghao Ye 等CVPR 2025
- Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction TuningJi Hyeok Jung, Eun Tae Kim, Seo Yeon Kim, Joo Ho Lee 等CVPR 2025
- Let Language Constrain Geometry: Vision–Language Models as Semantic and Spatial Critics for 3D GenerationWeimin Bai, Yubo Li, Weijian Luo, Zeqiang Lai 等ICML 2026 · 被引用 1 次
- INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction TuningWujian Peng, Lingchen Meng, Yitong Chen, Yiweng Xie 等NeurIPS 2025 · 被引用 7 次
