CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy
Zhibo Yang, Jun Tang, Zhaohai Li, Pengfei Wang, Jianqiang Wan, Humen Zhong, Xuejing Liu, Mingkun Yang, Peng Wang, Shuai Bai, Lianwen Jin, Junyang Lin
Abstract
Large Multimodal Models (LMMs) have demonstrated impressive performance in recognizing document images with natural language instructions. However, it remains unclear to what extent capabilities in literacy with rich structure and fine-grained visual challenges. The current landscape lacks a comprehensive benchmark to effectively measure the literate capabilities of LMMs. Existing benchmarks are often limited by narrow scenarios and specified tasks. To this end, we introduce CC-OCR, a comprehensive benchmark that possesses a diverse range of scenarios, tasks, and challenges. CC-OCR comprises four OCR-centric tracks: multi-scene text reading, multilingual text reading, document parsing, and key information extraction. It includes 39 subsets with 7,058 full annotated images, of which 41% are sourced from real applications, and released for the first time. We evaluate nine prominent LMMs and reveal both the strengths and weaknesses of these models, particularly in text grounding, multi-orientation, and hallucination of repetition. CC-OCR aims to comprehensively evaluate the capabilities of LMMs on OCR-centered tasks, facilitating continued progress in this crucial area.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image ReasoningMingxin Huang, Yongxin Shi, Dezhi Peng, Songxuan Lai et al.ICLR 2026 · 28 citations
- Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language ModelsZhentao He, Can Zhang, Ziheng Wu, Zhenghao Chen et al.NeurIPS 2025 · 14 citations
- Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCRYulong Zhang, Tianyi Liang, Erfei Cui, Guoqing Wang et al.CVPR 2026 · 14 citations
- Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware TrainingGengluo Li, Pengyuan Lyu, Chengquan Zhang, Huawen Shen et al.CVPR 2026 · 9 citations
- Why Keep Your Doubts to Yourself? Trading Visual Uncertainties among Vision-Language ModelsJusheng Zhang, Yijia Fan, Kaitong Cai, Jing Yang et al.ICLR 2026 · 6 citations
Builds on13
- Nougat: Neural Optical Understanding for Academic DocumentsLukas Blecher, Guillem Cucurull, Thomas Scialom, Robert StojnicICLR 2024 · 243 citations
- DPText-DETR: Towards Better Scene Text Detection with Dynamic Points in TransformerMaoyuan Ye, Jing Zhang, Shanshan Zhao, Juhua Liu et al.AAAI 2023 · 123 citations
- TabPedia: Towards Comprehensive Visual Table Understanding with Concept SynergyWeichao Zhao, Hao Feng, Qi Liu, Jingqun Tang et al.NeurIPS 2024 · 97 citations
- Towards Robust Visual Information Extraction in Real World: New Dataset and Novel SolutionJiapeng Wang, Chongyu Liu, Lianwen Jin, Guozhi Tang et al.AAAI 2021 · 97 citations
- Towards End-to-End Unified Scene Text Detection and Layout AnalysisShangbang Long, Siyang Qin, Dmitry Panteleev, Alessandro Bissacco et al.CVPR 2022 · 86 citations
Related papers
- OCR-Critic: Aligning Multimodal Large Language Models' Perception through Critical FeedbackQiuna Tan, Runqi Qiao, Guanting Dong, Yifan Zhang et al.ACM MM 2025
- UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual DocumentsYifan Ji, Zhipeng Xu, Zhenghao Liu, Zulong Chen et al.ACL 2026 · 3 citations
- MIBench: Evaluating Multimodal Large Language Models over Multiple ImagesHaowei Liu, Xi Zhang, Haiyang Xu, Yaya Shi et al.EMNLP 2024 · 7 citations
- UNICBench: UNIfied Counting Benchmark for MLLMChenggang Rong, Tao Han, Zhiyuan Zhao, Yaowu Fan et al.CVPR 2026 · 3 citations
- OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive AnnotationsLinke Ouyang, Yuan Qu, Hongbin Zhou, Jiawei Zhu et al.CVPR 2025
