OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation
Junyuan Zhang, Qintong Zhang, Bin Wang, Linke Ouyang, Zichen Wen, Ying Li, Ka-Ho Chow, Conghui He, Wentao Zhang
Abstract
Retrieval-augmented Generation (RAG) enhances Large Language Models (LLMs) by integrating external knowledge to reduce hallucinations and incorporate up-to-date information without retraining. As an essential part of RAG, external knowledge bases are commonly built by extracting structured data from unstructured PDF documents using Optical Character Recognition (OCR). However, given the imperfect prediction of OCR and the inherent non-uniform representation of structured data, knowledge bases inevitably contain various OCR noises. In this paper, we introduce OHRBench, the first benchmark for understanding the cascading impact of OCR on RAG systems. OHRBench includes 8,561 carefully selected unstructured document images from seven real-world RAG application domains, along with 8,498 Q&A pairs derived from multimodal elements in documents, challenging existing OCR solutions used for RAG. To better understand OCR's impact on RAG systems, we identify two primary types of OCR noise: Semantic Noise and Formatting Noise and apply perturbation to generate a set of structured data with varying degrees of each OCR noise. Using OHRBench, we first conduct a comprehensive evaluation of current OCR solutions and reveal that none is competent for constructing high-quality knowledge bases for RAG systems. We then systematically evaluate the impact of these two noise types and demonstrate the trend relationship between the degree of OCR noise and RAG performance. Our OHRBench, including PDF documents, Q&As, and the ground truth structured data are released at: https: //github.com/opendatalab/OHR-Bench
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 91ed06df-0cb3-48fd-8512-b85e52d99c4aCited by top-tier papers8
- HiChunk: Evaluating and Enhancing Retrieval Augmented Generation with Hierarchical ChunkingWensheng Lu, Keyu Chen, Zhifeng Shen, Ruizhi Qiao et al.ACL 2026 · 10 citations
- Confundo: Learning to Generate Robust Poison for Practical RAG SystemsHaoyang Hu, Zhejun Jiang, Yueming Lyu, Junyuan Zhang et al.USENIX Security 2026 · 5 citations
- TRivia: Self-supervised Fine-tuning of Vision-Language Models for Table RecognitionJunyuan Zhang, Bin Wang, Qintong Zhang, Fan Wu et al.CVPR 2026 · 5 citations
- ReAlign: Optimizing the Visual Document Retriever with Reasoning-Guided Fine-Grained AlignmentHao Yang, Yifan Ji, Zhipeng Xu, Zhenghao Liu et al.SIGIR 2026 · 4 citations
- OmniDocLayout: Towards Diverse Document Layout Generation via Coarse-to-Fine LLM LearningHengrui Kang, Zhuangcheng Gu, Zhiyuan Zhao, Zichen Wen et al.CVPR 2026 · 2 citations
Builds on13
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Benchmarking Large Language Models in Retrieval-Augmented GenerationJiawei Chen, Hongyu Lin, Xianpei Han, Le SunAAAI 2024 · 531 citations
- Nougat: Neural Optical Understanding for Academic DocumentsLukas Blecher, Guillem Cucurull, Thomas Scialom, Robert StojnicICLR 2024 · 243 citations
- Document Understanding Dataset and Evaluation (DUDE)Jordy Van Landeghem, Rafal Powalski, Rubèn Tito, Dawid Jurkiewicz et al.ICCV 2023 · 130 citations
- UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept TokensRuichuan An, Sihan Yang, Renrui Zhang, Zijun Shen et al.NeurIPS 2025 · 61 citations
Related papers
- Pandora's Box or Aladdin's Lamp: A Comprehensive Analysis Revealing the Role of RAG Noise in Large Language ModelsJinyang Wu, Shuai Zhang, Feihu Che, Mingkuan Feng et al.ACL 2025 · 12 citations
- Are We on the Right Way to Assess Document Retrieval-Augmented Generation?Wenxuan Shen, Mingjia Wang, Yaochen Wang, Dongping Chen et al.AAAI 2026
- PRGB Benchmark: A Robust Placeholder-Assisted Algorithm for Benchmarking Retrieval-Augmented GenerationZhehao Tan, Yihan Jiao, Dan Yang, Junwei Liu et al.AAAI 2026
- When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrieval-Augmented GenerationZhishang Xiang, Chuanjie Wu, Qinggang Zhang, Shengyuan Chen et al.ICLR 2026 · 56 citations
- REAL-MM-RAG: A Real-World Multi-Modal Retrieval BenchmarkNavve Wasserman, Roi Pony, Oshri Naparstek, Adi Raz Goldfarb et al.ACL 2025 · 33 citations
