PHD: Pixel-Based Language Modeling of Historical Documents
Nadav Borenstein, Phillip Rust, Desmond Elliott, Isabelle Augenstein
Abstract
The digitisation of historical documents has provided historians with unprecedented research opportunities. Yet, the conventional approach to analysing historical documents involves converting them from images to text using OCR, a process that overlooks the potential benefits of treating them as images and introduces high levels of noise. To bridge this gap, we take advantage of recent advancements in pixel-based language models trained to reconstruct masked patches of pixels instead of predicting token distributions. Due to the scarcity of real historical scans, we propose a novel method for generating synthetic scans to resemble real historical documents. We then pre-train our model, PHD, on a combination of synthetic scans and real historical newspapers from the 1700-1900 period. Through our experiments, we demonstrate that PHD exhibits high proficiency in reconstructing masked image patches and provide evidence of our model's noteworthy language understanding capabilities. Notably, we successfully apply our model to a historical QA task, highlighting its usefulness in this domain.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Pixology: Probing the Linguistic and Visual Capabilities of Pixel-based Language ModelsKushal Tatariya, Vladimir Araujo, Thomas Bauwens, Miryam de LhoneuxEMNLP 2024 · 2 citations
- Towards Natural Language-Based Document Image Retrieval: New Dataset and BenchmarkHao Guo, Xugong Qin, Jun Jie Ou Yang, Peng Zhang et al.CVPR 2025
Builds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingKenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu et al.ICML 2023 · 426 citations
- DocFormer: End-to-End Transformer for Document UnderstandingSrikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie et al.ICCV 2021 · 392 citations
- DiT: Self-supervised Pre-training for Document Image TransformerJunlong Li, Yiheng Xu, Tengchao Lv, Lei Cui et al.ACM MM 2022 · 184 citations
- Language Modelling with PixelsPhillip Rust, Jonas F. Lotz, Emanuele Bugliarello, Elizabeth Salesky et al.ICLR 2023 · 17 citations
Related papers
- PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR AccuracyShuhao Guan, Moule Lin, Cheng Xu, Xinyi Liu et al.ACL 2025
- Reviving Cultural Heritage: A Novel Approach for Comprehensive Historical Document RestorationYuyi Zhang, Peirong Zhang, Zhenhua Yang, Pengyu Yan et al.ACL 2025 · 5 citations
- Predicting the Original Appearance of Damaged Historical DocumentsZhenhua Yang, Dezhi Peng, Yongxin Shi, Yuyi Zhang et al.AAAI 2025 · 8 citations
- StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-trainingYuechen Yu, Yulin Li, Chengquan Zhang, Xiaoqiang Zhang et al.ICLR 2023 · 18 citations
- TrOCR: Transformer-Based Optical Character Recognition with Pre-trained ModelsMinghao Li, Tengchao Lv, Jingye Chen, Lei Cui et al.AAAI 2023 · 607 citations
