POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document Conversion
Yuan Liu, Zhongyin Zhao, Le Tian, Haicheng Wang, Xubing Ye, Yangxiu You, Zilin Yu, Chuhan Wu, Xiao Zhou, Yang Yu, Jie Zhou
Abstract
High-quality labeled data is essential for training accurate document conversion models, particularly in domains with complex formats such as tables, formulas, and multi-column text. However, manual annotation is both costly and time-consuming, while automatic labeling using existing models often lacks accuracy in handling such challenging scenarios. Consequently, training student models by distilling outputs from teacher models can significantly limit their performance in real-world applications. In this paper, we propose a fully automated, distillation-free framework comprising two stages for constructing high-quality document extraction datasets and models capable of handling diverse document formats and layouts. In the first stage, we introduce a method for generating large-scale, diverse synthetic data, which enables a model to extract key elements in a unified format with strong initial performance. In the second stage, we present a selfimprovement approach that further adapts the model, initially trained on synthetic data, to real-world documents. Specifically, we first use the fine-tuned model to annotate real documents, then apply a suite of filtering strategies to verify annotation quality, and finally retrain the model on the verified dataset. By iteratively repeating this process, we progressively enhance both the model's conversion capabilities and the quality of the generated data. We train a public POINTS-1.5 model to obtain POINTS-Reader, which surpasses many existing public and proprietary models of comparable or larger size. Our model is available at https: //github.com/Tencent/POINTS-Reader 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware TrainingGengluo Li, Pengyuan Lyu, Chengquan Zhang, Huawen Shen et al.CVPR 2026 · 9 citations
- Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCRYufeng Zhong, Lei Chen, Zhixiong Zeng, Xuanle Zhao et al.CVPR 2026 · 8 citations
- TRivia: Self-supervised Fine-tuning of Vision-Language Models for Table RecognitionJunyuan Zhang, Bin Wang, Qintong Zhang, Fan Wu et al.CVPR 2026 · 5 citations
- Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual ProcessingCheng Cui, Ting Sun, Suyin Liang, Tingquan Gao et al.CVPR 2026 · 3 citations
- SEA-Vision: A Multilingual Benchmark for Comprehensive Document and Scene Text Understanding in Southeast AsiaPengfei Yue, Xingran Zhao, Juntao Chen, Peng Hou et al.CVPR 2026 · 2 citations
Builds on6
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingKenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu et al.ICML 2023 · 426 citations
- Nougat: Neural Optical Understanding for Academic DocumentsLukas Blecher, Guillem Cucurull, Thomas Scialom, Robert StojnicICLR 2024 · 243 citations
- OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive AnnotationsLinke Ouyang, Yuan Qu, Hongbin Zhou, Jiawei Zhu et al.CVPR 2025
Related papers
- Document Registration: Towards Automated Labeling of Pixel-Level Alignment Between Warped-Flat DocumentsWeiguang Zhang, Qiufeng Wang, Kaizhu Huang, Xiaowei Huang et al.ACM MM 2024 · 1 citation
- S2ynRE: Two-stage Self-training with Synthetic data for Low-resource Relation ExtractionBenfeng Xu, Quan Wang, Yajuan Lyu, Dai Dai et al.ACL 2023 · 13 citations
- DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding ModelsSungnyun Kim, Haofu Liao, Srikar Appalaraju, Peng Tang et al.EMNLP 2024 · 3 citations
- OD: Optimization-free Dataset Distillation for Object DetectionSalwa K. Al Khatib, Ahmed Elhagry, Shitong Shao, Zhiqiang ShenICLR 2026 · 2 citations
- Distantly-Supervised Named Entity Recognition with Noise-Robust Learning and Language Model Augmented Self-TrainingYu Meng, Yunyi Zhang, Jiaxin Huang, Xuan Wang et al.EMNLP 2021 · 50 citations
