InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions
Ryota Tanaka, Taichi Iki, Kyosuke Nishida, Kuniko Saito, Jun Suzuki
Abstract
We study the problem of completing various visual document understanding (VDU) tasks, e.g., question answering and information extraction, on real-world documents through human-written instructions. To this end, we propose Instruct-Doc, the first large-scale collection of 30 publicly available VDU datasets, each with diverse instructions in a unified format, which covers a wide range of 12 tasks and includes open document types/formats. Furthermore, to enhance the generalization performance on VDU tasks, we design a new instruction-based document reading and understanding model, InstructDr, that connects document images, image encoders, and large language models (LLMs) through a trainable bridging module. Experiments demonstrate that InstructDr can effectively adapt to new VDU datasets, tasks, and domains via given instructions and outperforms existing multimodal LLMs and ChatGPT without specific training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 08f89b62-fc76-4bf7-bedd-89f3905195a0Cited by top-tier papers8
- DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual SlimmingJiaxin Zhang, Wentao Yang, Songxuan Lai, Zecheng Xie et al.AAAI 2025 · 39 citations
- A Token-Level Text Image Foundation Model for Document UnderstandingTongkun Guan, Zining Wang, Pei Fu, Zhengtao Guo et al.ICCV 2025 · 5 citations
- Modeling Layout Reading Order as Ordering Relations for Visually-rich Document UnderstandingChong Zhang, Yi Tu, Yixi Zhao, Chenshu Yuan et al.EMNLP 2024 · 4 citations
- DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding ModelsSungnyun Kim, Haofu Liao, Srikar Appalaraju, Peng Tang et al.EMNLP 2024 · 3 citations
- DocVLM: Make Your VLM an Efficient ReaderMor Shpigel Nacson, Aviad Aberdam, Roy Ganz, Elad Ben-Avraham et al.CVPR 2025
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
Related papers
- L-Man: A Large Multi-modal Model Unifying Human-centric TasksJialong Zuo, Ying Nie, Tianyu Guo, Huaxin Zhang et al.AAAI 2025 · 1 citation
- HRVDA: High-Resolution Visual Document AssistantChaohu Liu, Kun Yin, Haoyu Cao, Xinghua Jiang et al.CVPR 2024 · 10 citations
- ProcTag: Process Tagging for Assessing the Efficacy of Document Instruction DataYufan Shen, Chuwei Luo, Zhaoqing Zhu, Yang Chen et al.AAAI 2025 · 6 citations
- VDocRAG: Retrieval-Augmented Generation over Visually-Rich DocumentsRyota Tanaka, Taichi Iki, Taku Hasegawa, Kyosuke Nishida et al.CVPR 2025
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
