SmolDocling: An Ultra-Compact Vision-Language Model for End-To-End Multi-Modal Document Conversion
Ahmed S. Nassar, Matteo Omenetti, Maksym Lysak, Nikolaos Livathinos, Christoph Auer, Lucas Morin, Rafael Teixeira de Lima, Yusik Kim, A. Said Gurbuz, Michele Dolfi, Peter W. J. Staar
Abstract
We introduce SmolDocling, an ultra-compact vision-language model targeting end-to-end document conversion. Our model comprehensively processes entire pages by generating DocTags, a new universal markup format that captures all page elements in their full context with location. Unlike existing approaches that rely on large foundational models, or ensemble solutions that rely on handcrafted pipelines of multiple specialized models, SmolDocling offers an end-to-end conversion for accurately capturing content, structure and spatial location of document elements in a 256M parameters vision-language model. SmolDocling exhibits robust performance in correctly reproducing document features such as code listings, tables, equations, charts, lists, and more across a diverse range of document types including business documents, academic papers, technical reports, patents, and forms -- significantly extending beyond the commonly observed focus on scientific papers. Additionally, we contribute novel publicly sourced datasets for charts, tables, equations, and code recognition. Experimental results demonstrate that SmolDocling competes with other Vision Language Models that are up to 27 times larger in size, while reducing computational requirements substantially. The model is currently available, datasets will be publicly available soon.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7310b04d-3808-417c-8d56-537acecaadf7Cited by top-tier papers6
- TDATR: Improving End-to-End Table Recognition via Table Detail-Aware Learning and Cell-Level Visual AlignmentChunxia Qin, Chenyu Liu, Pengcheng Xia, Jun Du et al.CVPR 2026 · 2 citations
- DAVE: A VLM Vision Encoder for Document Understanding and Web AgentsBrandon Huang, Hang Hua, Zhuoran Yu, Trevor Darrell et al.ICLR 2026 · 2 citations
- MarkushGrapher-2: End-to-end Multimodal Recognition of Chemical StructuresTim Strohmeyer, Lucas Morin, Gerhard Ingmar Meijer, Valéry Weber et al.CVPR 2026 · 2 citations
- HCT-QA: A Benchmark for Question Answering on Human-Centric TablesMohammad Shahmeer Ahmad, Zan Ahmad Naeem, Michaël Aupetit, Ahmed K. Elmagarmid et al.ICDE 2026
- MarkushGrapher: Joint Visual and Textual Recognition of Markush StructuresLucas Morin, Valéry Weber, Ahmed Nassar, Gerhard Ingmar Meijer et al.CVPR 2025
Builds on22
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Grounding Multimodal Large Language Models to the WorldZhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao et al.ICLR 2024 · 1,170 citations
Related papers
- OvisOCR: End-to-End Document Parsing via Aligning Specialized Perception with General ReasoningJun-Peng Jiang, Shiyin Lu, An-Yang Ji, Yinglun Li et al.ICML 2026
- VDocRAG: Retrieval-Augmented Generation over Visually-Rich DocumentsRyota Tanaka, Taichi Iki, Taku Hasegawa, Kyosuke Nishida et al.CVPR 2025
- UniDoc: Unified Pretraining Framework for Document UnderstandingJiuxiang Gu, Jason Kuen, Vlad I. Morariu, Handong Zhao et al.NeurIPS 2021 · 118 citations
- DocFormer: End-to-End Transformer for Document UnderstandingSrikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie et al.ICCV 2021 · 392 citations
- SelfDoc: Self-Supervised Document Representation LearningPeizhao Li, Jiuxiang Gu, Jason Kuen, Vlad I. Morariu et al.CVPR 2021
