Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval
Hao Sun, Yingyan Hou, Jiayan Guo, Bo Wang, Chunyu Yang, Jinsong Ni, Yan Zhang
Abstract
Document retrieval in real-world scenarios faces significant challenges due to diverse document formats and modalities. Traditional textbased approaches rely on tailored parsing techniques that disregard layout information and are prone to errors, while recent parsing-free visual methods often struggle to capture fine-grained textual semantics in text-rich scenarios. To address these limitations, we propose Unveil, a novel visual-textual embedding framework that effectively integrates textual and visual features for robust document representation. Through knowledge distillation, we transfer the semantic understanding capabilities from the visual-textual embedding model to a purely visual model, enabling efficient parsing-free retrieval while preserving semantic fidelity. Experimental results demonstrate that our visualtextual embedding method surpasses existing approaches, while knowledge distillation successfully bridges the performance gap between visual-textual and visual-only methods, improving both retrieval accuracy and efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext daadf2d9-4b33-45af-985c-8f0f88d7b2bbCited by top-tier papers1
Ask how each one uses itBuilds on10
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- SlideVQA: A Dataset for Document Visual Question Answering on Multiple ImagesRyota Tanaka, Kyosuke Nishida, Kosuke Nishida, Taku Hasegawa et al.AAAI 2023 · 178 citations
- Large Dual Encoders Are Generalizable RetrieversJianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai et al.EMNLP 2022 · 145 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- ASQA: Factoid Questions Meet Long-Form AnswersIvan Stelmakh, Yi Luan, Bhuwan Dhingra, Ming-Wei ChangEMNLP 2022 · 51 citations
Related papers
- Unifying Multimodal Retrieval via Document Screenshot EmbeddingXueguang Ma, Sheng-Chieh Lin, Minghan Li, Wenhu Chen et al.EMNLP 2024 · 13 citations
- DocRetriever: A Plug-and-Play Framework for Multimodal Document Retrieval with Comprehensive BenchmarkRuofan Hu, Menghui Zhu, Jieming Zhu, Bo Chen et al.KDD 2026 · 1 citation
- HieRD: Hierarchical Relational Distillation for Vision-Language Embedding ModelsVinh Le, Nguyen Dang, Tu Vu, Linh Van et al.ICML 2026
- DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding ModelsSungnyun Kim, Haofu Liao, Srikar Appalaraju, Peng Tang et al.EMNLP 2024 · 3 citations
- ViSTA: Vision and Scene Text Aggregation for Cross-Modal RetrievalMengjun Cheng, Yipeng Sun, Longchao Wang, Xiongwei Zhu et al.CVPR 2022 · 86 citations
