Visually-Situated Natural Language Understanding with Contrastive Reading Model and Frozen Large Language Models
Geewook Kim, Hodong Lee, Daehee Kim, Haeji Jung, Sanghee Park, Yoonsik Kim, Sangdoo Yun, Taeho Kil, Bado Lee, Seunghyun Park
Abstract
Recent advances in Large Language Models (LLMs) have stimulated a surge of research aimed at extending their applications to the visual domain. While these models exhibit promise in generating abstract image captions and facilitating natural conversations, their performance on text-rich images still requires improvement. In this paper, we introduce Contrastive Reading Model (Cream), a novel neural architecture designed to enhance the languageimage understanding capability of LLMs by capturing intricate details that are often overlooked in existing methods. Cream combines vision and auxiliary encoders, fortified by a contrastive feature alignment technique, to achieve a more effective comprehension of language information in visually situated contexts within the images. Our approach bridges the gap between vision and language understanding, paving the way for the development of more sophisticated Document Intelligence Assistants. Through rigorous evaluations across diverse visually-situated language understanding tasks that demand reasoning capabilities, we demonstrate the compelling performance of Cream, positioning it as a prominent model in the field of visual document understanding. We provide our codebase and newly-generated datasets at https://github.com/naver-ai/cream .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual SlimmingJiaxin Zhang, Wentao Yang, Songxuan Lai, Zecheng Xie et al.AAAI 2025 · 39 citations
- A Token-Level Text Image Foundation Model for Document UnderstandingTongkun Guan, Zining Wang, Pei Fu, Zhengtao Guo et al.ICCV 2025 · 5 citations
- Qualitative Study for LLM-assisted Design Study Process: Strategies, Challenges, and RolesShaolun Ruan, Rui Sheng, Xiaolin Wen, Jiachen Wang et al.IEEE VIS 2025 · 2 citations
- On Efficient Language and Vision Assistants for Visually-Situated Natural Language Understanding: What Matters in Reading and ReasoningGeewook Kim, Minjoon SeoEMNLP 2024
- Marten: Visual Question Answering with Mask Generation for Multi-modal Document UnderstandingZining Wang, Tongkun Guan, Pei Fu, Chen Duan et al.CVPR 2025
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Enhancing Visual Document Understanding with Contrastive Learning in Large Visual-Language ModelsXin Li, Yunfei Wu, Xinghua Jiang, Zhihao Guo et al.CVPR 2024
- CREAM: Coarse-to-Fine Retrieval and Multi-modal Efficient Tuning for Document VQAJinxu Zhang, Yongqi Yu, Yu ZhangACM MM 2024 · 4 citations
- ProReason: Multi-Modal Proactive Reasoning with Decoupled Eyesight and WisdomJingqi Zhou, Sheng Wang, Jingwei Dong, Kai Liu et al.EMNLP 2025
- Enhancing Advanced Visual Reasoning Ability of Large Language ModelsZhiyuan Li, Dongnan Liu, Chaoyi Zhang, Heng Wang et al.EMNLP 2024 · 10 citations
- BLIVA: A Simple Multimodal LLM for Better Handling of Text-Rich Visual QuestionsWenbo Hu, Yifan Xu, Yi Li, Weiyue Li et al.AAAI 2024 · 209 citations
