GIVL: Improving Geographical Inclusivity of Vision-Language Models with Pre-Training Methods
Da Yin, Feng Gao, Govind Thattai, Michael Johnston, Kai-Wei Chang
Abstract
West non-West 11% 5% 66.8% 70.4% Diverse Scenarios in Different Regions Wedding Wedding Figure 1. Scenarios around the world including festivals and weddings. Even the same scenarios have distinct visual characteristics across regions (a.k.a. geographically diverse). Compared with prior Vision-Language Pre-trained Models (VLPs), GIVL achieves much better performance on non-Western data in GD-VCR [48]. GIVL can also make the gap between Western and non-Western cases much closer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 285f7b20-7559-4152-b605-94ef244a92d3Cited by top-tier papers8
- Evaluating Visual and Cultural Interpretation: The K-Viscuit Benchmark with Human-VLM CollaborationChaeHun Park, Yujin Baek, Jaeseok Kim, Yu-Jung Heo et al.ACL 2025 · 17 citations
- Culture in Action: Evaluating Text-to-Image Models through Social ActivitiesSina Malakouti, Boqing Gong, Adriana KovashkaICLR 2026 · 9 citations
- PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal RetrieversWeizhe Lin, Jingbiao Mei, Jinghong Chen, Bill ByrneACL 2024 · 8 citations
- See It from My Perspective: How Language Affects Cultural Bias in Image UnderstandingAmith Ananthram, Elias Stengel-Eskin, Mohit Bansal, Kathleen McKeownICLR 2025 · 3 citations
- Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning DistractorJiali Chen, Xusen Hei, Yuqi Xue, Yuancheng Wei et al.ACM MM 2024 · 3 citations
Builds on20
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu et al.AAAI 2020 · 1,047 citations
Related papers
- Broaden the Vision: Geo-Diverse Visual Commonsense ReasoningDa Yin, Liunian Harold Li, Ziniu Hu, Nanyun Peng et al.EMNLP 2021 · 32 citations
- From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language ModelsMehar Bhatia, Sahithya Ravi, Aditya Chinchure, Eunjeong Hwang et al.EMNLP 2024 · 6 citations
- No Filter: Cultural and Socioeconomic Diversity in Contrastive Vision-Language ModelsAngéline Pouget, Lucas Beyer, Emanuele Bugliarello, Xiao Wang et al.NeurIPS 2024 · 17 citations
- GeoMLAMA: Geo-Diverse Commonsense Probing on Multilingual Pre-Trained Language ModelsDa Yin, Hritik Bansal, Masoud Monajatipoor, Liunian Harold Li et al.EMNLP 2022 · 27 citations
- LDP: Generalizing to Multilingual Visual Information Extraction by Language Decoupled PretrainingHuawen Shen, Gengluo Li, Jinwen Zhong, Yu ZhouAAAI 2025 · 4 citations
