Babel-ImageNet: Massively Multilingual Evaluation of Vision-and-Language Representations
Gregor Geigle, Radu Timofte, Goran Glavas
摘要
Vision-and-language (VL) models with separate encoders for each modality (e.g., CLIP) have become the go-to models for zero-shot image classification and image-text retrieval. They are, however, mostly evaluated in English as multilingual benchmarks are limited in availability. We introduce Babel-ImageNet, a massively multilingual benchmark that offers (partial) translations of ImageNet labels to 100 languages, built without machine translation or manual annotation. We instead automatically obtain reliable translations by linking them -via shared WordNet synsets -to Babel-Net, a massively multilingual lexico-semantic network. We evaluate 11 public multilingual CLIP models on zero-shot image classification (ZS-IC) on our benchmark, demonstrating a significant gap between English ImageNet performance and that of high-resource languages (e.g., German or Chinese), and an even bigger gap for low-resource languages (e.g., Sinhala or Lao). Crucially, we show that the models' ZS-IC performance highly correlates with their performance in image-text retrieval, validating the use of Babel-ImageNet to evaluate multilingual models for the vast majority of languages without gold image-text data. Finally, we show that the performance of multilingual CLIP can be drastically improved for low-resource languages with parameter-efficient languagespecific training. We make our code and data publicly available: https://github. com/gregor-ge/Babel-ImageNet
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Meta CLIP 2: A Worldwide Scaling RecipeYung-Sung Chuang, Yang Li, Dong Wang, Ching-Feng Yeh 等NeurIPS 2025 · 被引用 72 次
- Multilingual Diversity Improves Vision-Language RepresentationsThao Nguyen, Matthew Wallingford, Sebastin Santy, Wei-Chiu Ma 等NeurIPS 2024 · 被引用 19 次
- LaoBench: A Large-Scale Multidimensional Lao Benchmark for Large Language ModelsJian Gao, Richeng Xuan, Zhaolu Kang, Dingshi Liao 等ACL 2026 · 被引用 1 次
- Semantic and Expressive Variations in Image Captions Across LanguagesAndre Ye, Sebastin Santy, Jena D. Hwang, Amy X. Zhang 等CVPR 2025
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- IGLUE: A Benchmark for Transfer Learning across Modalities, Tasks, and LanguagesEmanuele Bugliarello, Fangyu Liu, Jonas Pfeiffer, Siva Reddy 等ICML 2022 · 被引用 71 次
- ViTamin: Designing Scalable Vision Models in the Vision-Language EraJieneng Chen, Qihang Yu, Xiaohui Shen, Alan L. Yuille 等CVPR 2024
- Remote Sensing Vision-Language Foundation Models without Annotations via Ground Remote AlignmentUtkarsh Mall, Cheng Perng Phoo, Meilin Kelsey Liu, Carl Vondrick 等ICLR 2024 · 被引用 90 次
- Language-Driven Cross-Modal Classifier for Zero-Shot Multi-Label Image RecognitionYicheng Liu, Jie Wen, Chengliang Liu, Xiaozhao Fang 等ICML 2024 · 被引用 7 次
- Embracing Language Inclusivity and Diversity in CLIP through Continual Language LearningBang Yang, Yong Dai, Xuxin Cheng, Yaowei Li 等AAAI 2024 · 被引用 9 次
