Multilingual Conceptual Coverage in Text-to-Image Models
Michael Saxon, William Yang Wang
摘要
We propose “Conceptual Coverage Across Languages” (CoCo-CroLa), a technique for benchmarking the degree to which any generative text-to-image system provides multilingual parity to its training language in terms of tangible nouns. For each model we can assess “conceptual coverage” of a given target language relative to a source language by comparing the population of images generated for a series of tangible nouns in the source language to the population of images generated for each noun under translation in the target language. This technique allows us to estimate how well-suited a model is to a target language as well as identify model-specific weaknesses, spurious correlations, and biases without a-priori assumptions. We demonstrate how it can be used to benchmark T2I models in terms of multilinguality, and how despite its simplicity it is a good proxy for impressive generalization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Inspecting the Geographical Representativeness of Images from Text-to-Image ModelsAbhipsa Basu, R. Venkatesh Babu, Danish PruthiICCV 2023 · 被引用 54 次
- Who Evaluates the Evaluations? Objectively Scoring Text-to-Image Prompt Coherence Metrics with T2IScoreScore (TS2)Michael Saxon, Fatima Jahara, Mahsa Khoshnoodi, Yujie Lu 等NeurIPS 2024 · 被引用 19 次
- ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion ModelsMaitreya Patel, Tejas Gokhale, Chitta Baral, Yezhou YangAAAI 2024 · 被引用 16 次
- Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-ThoughtVaishnavi Himakunthala, Andy Ouyang, Daniel Rose, Ryan He 等EMNLP 2023 · 被引用 5 次
- Multilingual Text-to-Image Generation Magnifies Gender StereotypesFelix Friedrich, Katharina Hämmerl, Patrick Schramowski, Manuel Brack 等ACL 2025
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
相关 Paper
- Post-Training Language Models for Crosslingual ConsistencyTianyu Liu, Jirui Qi, Mrinmaya Sachan, Ryan Cotterell 等ICML 2026
- Translation-Enhanced Multilingual Text-to-Image GenerationYaoyiran Li, Ching-Yun Chang, Stephen Rawls, Ivan Vulic 等ACL 2023 · 被引用 8 次
- DIMCIM: A Quantitative Evaluation Framework for Default-Mode Diversity and Generalization in Text-to-Image Generative ModelsRevant Teotia, Candace Ross, Karen Ullrich, Sumit Chopra 等ICCV 2025 · 被引用 1 次
- Generalising Multilingual Concept-to-Text NLG with Language Agnostic DelexicalisationGiulio Zhou, Gerasimos LampourasACL 2021
- XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralisationJunjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig 等ICML 2020 · 被引用 1,132 次
