Where Culture Fades: Revealing the Cultural Gap in Text-to-Image Generation
Chuancheng Shi, Shangze Li, Shiming Guo, Simiao Xie, Wenhua Wu, Jingtong Dou, Chao Wu, Canran Xiao, Cong Wang, Zifeng Cheng, Fei Shen, Tat-Seng Chua
摘要
Multilingual text-to-image (T2I) models have advanced rapidly in terms of visual realism and semantic alignment, and are now widely utilized. Yet outputs vary across cultural contexts: because language carries cultural connotations, images synthesized from multilingual prompts should preserve cross-lingual cultural consistency. We conduct a comprehensive analysis showing that current T2I models often produce culturally neutral or English-biased results under multilingual prompts. Analyses of two representative models indicate that the issue stems not from missing cultural knowledge but from insufficient activation of culture-related representations. We propose a probing method that localizes culture-sensitive signals to a small set of neurons in a few fixed layers. Guided by this finding, we introduce two complementary alignment strategies: (1) inference-time cultural activation that amplifies the identified neurons without backbone fine-tuned; and (2) layer-targeted cultural enhancement that updates only culturally relevant layers. Experiments on our CultureBench demonstrate consistent improvements over strong baselines in cultural consistency while preserving fidelity and diversity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- DNA: Uncovering Universal Latent Forgery KnowledgeJingtong Dou, Chuancheng Shi, Anqi Yi, Shiming Guo 等ICML 2026 · 被引用 8 次
- Accelerating Diffusion Model Training under Minimal Budgets: A Condensation-Based PerspectiveRui Huang, Shitong Shao, Zikai Zhou, Pukun Zhao 等CVPR 2026 · 被引用 7 次
- Bias Is a Subspace, Not a Coordinate: A Geometric Rethinking of Post‑hoc Debiasing in Vision-Language ModelsDachuan Zhao, Weiyue Li, Zhenda Shen, Yushu Qiu 等CVPR 2026 · 被引用 5 次
- PhenoYieldNet: Learning Crop-Aware Phenological Responses for Multi-Crop Yield PredictionYu Luo, Xiaogang Zhu, Shan Zeng, Wei Xiang 等CVPR 2026 · 被引用 1 次
- SGMHand: Structure-Guided Modulation for Structure-Aware Hand InpaintingChuancheng Shi, Shiming Guo, Ke Shui, Yixiang Chen 等AAAI 2026 · 被引用 1 次
它引用的顶会 Paper29
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationJiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong 等NeurIPS 2023 · 被引用 1,310 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
相关 Paper
- Multilingual Text-to-Image Generation Magnifies Gender StereotypesFelix Friedrich, Katharina Hämmerl, Patrick Schramowski, Manuel Brack 等ACL 2025
- From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language ModelsMehar Bhatia, Sahithya Ravi, Aditya Chinchure, Eunjeong Hwang 等EMNLP 2024 · 被引用 6 次
- MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding EvaluationWeihua Zheng, Zhengyuan Liu, Tanmoy Chakraborty, Weiwen Xu 等ACL 2026
- AltDiffusion: A Multilingual Text-to-Image Diffusion ModelFulong Ye, Guang Liu, Xinya Wu, Ledell WuAAAI 2024 · 被引用 54 次
- Evaluating and Improving Cultural Awareness of Reward Models for LLM AlignmentHongbin Zhang, Kehai Chen, Xuefeng Bai, Yang Xiang 等ICLR 2026 · 被引用 4 次
