AltDiffusion: A Multilingual Text-to-Image Diffusion Model
Fulong Ye, Guang Liu, Xinya Wu, Ledell Wu
摘要
Large Text-to-Image(T2I) diffusion models have shown a remarkable capability to produce photorealistic and diverse images based on text inputs. However, existing works only support limited language input, e.g., English, Chinese, and Japanese, leaving users beyond these languages underserved and blocking the global expansion of T2I models. Therefore, this paper presents AltDiffusion, a novel multilingual T2I diffusion model that supports eighteen different languages 1 . Specifically, we first train a multilingual text encoder based on the knowledge distillation. Then we plug it into a pretrained English-only diffusion model and train the model with a two-stage schema to enhance the multilingual capability, including concept alignment and quality improvement stage on a large-scale multilingual dataset. Furthermore, we introduce a new benchmark, which includes Multilingual-General-18(MG-18) and Multilingual-Cultural-18(MC-18) datasets, to evaluate the capabilities of T2I diffusion models for generating high-quality images and capturing culture-specific concepts in different languages. Experimental results on both MG-18 and MC-18 demonstrate that AltDiffusion outperforms current state-of-the-art T2I models, e.g., Stable Diffusion in multilingual understanding, especially with respect to culture-specific concepts, while still having comparable capability for generating high-quality images. All source code and checkpoints could be found in https://github.com/superhero-7/AltDiffuson .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Where Culture Fades: Revealing the Cultural Gap in Text-to-Image GenerationChuancheng Shi, Shangze Li, Shiming Guo, Simiao Xie 等CVPR 2026 · 被引用 14 次
- TraceRouter: Robust Safety for Large Foundation Models via Path-Level InterventionChuancheng Shi, shangze li, Wenjun Lu, Wenhua Wu 等ICML 2026 · 被引用 12 次
- Neodragon: Mobile Video Generation Using Diffusion TransformerAnimesh Karnewar, Denis Korzhenkov, Ioannis Lelekas, Noor Fathima 等ICLR 2026 · 被引用 10 次
- CGI-DM: Digital Copyright Authentication for Diffusion Models via Contrasting Gradient InversionXiaoyu Wu, Yang Hua, Chumeng Liang, Jiaru Zhang 等CVPR 2024 · 被引用 2 次
- PIXELS: Progressive Image Xemplar-based Editing with Latent SurgeryShristi Das Biswas, Matthew Shreve, Xuelu Li, Prateek Singhal 等AAAI 2025 · 被引用 2 次
它引用的顶会 Paper11
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- CogView: Mastering Text-to-Image Generation via TransformersMing Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng 等NeurIPS 2021 · 被引用 1,026 次
相关 Paper
- X2i: Seamless Integration of Multimodal Understanding Into Diffusion Transformer Via Attention DistillationJian Ma, Qirong Peng, Xu Guo, Chen Chen 等ICCV 2025 · 被引用 1 次
- MultiFusion: Fusing Pre-Trained Models for Multi-Lingual, Multi-Modal Image GenerationMarco Bellagente, Manuel Brack, Hannah Teufel, Felix Friedrich 等NeurIPS 2023 · 被引用 31 次
- One Transformer Fits All Distributions in Multi-Modal Diffusion at ScaleFan Bao, Shen Nie, Kaiwen Xue, Chongxuan Li 等ICML 2023 · 被引用 236 次
- AnyText: Multilingual Visual Text Generation and EditingYuxiang Tuo, Wangmeng Xiang, Jun-Yan He, Yifeng Geng 等ICLR 2024 · 被引用 148 次
- AcT2I: Evaluating and Improving Action Depiction in Text-to-Image ModelsVatsal Malaviya, Agneet Chatterjee, Maitreya Patel, Yezhou Yang 等EMNLP 2025
