Cross-modal Language Generation using Pivot Stabilization for Web-scale Language Coverage
Ashish V. Thapliyal, Radu Soricut
Abstract
Cross-modal language generation tasks such as image captioning are directly hurt in their ability to support non-English languages by the trend of data-hungry models combined with the lack of non-English annotations. We investigate potential solutions for combining existing language-generation annotations in English with translation capabilities in order to create solutions at web-scale in both domain and language coverage. We describe an approach called Pivot-Language Generation Stabilization (PLuGS), which leverages directly at training time both existing English annotations (gold data) as well as their machine-translated versions (silver data); at run-time, it generates first an English caption and then a corresponding target-language caption. We show that PLuGS models outperform other candidate solutions in evaluations performed over 5 different target languages, under a large-domain testset using images from the Open Images dataset. Furthermore, we find an interesting effect where the English captions generated by the PLuGS models are better than the captions generated by the original, monolingual English model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Crossmodal-3600: A Massively Multilingual Multimodal Evaluation DatasetAshish V. Thapliyal, Jordi Pont-Tuset, Xi Chen, Radu SoricutEMNLP 2022 · 31 citations
- CrossVQA: Scalably Generating Benchmarks for Systematically Testing VQA GeneralizationArjun R. Akula, Soravit Changpinyo, Boqing Gong, Piyush Sharma et al.EMNLP 2021 · 18 citations
- Semantic and Expressive Variations in Image Captions Across LanguagesAndre Ye, Sebastin Santy, Jena D. Hwang, Amy X. Zhang et al.CVPR 2025
Related papers
- Translation-Enhanced Multilingual Text-to-Image GenerationYaoyiran Li, Ching-Yun Chang, Stephen Rawls, Ivan Vulic et al.ACL 2023 · 8 citations
- PLUG: Leveraging Pivot Language in Cross-Lingual Instruction TuningZhihan Zhang, Dong-Ho Lee, Yuwei Fang, Wenhao Yu et al.ACL 2024
- Synthetic Data Augmentation for Zero-Shot Cross-Lingual Question AnsweringArij Riabi, Thomas Scialom, Rachel Keraron, Benoît Sagot et al.EMNLP 2021
- Model Selection for Cross-lingual TransferYang Chen, Alan RitterEMNLP 2021
- The Visual Prism: Refracting Images into Parallel Multilingual Descriptions with Structured Visual GuidanceChengpeng Fu, Xiaocheng Feng, Yichong Huang, Wenshuai Huo et al.AAAI 2026
