Cross-modal Language Generation using Pivot Stabilization for Web-scale Language Coverage
Ashish V. Thapliyal, Radu Soricut
摘要
Cross-modal language generation tasks such as image captioning are directly hurt in their ability to support non-English languages by the trend of data-hungry models combined with the lack of non-English annotations. We investigate potential solutions for combining existing language-generation annotations in English with translation capabilities in order to create solutions at web-scale in both domain and language coverage. We describe an approach called Pivot-Language Generation Stabilization (PLuGS), which leverages directly at training time both existing English annotations (gold data) as well as their machine-translated versions (silver data); at run-time, it generates first an English caption and then a corresponding target-language caption. We show that PLuGS models outperform other candidate solutions in evaluations performed over 5 different target languages, under a large-domain testset using images from the Open Images dataset. Furthermore, we find an interesting effect where the English captions generated by the PLuGS models are better than the captions generated by the original, monolingual English model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Crossmodal-3600: A Massively Multilingual Multimodal Evaluation DatasetAshish V. Thapliyal, Jordi Pont-Tuset, Xi Chen, Radu SoricutEMNLP 2022 · 被引用 31 次
- CrossVQA: Scalably Generating Benchmarks for Systematically Testing VQA GeneralizationArjun R. Akula, Soravit Changpinyo, Boqing Gong, Piyush Sharma 等EMNLP 2021 · 被引用 18 次
- Semantic and Expressive Variations in Image Captions Across LanguagesAndre Ye, Sebastin Santy, Jena D. Hwang, Amy X. Zhang 等CVPR 2025
相关 Paper
- Translation-Enhanced Multilingual Text-to-Image GenerationYaoyiran Li, Ching-Yun Chang, Stephen Rawls, Ivan Vulic 等ACL 2023 · 被引用 8 次
- PLUG: Leveraging Pivot Language in Cross-Lingual Instruction TuningZhihan Zhang, Dong-Ho Lee, Yuwei Fang, Wenhao Yu 等ACL 2024
- Synthetic Data Augmentation for Zero-Shot Cross-Lingual Question AnsweringArij Riabi, Thomas Scialom, Rachel Keraron, Benoît Sagot 等EMNLP 2021
- Model Selection for Cross-lingual TransferYang Chen, Alan RitterEMNLP 2021
- The Visual Prism: Refracting Images into Parallel Multilingual Descriptions with Structured Visual GuidanceChengpeng Fu, Xiaocheng Feng, Yichong Huang, Wenshuai Huo 等AAAI 2026
