The Visual Prism: Refracting Images into Parallel Multilingual Descriptions with Structured Visual Guidance
Chengpeng Fu, Xiaocheng Feng, Yichong Huang, Wenshuai Huo, Baohang Li, Yang Xiang, Ting Liu
摘要
Parallel corpora, as the foundation of machine translation, remain crucial even in the era of large language models (LLMs) for pre-training and fine-tuning. However, annotating parallel corpora is extremely costly, as it requires annotators to be proficient in multiple languages. To reduce this cost, prior work has explored image-pivoted corpus synthesis, generating multilingual captions for the same image as pseudo-parallel data. Unfortunately, these pseudo corpora suffer from the serious issue of multilingual focus divergence, i.e., the model attending to distinct aspects of the image when generating captions in different languages. To address this problem, we propose a method called PRISMS (Parallel Refracting ImageS into Multilingual descriptions with Structured visual guidance), which leverages semantic graphs as structured visual guidance to unify the focus of multilingual captions. To ensure adherence to this guidance, we introduce two key techniques: supervised fine-tuning using self-generated instructional data, and reinforcement learning with a reward signal based on semantic graph consistency. Experimental results on five languages show that our PRISMS significantly improves the image-pivot parallel corpora synthesis, enabling LLMs to achieve translation performance comparable to that of models trained on manually annotated corpora.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- The Unreasonable Effectiveness of Few-shot Learning for Machine TranslationXavier Garcia, Yamini Bansal, Colin Cherry, George F. Foster 等ICML 2023 · 被引用 133 次
- Unsupervised Multimodal Neural Machine Translation with Pseudo Visual PivotingPo-Yao Huang, Junjie Hu, Xiaojun Chang, Alexander G. HauptmannACL 2020 · 被引用 43 次
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 被引用 6 次
相关 Paper
- CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement LearningLong Xing, Xiaoyi Dong, Yuhang Zang, Yuhang Cao 等ICLR 2026 · 被引用 37 次
- Enhancing Non-English Capabilities of English-Centric Large Language Models Through Deep Supervision Fine-TuningWenshuai Huo, Xiaocheng Feng, Yichong Huang, Chengpeng Fu 等AAAI 2025 · 被引用 11 次
- Cross2StrA: Unpaired Cross-lingual Image Captioning with Cross-lingual Cross-modal Structure-pivoted AlignmentShengqiong Wu, Hao Fei, Wei Ji, Tat-Seng ChuaACL 2023 · 被引用 42 次
- Prompt Refinement with Image Pivot for Text-to-Image GenerationJingtao Zhan, Qingyao Ai, Yiqun Liu, Yingwei Pan 等ACL 2024
- From Utterance to Vividity: Training Expressive Subtitle Translation LLM via Adaptive Local Preference OptimizationChaoqun Cui, Shijing Wang, Liangbin Huang, Qingqing Gu 等ICLR 2026 · 被引用 1 次
