Seeing Through Ambiguity: Effective Video-guided Machine Translation via Chaotic Fusion and Causally Aligned Spatio-temporal Attention
Jiawei Zheng, Feiyan Liu, Xiaoli Wang
Abstract
Video-guided machine translation (VMT) involves taking text and video modalities as inputs, leveraging visual context to resolve the semantic ambiguities for improving the translation quality. This task remains challenging due to the difficulty of effective cross-modal integration and visual grounding. To address the issues, we propose a novel VMT model that combines temporal video and spatial keyframe streams by providing complementary visual cues. We develop a chaotic fusion mechanism to integrate modality-specific representations from various modalities that help capture semantic interactions between visual and textual cues. To improve visual grounding, a causally aligned spatio-temporal attention mechanism is also designed to enhance semantic alignment by refining decoder-side attention over the video and keyframe streams, respectively. We further propose PolyVTE, an evaluation dataset targeting polysemous ambiguities in VMT. Results on VATEX and PolyVTE datasets show that our model outperforms state-of-the-art models. The results also prove that using keyframe and video modalities significantly improves disambiguation capabilities. The PolyVTE dataset is available at https://github.com/zheng5d/PolyVTE.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 705c62b9-bb95-404b-a8a4-233070bb6d99Cited by top-tier papers1
Ask how each one uses itRelated papers
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li et al.ICCV 2019 · 688 citations
- Video-Helpful Multimodal Machine TranslationYihang Li, Shuichiro Shimizu, Chenhui Chu, Sadao Kurohashi et al.EMNLP 2023 · 1 citation
- Let Me Finish My Sentence: Video Temporal Grounding with Holistic Text UnderstandingJongbhin Woo, Hyeonggon Ryu, Youngjoon Jang, Jae-Won Cho et al.ACM MM 2024 · 3 citations
- SHIFT: Selected Helpful Informative Frame for Video-guided Machine TranslationBoyu Guan, Chuang Han, Yining Zhang, Yupu Liang et al.EMNLP 2025
- Virtual Visual-Guided Domain-Shadow Fusion via Modal Exchanging for Domain-Specific Multi-Modal Neural Machine TranslationZhenyu Hou, Junjun GuoACM MM 2024 · 4 citations
