Lune

ACM MM2025顶会

Seeing Through Ambiguity: Effective Video-guided Machine Translation via Chaotic Fusion and Causally Aligned Spatio-temporal Attention

Jiawei Zheng, Feiyan Liu, Xiaoli Wang

2025年份
1顶会引用

摘要

Video-guided machine translation (VMT) involves taking text and video modalities as inputs, leveraging visual context to resolve the semantic ambiguities for improving the translation quality. This task remains challenging due to the difficulty of effective cross-modal integration and visual grounding. To address the issues, we propose a novel VMT model that combines temporal video and spatial keyframe streams by providing complementary visual cues. We develop a chaotic fusion mechanism to integrate modality-specific representations from various modalities that help capture semantic interactions between visual and textual cues. To improve visual grounding, a causally aligned spatio-temporal attention mechanism is also designed to enhance semantic alignment by refining decoder-side attention over the video and keyframe streams, respectively. We further propose PolyVTE, an evaluation dataset targeting polysemous ambiguities in VMT. Results on VATEX and PolyVTE datasets show that our model outperforms state-of-the-art models. The results also prove that using keyframe and video modalities significantly improves disambiguation capabilities. The PolyVTE dataset is available at https://github.com/zheng5d/PolyVTE.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖