Aligning Subtitles in Sign Language Videos
Hannah Bull, Triantafyllos Afouras, Gül Varol, Samuel Albanie, Liliane Momeni, Andrew Zisserman
Abstract
The goal of this work is to temporally align asynchronous subtitles in sign language videos. In particular, we focus on sign-language interpreted TV broadcast data comprising (i) a video of continuous signing, and (ii) subtitles corresponding to the audio content. Previous work exploiting such weakly-aligned data only considered finding keyword-sign correspondences, whereas we aim to localise a complete subtitle text in continuous signing. We propose a Transformer architecture tailored for this task, which we train on manually annotated alignments covering over 15K subtitles that span 17.7 hours of video. We use BERT subtitle embeddings and CNN video representations learned for sign recognition to encode the two signals, which interact through a series of attention layers. Our model outputs frame-level predictions, i.e., for each video frame, whether it belongs to the queried subtitle or not. Through extensive evaluations, we show substantial improvements over existing alignment baselines that do not make use of subtitle text embeddings for learning. Our automatic alignment model opens up possibilities for advancing machine translation of sign languages via providing continuously synchronized video-text data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 445c1059-c993-4a14-a0fc-1e17da68b635Cited by top-tier papers6
- Open-Domain Sign Language Translation Learned from Online VideoBowen Shi, Diane Brentari, Gregory Shakhnarovich, Karen LivescuEMNLP 2022 · 39 citations
- Sentence-level Segmentation for Long Sign Language Videos with CaptionsBowen Guo, Shiwei Gan, Yafeng Yin, Xiao Liu et al.ACM MM 2025
- Lost in Translation, Found in Context: Sign Language Translation with Contextual CuesYoungjoon Jang, Haran Raajesh, Liliane Momeni, Gül Varol et al.CVPR 2025
- BoostSLT: Boosting Sign Language Translation via a Plug-and-Play Diffusion-Based Semantic EnhancerChangzhou Han, Wanlun Ma, Xi Tang, Kun Hu et al.CVPR 2026
- Transformer with Controlled Attention for Synchronous Motion CaptioningKarim Radouane, Sylvie Ranwez, Julien Lagarde, Andon TchechmedjievAAAI 2026
Builds on7
- TSPNet: Hierarchical Feature Learning via Temporal Semantic Pyramid for Sign Language TranslationDongxu Li, Chenchen Xu, Xin Yu, Kaihao Zhang et al.NeurIPS 2020 · 171 citations
- Sign Language Transformers: Joint End-to-End Sign Language Recognition and TranslationNecati Cihan Camgöz, Oscar Koller, Simon Hadfield, Richard BowdenCVPR 2020
- Transferring Cross-Domain Knowledge for Video Sign Language RecognitionDongxu Li, Xin Yu, Chenchen Xu, Lars Petersson et al.CVPR 2020
- VirTex: Learning Visual Representations From Textual AnnotationsKaran Desai, Justin JohnsonCVPR 2021
- Read and Attend: Temporal Localisation in Sign Language VideosGül Varol, Liliane Momeni, Samuel Albanie, Triantafyllos Afouras et al.CVPR 2021
Related papers
- Segment, Embed, and Align: A Universal Recipe for Aligning Subtitles to SigningZifan Jiang, Youngjoon Jang, Liliane Momeni, Gül Varol et al.ACL 2026
- SHuBERT: Self-Supervised Sign Language Representation Learning via Multi-Stream Cluster PredictionShester Gueuwou, Xiaodan Du, Greg Shakhnarovich, Karen Livescu et al.ACL 2025
- Gloss Attention for Gloss-free Sign Language TranslationAoxiong Yin, Tianyun Zhong, Li Tang, Weike Jin et al.CVPR 2023
- Sign Language Video Retrieval with Free-Form Textual QueriesAmanda Cardoso Duarte, Samuel Albanie, Xavier Giró-i-Nieto, Gül VarolCVPR 2022 · 27 citations
- LLMs are Good Sign Language TranslatorsJia Gong, Lin Geng Foo, Yixuan He, Hossein Rahmani et al.CVPR 2024
