Read and Attend: Temporal Localisation in Sign Language Videos
Gül Varol, Liliane Momeni, Samuel Albanie, Triantafyllos Afouras, Andrew Zisserman
Abstract
The objective of this work is to annotate sign instances across a broad vocabulary in continuous sign language. We train a Transformer model to ingest a continuous signing stream and output a sequence of written tokens on a largescale collection of signing footage with weakly-aligned subtitles. We show that through this training it acquires the ability to attend to a large vocabulary of sign instances in the input sequence, enabling their localisation. Our contributions are as follows: (1) we demonstrate the ability to leverage large quantities of continuous signing videos with weakly-aligned subtitles to localise signs in continuous sign language; (2) we employ the learned attention to automatically generate hundreds of thousands of annotations for a large sign vocabulary; (3) we collect a set of 37K manually verified sign instances across a vocabulary of 950 sign classes to support our study of sign language recognition; (4) by training on the newly annotated data from our method, we outperform the prior state of the art on the BSL-1K sign language recognition benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cfe90c14-b255-45bf-b6bf-3fa1697e45d8Cited by top-tier papers9
- Aligning Subtitles in Sign Language VideosHannah Bull, Triantafyllos Afouras, Gül Varol, Samuel Albanie et al.ICCV 2021 · 39 citations
- Sign Language Video Retrieval with Free-Form Textual QueriesAmanda Cardoso Duarte, Samuel Albanie, Xavier Giró-i-Nieto, Gül VarolCVPR 2022 · 27 citations
- Geo-Sign: Hyperbolic Contrastive Regularisation for Geometrically Aware Sign Language TranslationEdward Fish, Richard BowdenNeurIPS 2025 · 15 citations
- Text-Driven 3D Hand Motion Generation from Sign Language DataLéore Bensabath, Mathis Petrovich, Gül VarolCVPR 2026 · 5 citations
- SEDS: Semantically Enhanced Dual-Stream Encoder for Sign Language RetrievalLongtao Jiang, Min Wang, Zecheng Li, Yao Fang et al.ACM MM 2024 · 2 citations
Builds on3
- HACS: Human Action Clips and Segments Dataset for Recognition and Temporal LocalizationHang Zhao, Antonio Torralba, Lorenzo Torresani, Zhicheng YanICCV 2019 · 298 citations
- Sign Language Transformers: Joint End-to-End Sign Language Recognition and TranslationNecati Cihan Camgöz, Oscar Koller, Simon Hadfield, Richard BowdenCVPR 2020
- Transferring Cross-Domain Knowledge for Video Sign Language RecognitionDongxu Li, Xin Yu, Chenchen Xu, Lars Petersson et al.CVPR 2020
Related papers
- Gloss Attention for Gloss-free Sign Language TranslationAoxiong Yin, Tianyun Zhong, Li Tang, Weike Jin et al.CVPR 2023
- Lost in Translation, Found in Context: Sign Language Translation with Contextual CuesYoungjoon Jang, Haran Raajesh, Liliane Momeni, Gül Varol et al.CVPR 2025
- Open-Domain Sign Language Translation Learned from Online VideoBowen Shi, Diane Brentari, Gregory Shakhnarovich, Karen LivescuEMNLP 2022 · 39 citations
- YouTube-SL-25: A Large-Scale, Open-Domain Multilingual Sign Language Parallel CorpusGarrett Tanzer, Biao ZhangICLR 2025
- Segment, Embed, and Align: A Universal Recipe for Aligning Subtitles to SigningZifan Jiang, Youngjoon Jang, Liliane Momeni, Gül Varol et al.ACL 2026
