SimulSLT: End-to-End Simultaneous Sign Language Translation
Aoxiong Yin, Zhou Zhao, Jinglin Liu, Weike Jin, Meng Zhang, Xingshan Zeng, Xiaofei He
Abstract
Sign language translation as a kind of technology with profound social significance has attracted growing researchers' interest in recent years. However, the existing sign language translation methods need to read all the videos before starting the translation, which leads to a high inference latency and also limits their application in real-life scenarios. To solve this problem, we propose Simul-SLT, the first end-to-end simultaneous sign language translation model, which can translate sign language videos into target text concurrently. SimulSLT is composed of a text decoder, a boundary predictor, and a masked encoder. We 1) use the wait-k strategy for simultaneous translation. 2) design a novel boundary predictor based on the integrate-and-fire module to output the gloss boundary, which is used to model the correspondence between the sign language video and the gloss. 3) propose an innovative re-encode method to help the model obtain more abundant contextual information, which allows the existing video features to interact fully. The experimental results conducted on the RWTH-PHOENIX-Weather 2014T dataset show that SimulSLT achieves BLEU scores that exceed the latest end-to-end non-simultaneous sign language translation model while maintaining low latency, which proves the effectiveness of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8b8e0529-6ee1-420e-88c5-4eac70ebca9aCited by top-tier papers13
- MLSLT: Towards Multilingual Sign Language TranslationAoxiong Yin, Zhou Zhao, Weike Jin, Meng Zhang et al.CVPR 2022 · 46 citations
- TranSpeech: Speech-to-Speech Translation With Bilateral PerturbationRongjie Huang, Jinglin Liu, Huadai Liu, Yi Ren et al.ICLR 2023 · 17 citations
- Towards Online Continuous Sign Language Recognition and TranslationRonglai Zuo, Fangyun Wei, Brian MakEMNLP 2024 · 14 citations
- Towards Real-Time Sign Language Recognition and Translation on Edge DevicesShiwei Gan, Yafeng Yin, Zhiwei Jiang, Lei Xie et al.ACM MM 2023 · 12 citations
- OpenSR: Open-Modality Speech Recognition via Maintaining Multi-Modality AlignmentXize Cheng, Tao Jin, Linjun Li, Wang Lin et al.ACL 2023 · 10 citations
Builds on7
- TSPNet: Hierarchical Feature Learning via Temporal Semantic Pyramid for Sign Language TranslationDongxu Li, Chenchen Xu, Xin Yu, Kaihao Zhang et al.NeurIPS 2020 · 171 citations
- Monotonic Multihead AttentionXutai Ma, Juan Miguel Pino, James Cross, Liezl Puzon et al.ICLR 2020 · 148 citations
- SimulSpeech: End-to-End Simultaneous Speech to Text TranslationYi Ren, Jinglin Liu, Xu Tan, Chen Zhang et al.ACL 2020 · 81 citations
- Towards Fast and High-Quality Sign Language ProductionWencan Huang, Wenwen Pan, Zhou Zhao, Qi TianACM MM 2021 · 43 citations
- Contrastive Disentangled Meta-Learning for Signer-Independent Sign Language TranslationTao Jin, Zhou ZhaoACM MM 2021 · 23 citations
Related papers
- Gloss Attention for Gloss-free Sign Language TranslationAoxiong Yin, Tianyun Zhong, Li Tang, Weike Jin et al.CVPR 2023
- Sign Language Transformers: Joint End-to-End Sign Language Recognition and TranslationNecati Cihan Camgöz, Oscar Koller, Simon Hadfield, Richard BowdenCVPR 2020
- SLTUNET: A Simple Unified Model for Sign Language TranslationBiao Zhang, Mathias Müller, Rico SennrichICLR 2023 · 14 citations
- Sign2GPT: Leveraging Large Language Models for Gloss-Free Sign Language TranslationRyan Wong, Necati Cihan Camgöz, Richard BowdenICLR 2024 · 58 citations
- A Simple Multi-Modality Transfer Learning Baseline for Sign Language TranslationYutong Chen, Fangyun Wei, Xiao Sun, Zhirong Wu et al.CVPR 2022 · 137 citations
