Conditional Variational Autoencoder for Sign Language Translation with Cross-Modal Alignment
Rui Zhao, Liang Zhang, Biao Fu, Cong Hu, Jinsong Su, Yidong Chen
Abstract
Sign language translation (SLT) aims to convert continuous sign language videos into textual sentences. As a typical multi-modal task, there exists an inherent modality gap between sign language videos and spoken language text, which makes the cross-modal alignment between visual and textual modalities crucial. However, previous studies tend to rely on an intermediate sign gloss representation to help alleviate the cross-modal problem thereby neglecting the alignment across modalities that may lead to compromised results. To address this issue, we propose a novel framework based on Conditional Variational autoencoder for SLT (CV-SLT) that facilitates direct and sufficient cross-modal alignment between sign language videos and spoken language text. Specifically, our CV-SLT consists of two paths with two Kullback-Leibler (KL) divergences to regularize the outputs of the encoder and decoder, respectively. In the prior path, the model solely relies on visual information to predict the target text; whereas in the posterior path, it simultaneously encodes visual information and textual knowledge to reconstruct the target text. The first KL divergence optimizes the conditional variational autoencoder and regularizes the encoder outputs, while the second KL divergence performs a self-distillation from the posterior path to the prior path, ensuring the consistency of decoder outputs. We further enhance the integration of textual information to the posterior path by employing a shared Attention Residual Gaussian Distribution (ARGD), which considers the textual information in the posterior path as a residual component relative to the prior path. Extensive experiments conducted on public datasets (PHOENIX14T and CSL-daily) demonstrate the effectiveness of our framework, achieving new state-of-the-art results while significantly alleviating the cross-modal representation discrepancy. The code and models are available at https://github.com/rzhao-zhsq/CV-SLT . Intruduction Sign language serves as the primary mode of communication within the deaf community. It possesses a unique grammatical structure and lexicon that conveys semantic meanings through coordinated movements of the torso, head, hands, and other body parts. This characteristic distinguishes
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 69a17894-f617-4eb4-b7c8-91ba18bbfaeaCited by top-tier papers11
- Geo-Sign: Hyperbolic Contrastive Regularisation for Geometrically Aware Sign Language TranslationEdward Fish, Richard BowdenNeurIPS 2025 · 15 citations
- SCOPE: Sign Language Contextual Processing with Embedding from LLMsYuqi Liu, Wenqian Zhang, Sihan Ren, Chengyu Huang et al.AAAI 2025 · 7 citations
- Multi-Level Cross-Modal Alignment for Speech Relation ExtractionLiang Zhang, Zhen Yang, Biao Fu, Ziyao Lu et al.EMNLP 2024 · 2 citations
- Sign-Language Datasets at Scale: A Comprehensive Survey on Resources, Benchmarks, and Annotation StandardsYiming Ni, Zhi-Qi Cheng, Jiayu Li, Wei ChengACL 2026 · 1 citation
- Think in Latent Thoughts: A New Paradigm for Gloss-Free Sign Language TranslationYiyang Jiang, Li Zhang, Xiao-Yong Wei, Li QingACL 2026 · 1 citation
Builds on10
- NVAE: A Deep Hierarchical Variational AutoencoderArash Vahdat, Jan KautzNeurIPS 2020 · 1,141 citations
- TSPNet: Hierarchical Feature Learning via Temporal Semantic Pyramid for Sign Language TranslationDongxu Li, Chenchen Xu, Xin Yu, Kaihao Zhang et al.NeurIPS 2020 · 171 citations
- Latent-Variable Non-Autoregressive Neural Machine Translation with Deterministic Inference Using a Delta PosteriorRaphael Shu, Jason Lee, Hideki Nakayama, Kyunghyun ChoAAAI 2020 · 125 citations
- A Batch Normalized Inference Network Keeps the KL Vanishing AwayQile Zhu, Wei Bi, Xiaojiang Liu, Xiyao Ma et al.ACL 2020 · 70 citations
- SLTUNET: A Simple Unified Model for Sign Language TranslationBiao Zhang, Mathias Müller, Rico SennrichICLR 2023 · 14 citations
Related papers
- CVT-SLR: Contrastive Visual-Textual Transformation for Sign Language Recognition with Variational AlignmentJiangbin Zheng, Yile Wang, Cheng Tan, Siyuan Li et al.CVPR 2023
- C2ST: Cross-modal Contextualized Sequence Transduction for Continuous Sign Language RecognitionHuaiwen Zhang, Zihang Guo, Yang Yang, Xin Liu et al.ICCV 2023 · 21 citations
- Bridging Sign and Spoken Languages: Pseudo Gloss Generation for Sign Language TranslationJianyuan Guo, Peike Li, Trevor CohnNeurIPS 2025 · 17 citations
- Leveraging the Power of MLLMs for Gloss-Free Sign Language TranslationJungeun Kim, Hyeongwoo Jeon, Jongseong Bae, Ha Young KimICCV 2025 · 10 citations
- Two-Stream Network for Sign Language Recognition and TranslationYutong Chen, Ronglai Zuo, Fangyun Wei, Yu Wu et al.NeurIPS 2022 · 288 citations
