Stochastic Transformer Networks with Linear Competing Units: Application to end-to-end SL Translation
Andreas Voskou, Konstantinos P. Panousis, Dimitrios I. Kosmopoulos, Dimitris N. Metaxas, Sotirios Chatzis
摘要
Automating sign language translation (SLT) is a challenging real-world application. Despite its societal importance, though, research progress in the field remains rather poor. Crucially, existing methods that yield viable performance necessitate the availability of laborious to obtain gloss sequence groundtruth. In this paper, we attenuate this need, by introducing an end-to-end SLT model that does not entail explicit use of glosses; the model only needs text groundtruth. This is in stark contrast to existing end-to-end models that use gloss sequence groundtruth, either in the form of a modality that is recognized at an intermediate model stage, or in the form of a parallel output process, jointly trained with the SLT model. Our approach constitutes a Transformer network with a novel type of layers that combines: (i) local winner-takes-all (LWTA) layers with stochastic winner sampling, instead of conventional ReLU layers, (ii) stochastic weights with posterior distributions estimated via variational inference, and (iii) a weight compression technique at inference time that exploits estimated posterior variance to perform massive, almost lossless compression. We demonstrate that our approach can reach the currently best reported BLEU-4 score on the PHOENIX 2014T benchmark, but without making use of glosses for model training, and with a memory footprint reduced by more than 70%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Gloss-free Sign Language Translation: Improving from Visual-Language PretrainingBenjia Zhou, Zhigang Chen, Albert Clapés, Jun Wan 等ICCV 2023 · 被引用 123 次
- Stochastic Deep Networks with Linear Competing Units for Model-Agnostic Meta-LearningKonstantinos Kalais, Sotirios ChatzisICML 2022 · 被引用 10 次
- Leveraging the Power of MLLMs for Gloss-Free Sign Language TranslationJungeun Kim, Hyeongwoo Jeon, Jongseong Bae, Ha Young KimICCV 2025 · 被引用 10 次
- DISCOVER: Making Vision Networks Interpretable via Competition and DissectionKonstantinos P. Panousis, Sotirios ChatzisNeurIPS 2023 · 被引用 9 次
- SCOPE: Sign Language Contextual Processing with Embedding from LLMsYuqi Liu, Wenqian Zhang, Sihan Ren, Chengyu Huang 等AAAI 2025 · 被引用 7 次
它引用的顶会 Paper2
- Spatial-Temporal Multi-Cue Network for Continuous Sign Language RecognitionHao Zhou, Wengang Zhou, Yun Zhou, Houqiang LiAAAI 2020 · 被引用 249 次
- Sign Language Transformers: Joint End-to-End Sign Language Recognition and TranslationNecati Cihan Camgöz, Oscar Koller, Simon Hadfield, Richard BowdenCVPR 2020
相关 Paper
- SLTUNET: A Simple Unified Model for Sign Language TranslationBiao Zhang, Mathias Müller, Rico SennrichICLR 2023 · 被引用 14 次
- Gloss Matters: Unlocking the Potential of Non-Autoregressive Sign Language TranslationZhihao Wang, Shiyu Liu, Zhiwei He, Kangjie Zheng 等ACM MM 2025
- SimulSLT: End-to-End Simultaneous Sign Language TranslationAoxiong Yin, Zhou Zhao, Jinglin Liu, Weike Jin 等ACM MM 2021 · 被引用 35 次
- Sign2GPT: Leveraging Large Language Models for Gloss-Free Sign Language TranslationRyan Wong, Necati Cihan Camgöz, Richard BowdenICLR 2024 · 被引用 58 次
- C2ST: Cross-modal Contextualized Sequence Transduction for Continuous Sign Language RecognitionHuaiwen Zhang, Zihang Guo, Yang Yang, Xin Liu 等ICCV 2023 · 被引用 21 次
