Weakly-supervised Disentanglement Network for Video Fingerspelling Detection
Ziqi Jiang, Shengyu Zhang, Siyuan Yao, Wenqiao Zhang, Sihan Zhang, Juncheng Li, Zhou Zhao, Fei Wu
Abstract
Fingerspelling detection, which aims to localize and recognize fingerspelling gestures in raw, untrimmed videos, is a nascent but important research area that could help bridge the communication gap between deaf people and others. Many existing works tend to exploit additional knowledge, such as pose annotations, and newly datasets for performance improvement. However, in real-world applications, additional data collection and annotation require tremendous human efforts that are not always affordable. In this paper, we propose the Weakly-supervised Disentanglement Network, namely WED, that requires no additional knowledge, and better exploits the video-sentence weak supervisions. Specifically, WED incorporates two critical components: 1) Masked Disentanglement Module, which employs a Variational Autoencoder for signed letters disentanglement. Each latent factor in the VAE corresponds to a particular signed letter, and we mask latent factors corresponding to letters that do not appear in the video during decoding. Compared to the vanilla VAE, the masked reconstruction leverages the video-sentence weak supervision, leading to a better sign language oriented disentanglement; and 2) the Dynamic Memory Network module, which leverages the disentangled sign knowledge as prior knowledge and reference for sign-related frame identification and gesture recognition through a carefully designed memory reading component. We conduct extensive experiments on the benchmark ChicagoFSWild and ChicagoFSWild+ datasets. Empirical studies validate that the WED network achieves effective sign gesture disentanglement, contributing to the state-of-the-art performance for fingerspelling detection and recognition.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers2
- Label-Efficient Domain Generalization via Collaborative Exploration and GeneralizationJunkun Yuan, Xu Ma, Defang Chen, Kun Kuang et al.ACM MM 2022 · 21 citations
- Dilated Context Integrated Network with Cross-Modal Consensus for Temporal Emotion Localization in VideosJuncheng Li, Junlin Xie, Linchao Zhu, Long Qian et al.ACM MM 2022 · 8 citations
Related papers
- Fingerspelling Recognition in the Wild with Fixed-Query based Visual AttentionSrinivas Kruthiventi S. S, George Jose, Nitya Tandon, Rajesh Roshan Biswal et al.ACM MM 2021 · 5 citations
- Searching for fingerspelled content in American Sign LanguageBowen Shi, Diane Brentari, Greg Shakhnarovich, Karen LivescuACL 2022 · 8 citations
- Fingerspelling Recognition in the Wild With Iterative Visual AttentionBowen Shi, Aurora Martinez Del Rio, Jonathan Keane, Diane Brentari et al.ICCV 2019 · 76 citations
- Hand-Model-Aware Sign Language RecognitionHezhen Hu, Wengang Zhou, Houqiang LiAAAI 2021 · 79 citations
- Signing at Scale: Learning to Co-Articulate Signs for Large-Scale Photo-Realistic Sign Language ProductionBen Saunders, Necati Cihan Camgöz, Richard BowdenCVPR 2022 · 64 citations
