Weakly-supervised Disentanglement Network for Video Fingerspelling Detection
Ziqi Jiang, Shengyu Zhang, Siyuan Yao, Wenqiao Zhang, Sihan Zhang, Juncheng Li, Zhou Zhao, Fei Wu
摘要
Fingerspelling detection, which aims to localize and recognize fingerspelling gestures in raw, untrimmed videos, is a nascent but important research area that could help bridge the communication gap between deaf people and others. Many existing works tend to exploit additional knowledge, such as pose annotations, and newly datasets for performance improvement. However, in real-world applications, additional data collection and annotation require tremendous human efforts that are not always affordable. In this paper, we propose the Weakly-supervised Disentanglement Network, namely WED, that requires no additional knowledge, and better exploits the video-sentence weak supervisions. Specifically, WED incorporates two critical components: 1) Masked Disentanglement Module, which employs a Variational Autoencoder for signed letters disentanglement. Each latent factor in the VAE corresponds to a particular signed letter, and we mask latent factors corresponding to letters that do not appear in the video during decoding. Compared to the vanilla VAE, the masked reconstruction leverages the video-sentence weak supervision, leading to a better sign language oriented disentanglement; and 2) the Dynamic Memory Network module, which leverages the disentangled sign knowledge as prior knowledge and reference for sign-related frame identification and gesture recognition through a carefully designed memory reading component. We conduct extensive experiments on the benchmark ChicagoFSWild and ChicagoFSWild+ datasets. Empirical studies validate that the WED network achieves effective sign gesture disentanglement, contributing to the state-of-the-art performance for fingerspelling detection and recognition.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Label-Efficient Domain Generalization via Collaborative Exploration and GeneralizationJunkun Yuan, Xu Ma, Defang Chen, Kun Kuang 等ACM MM 2022 · 被引用 21 次
- Dilated Context Integrated Network with Cross-Modal Consensus for Temporal Emotion Localization in VideosJuncheng Li, Junlin Xie, Linchao Zhu, Long Qian 等ACM MM 2022 · 被引用 8 次
相关 Paper
- Fingerspelling Recognition in the Wild with Fixed-Query based Visual AttentionSrinivas Kruthiventi S. S, George Jose, Nitya Tandon, Rajesh Roshan Biswal 等ACM MM 2021 · 被引用 5 次
- Searching for fingerspelled content in American Sign LanguageBowen Shi, Diane Brentari, Greg Shakhnarovich, Karen LivescuACL 2022 · 被引用 8 次
- Fingerspelling Recognition in the Wild With Iterative Visual AttentionBowen Shi, Aurora Martinez Del Rio, Jonathan Keane, Diane Brentari 等ICCV 2019 · 被引用 76 次
- Hand-Model-Aware Sign Language RecognitionHezhen Hu, Wengang Zhou, Houqiang LiAAAI 2021 · 被引用 79 次
- Signing at Scale: Learning to Co-Articulate Signs for Large-Scale Photo-Realistic Sign Language ProductionBen Saunders, Necati Cihan Camgöz, Richard BowdenCVPR 2022 · 被引用 64 次
