Fingerspelling Recognition in the Wild with Fixed-Query based Visual Attention
Srinivas Kruthiventi S. S, George Jose, Nitya Tandon, Rajesh Roshan Biswal, Aashish Kumar
摘要
We propose an end-to-end solution for recognizing fingerspelling using multi-scale attention with fixed-queries. Fingerspelling recognition in the wild gets challenging because of the multiple sub-problems involved - detecting the signing hand, tracking it across frames, and recognizing subtle variations in a hand gesture. While the current state-of-the-art handles these with external face/hand detectors, optical flow features, and iteratively refining the attention maps, our work proposes a deep learning model that takes in the RGB videos and recognizes fingerspelling with a single forward pass. Without any frame-level supervision, our proposed model learns to pay attention to informative regions in each frame, such as fingers, hand, and face, to recognize signs. Multi-scale features from these attended regions are then processed using a recurrent neural network to recognize the alphabet sequentially. We train our model using a curriculum learning strategy with simpler samples at the beginning, followed by challenging samples at a later stage. We have evaluated our approach on Chicago Fingerspelling Wild and WildPlus datasets and have achieved about 8% and 4% improvements, respectively, compared to the current state-of-the-art methods. Further analysis of our method shows that our attention mechanism is intuitive from a human perspective, and visualizing it offers useful insights into the working of the model.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Fingerspelling Recognition in the Wild With Iterative Visual AttentionBowen Shi, Aurora Martinez Del Rio, Jonathan Keane, Diane Brentari 等ICCV 2019 · 被引用 76 次
- Weakly-supervised Disentanglement Network for Video Fingerspelling DetectionZiqi Jiang, Shengyu Zhang, Siyuan Yao, Wenqiao Zhang 等ACM MM 2022 · 被引用 5 次
- Searching for fingerspelled content in American Sign LanguageBowen Shi, Diane Brentari, Greg Shakhnarovich, Karen LivescuACL 2022 · 被引用 8 次
- OpenFS: Multi-Hand-Capable Fingerspelling Recognition with Implicit Signing-Hand Detection and Frame-Wise Letter-Conditioned SynthesisJunuk Cha, Jihyeon Kim, Han-Mu ParkCVPR 2026
- HyperSign: Saliency-Aware Spatial Graphs and Temporal Hypergraphs for Continuous Sign Language RecognitionWeiyi Ye, Xu-Hua Yang, Dong Wei, Gang-Feng Ma 等AAAI 2026
