FSboard: Over 3 Million Characters of ASL Fingerspelling Collected via Smartphones
Manfred Georg, Garrett Tanzer, Esha Uboweja, Saad Hassan, Maximus Shengelia, Sam S. Sepah, Sean Forbes, Thad Starner
Abstract
Progress in machine understanding of sign languages has been slow and hampered by limited data. In this paper, we present FSboard, an American Sign Language fingerspelling dataset situated in a mobile text entry use case, collected from 147 paid and consenting Deaf signers using Pixel 4A selfie cameras in a variety of environments. Fingerspelling recognition is an incomplete solution that is only one small part of sign language translation, but it could provide some immediate benefit to Deaf/Hard of Hearing signers as more broadly capable technology develops. At >3 million characters in length and >250 hours in duration, FSboard is the largest fingerspelling recognition dataset to date by a factor of >10x. As a simple baseline, we finetune 30 Hz MediaPipe Holistic landmark inputs into ByT5-Small and achieve 11.1% Character Error Rate (CER) on a test set with unique phrases and signers. This quality degrades gracefully when decreasing frame rate and excluding face/body landmarks-plausible optimizations to help models run on device in real time. 2
- equal contribution † equal advising ‡ work conducted while at Google 2 We publicly release FSboard at this link under a CC BY 4.0 license. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4fdb12ab-2591-4be5-8aa2-1759a4f1f792Cited by top-tier papers4
- Sign-Language Datasets at Scale: A Comprehensive Survey on Resources, Benchmarks, and Annotation StandardsYiming Ni, Zhi-Qi Cheng, Jiayu Li, Wei ChengACL 2026 · 1 citation
- SHuBERT: Self-Supervised Sign Language Representation Learning via Multi-Stream Cluster PredictionShester Gueuwou, Xiaodan Du, Greg Shakhnarovich, Karen Livescu et al.ACL 2025
- OpenFS: Multi-Hand-Capable Fingerspelling Recognition with Implicit Signing-Hand Detection and Frame-Wise Letter-Conditioned SynthesisJunuk Cha, Jihyeon Kim, Han-Mu ParkCVPR 2026
- BANZ-FS: BANZSL Fingerspelling DatasetXin Shen, Yan Ke, Xinyu Wang, Xin YuICLR 2026
Builds on5
- Fingerspelling Recognition in the Wild With Iterative Visual AttentionBowen Shi, Aurora Martinez Del Rio, Jonathan Keane, Diane Brentari et al.ICCV 2019 · 76 citations
- Open-Domain Sign Language Translation Learned from Online VideoBowen Shi, Diane Brentari, Gregory Shakhnarovich, Karen LivescuEMNLP 2022 · 39 citations
- Reconsidering Sentence-Level Sign Language TranslationGarrett Tanzer, Maximus Shengelia, Ken Harrenstien, David UthusEMNLP 2024 · 4 citations
- Towards Privacy-Aware Sign Language Translation at ScalePhillip Rust, Bowen Shi, Skyler Wang, Necati Cihan Camgöz et al.ACL 2024
- How2Sign: A Large-Scale Multimodal Dataset for Continuous American Sign LanguageAmanda Cardoso Duarte, Shruti Palaskar, Lucas Ventura, Deepti Ghadiyaram et al.CVPR 2021
Related papers
- Searching for fingerspelled content in American Sign LanguageBowen Shi, Diane Brentari, Greg Shakhnarovich, Karen LivescuACL 2022 · 8 citations
- YouTube-SL-25: A Large-Scale, Open-Domain Multilingual Sign Language Parallel CorpusGarrett Tanzer, Biao ZhangICLR 2025
- WearSign: Pushing the Limit of Sign Language Translation Using Inertial and EMG WearablesQian Zhang, JiaZhen Jing, Dong Wang, Run ZhaoUbiComp 2022 · 26 citations
- SCOPE: Sign Language Contextual Processing with Embedding from LLMsYuqi Liu, Wenqian Zhang, Sihan Ren, Chengyu Huang et al.AAAI 2025 · 7 citations
- INCLUDE: A Large Scale Dataset for Indian Sign Language RecognitionAdvaith Sridhar, Rohith Gandhi Ganesan, Pratyush Kumar, Mitesh M. KhapraACM MM 2020 · 144 citations
