SIGMA-ASL: Sensor-Integrated Multimodal Dataset for Sign Language Recognition
Xiaofang Xiao, Guangchao Li, Guangrong Zhao, Qi Lin, Wen Ma, Hongkai Wen, Yanxiang Wang, Yiran Shen
Abstract
Automatic sign language recognition (SLR) has become a key enabler of inclusive human-computer interaction, fostering seamless communication between deaf individuals and hearing communities. Despite significant advances in multimodal learning, existing SLR research remains dominated by vision-based datasets, which are limited by sensitivity to lighting and occlusion, privacy concerns, and a lack of cross-modal diversity. To address these challenges, we introduce SIGMA-ASL, a large-scale multimodal dataset for SLR. The dataset integrates an Azure Kinect RGB-D camera, a millimeter-wave (mmWave) radar, and two wrist-worn inertial measurement units (IMUs) to capture complementary visual, radio-reflection, and kinematic information. Collected in a controlled studio environment with 20 participants performing 160 common American sign language (ASL) signs, SIGMA-ASL provides 93,545 temporally synchronized word-level multimodal clips. A unified sensing framework achieves millisecond-level alignment across modalities, enabling reliable sensor fusion and cross-modal learning. We further design standardized preprocessing pipelines and benchmarking protocols under both user-dependent and userindependent settings, offering a comprehensive foundation for evaluating single and multimodal SLR. Extensive experiments validate the dataset's quality and demonstrate its potential as a valuable resource for developing robust, privacy-preserving, and ubiquitous sign language recognition systems.
CCS Concepts: • Human-centered computing → Ubiquitous and mobile computing systems and tools.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf160d46-755f-4833-9e6e-ab18f54c7f0eBuilds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
Related papers
- INCLUDE: A Large Scale Dataset for Indian Sign Language RecognitionAdvaith Sridhar, Rohith Gandhi Ganesan, Pratyush Kumar, Mitesh M. KhapraACM MM 2020 · 144 citations
- m3ASL: ASL Gesture Recognition with Moving mmWave RadarGuiyun Fan, Rong Ding, Xiaocheng Wang, Yichen Zhu et al.INFOCOM 2025 · 3 citations
- M4Human: A Large-Scale Multimodal mmWave Radar Benchmark for Human Mesh ReconstructionJunqiao Fan, Yunjiao Zhou, Yizhuo Yang, Xinyuan Cui et al.CVPR 2026 · 11 citations
- mmASL: Environment-Independent ASL Gesture Recognition Using 60 GHz Millimeter-wave SignalsPanneer Selvam Santhalingam, Al Amin Hosain, Ding Zhang, Parth H. Pathak et al.UbiComp 2020 · 89 citations
- RF-CM: Cross-Modal Framework for RF-enabled Few-Shot Human Activity RecognitionXuan Wang, Tong Liu, Chao Feng, Dingyi Fang et al.UbiComp 2023 · 18 citations
