mmPencil: Toward Writing-Style-Independent In-Air Handwriting Recognition via mmWave Radar and Large Vision-Language Model
Yifan Guo, Zhu Wang, Qian Qin, Yangqian Lei, Qiwen Gan, Zhuo Sun, Chao Chen, Bin Guo, Zhiwen Yu
Abstract
With the rapid advancement of wireless technologies, in-air handwriting recognition based on radio frequency (RF) signals emerges as a promising solution for human-computer interaction. However, existing methods often constrain writing gestures to predefined planes, impose strict requirements on writing order and direction, and limit recognition tasks to a small set of digits or letters. To overcome these limitations, we propose m2VLMs, a modality mapping architecture, which serves as a bridge between the millimeter-wave (mmWave) radar sensing technologies and large vision-language models (VLMs). Based on the proposed architecture, a three-dimensional (3D) in-air handwriting word recognition system named mmPencil is developed. Specifically, we design a multi-stage spatial trajectory reconstruction algorithm that extracts frequency-domain features to identify handwriting regions and achieve high-precision reconstruction of 3D word trajectories. Furthermore, we introduce a novel spatial-to-visual mapping algorithm, which bridges the gap between spatial trajectory information captured by mmWave radar and vision-language representations, providing a foundation for cross-modality understanding in the large vision-language model. As a result, mmPencil tackles the limitations of current solutions that are heavily reliant on handwriting styles and environmental factors, expanding RF-based in-air handwriting recognition to more complex word-level scenarios. We collect and release a 3D mmWave handwriting dataset comprising 200 distinct words (ranging from 2 to 9 letters), contributions from 12 users, and 22 different writing scenarios, totaling 7,664 samples with an overall size of 31.66 GB. Extensive experiments demonstrate that mmPencil achieves accurate and robust word recognition in real-world scenarios, remaining unaffected by variations in word categories, length, as well as writing position, range, angle, speed, size, direction, and user movement. Specifically, for 4 seen users, mmPencil achieves a recognition accuracy of approximately 97.60% across 200 word classes. Moreover, benefiting from the generalization capability of VLMs, the zero-shot recognition accuracy for 4 unseen users can also reach 92.50% across 50 word classes using training data from only 8 users, which outperforming state-of-the-art baselines.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 9db53d00-0f58-4850-9643-9ea689fb0b04Cited by top-tier papers2
- Foundation Models Defining A New Era In Sensor-based Human Activity Recognition: A Survey And OutlookSizhen Bian, Mengxi Liu, Lala Shakti Swarup Ray, Bo Zhou et al.UbiComp 2026 · 2 citations
- XGait: A Multi-Modality Wireless Sensing Dataset for Indoor Human Tracking and IdentificationWei Xu, Zhu Wang, Yifan Guo, Changlong Cheng et al.UbiComp 2026
Related papers
- IndexPen: Two-Finger Text Input with Millimeter-Wave RadarHaowen Wei, Ziheng Li, Alexander D. Galvan, Zhuoran Su et al.UbiComp 2022 · 26 citations
- Pantomime: Mid-Air Gesture Recognition with Sparse Millimeter-Wave Radar Point CloudsSameera Palipana, Dariush Salami, Luis A. Leiva, Stephan SiggUbiComp 2021 · 169 citations
- Real-time Arm Gesture Recognition in Smart Home Scenarios via Millimeter Wave SensingHaipeng Liu, Yuheng Wang, Anfu Zhou, Hanyue He et al.UbiComp 2021 · 149 citations
- m3ASL: ASL Gesture Recognition with Moving mmWave RadarGuiyun Fan, Rong Ding, Xiaocheng Wang, Yichen Zhu et al.INFOCOM 2025 · 3 citations
- RF-CM: Cross-Modal Framework for RF-enabled Few-Shot Human Activity RecognitionXuan Wang, Tong Liu, Chao Feng, Dingyi Fang et al.UbiComp 2023 · 18 citations
