mmPencil: Toward Writing-Style-Independent In-Air Handwriting Recognition via mmWave Radar and Large Vision-Language Model
Yifan Guo, Zhu Wang, Qian Qin, Yangqian Lei, Qiwen Gan, Zhuo Sun, Chao Chen, Bin Guo, Zhiwen Yu
摘要
With the rapid advancement of wireless technologies, in-air handwriting recognition based on radio frequency (RF) signals emerges as a promising solution for human-computer interaction. However, existing methods often constrain writing gestures to predefined planes, impose strict requirements on writing order and direction, and limit recognition tasks to a small set of digits or letters. To overcome these limitations, we propose m2VLMs, a modality mapping architecture, which serves as a bridge between the millimeter-wave (mmWave) radar sensing technologies and large vision-language models (VLMs). Based on the proposed architecture, a three-dimensional (3D) in-air handwriting word recognition system named mmPencil is developed. Specifically, we design a multi-stage spatial trajectory reconstruction algorithm that extracts frequency-domain features to identify handwriting regions and achieve high-precision reconstruction of 3D word trajectories. Furthermore, we introduce a novel spatial-to-visual mapping algorithm, which bridges the gap between spatial trajectory information captured by mmWave radar and vision-language representations, providing a foundation for cross-modality understanding in the large vision-language model. As a result, mmPencil tackles the limitations of current solutions that are heavily reliant on handwriting styles and environmental factors, expanding RF-based in-air handwriting recognition to more complex word-level scenarios. We collect and release a 3D mmWave handwriting dataset comprising 200 distinct words (ranging from 2 to 9 letters), contributions from 12 users, and 22 different writing scenarios, totaling 7,664 samples with an overall size of 31.66 GB. Extensive experiments demonstrate that mmPencil achieves accurate and robust word recognition in real-world scenarios, remaining unaffected by variations in word categories, length, as well as writing position, range, angle, speed, size, direction, and user movement. Specifically, for 4 seen users, mmPencil achieves a recognition accuracy of approximately 97.60% across 200 word classes. Moreover, benefiting from the generalization capability of VLMs, the zero-shot recognition accuracy for 4 unseen users can also reach 92.50% across 50 word classes using training data from only 8 users, which outperforming state-of-the-art baselines.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Foundation Models Defining A New Era In Sensor-based Human Activity Recognition: A Survey And OutlookSizhen Bian, Mengxi Liu, Lala Shakti Swarup Ray, Bo Zhou 等UbiComp 2026 · 被引用 2 次
- XGait: A Multi-Modality Wireless Sensing Dataset for Indoor Human Tracking and IdentificationWei Xu, Zhu Wang, Yifan Guo, Changlong Cheng 等UbiComp 2026
相关 Paper
- IndexPen: Two-Finger Text Input with Millimeter-Wave RadarHaowen Wei, Ziheng Li, Alexander D. Galvan, Zhuoran Su 等UbiComp 2022 · 被引用 26 次
- Pantomime: Mid-Air Gesture Recognition with Sparse Millimeter-Wave Radar Point CloudsSameera Palipana, Dariush Salami, Luis A. Leiva, Stephan SiggUbiComp 2021 · 被引用 169 次
- Real-time Arm Gesture Recognition in Smart Home Scenarios via Millimeter Wave SensingHaipeng Liu, Yuheng Wang, Anfu Zhou, Hanyue He 等UbiComp 2021 · 被引用 149 次
- m3ASL: ASL Gesture Recognition with Moving mmWave RadarGuiyun Fan, Rong Ding, Xiaocheng Wang, Yichen Zhu 等INFOCOM 2025 · 被引用 3 次
- RF-CM: Cross-Modal Framework for RF-enabled Few-Shot Human Activity RecognitionXuan Wang, Tong Liu, Chao Feng, Dingyi Fang 等UbiComp 2023 · 被引用 18 次
