OpenMAE: Efficient Masked Autoencoder for Vibration Sensing with Open-domain Data Enrichment
Chenzhi Hu, Yatong Chen, Denizhan Kara, Shengzhong Liu, Tarek F. Abdelzaher, Fan Wu, Guihai Chen
Abstract
This paper introduces OpenMAE, a novel data enrichment framework utilizing open-world sensor data streams to facilitate efficient masked autoencoder (MAE) pretraining on vibration signals. Due to highly sparse event occurrences and inevitable distributional shifts from downstream tasks, directly concatenating large-scale open-domain data with limited in-domain data during pretraining leads to degraded downstream task performance. The problem is further complicated by missing knowledge of open-world sensor environments and associated physical event semantics. Against these challenges, OpenMAE makes the following contributions to vibration MAE pretraining with open-domain data: First, it automatically filters out uninformative samples based on the event activeness and information consistency without relying on human annotations; Second, to mind the gap between open-domain and in-domain distributions, OpenMAE develops a novel data mixing method, FreqCutMix, that combines two data types in the frequency domain as augmented pretraining samples, preserving both events-of-interest semantics from in-domain data and real-world diversity from open-domain data. The open-domain data scale in data mixing is dynamically increased as pretraining progresses to stabilize the model convergence. We download over 5 million open-world vibration samples from the Raspberry Shake datacenter1 and conduct extensive experiments with two applications (i.e., indoor activity and outdoor transportation analysis). The evaluation results show OpenMAE improves downstream task accuracies by up to 23% and achieves enhanced generalizability into diverse downstream tasks, domain variations, and sensor-to-target distances.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- DAP-MAE: Domain-Adaptive Point Cloud Masked Autoencoder for Effective Cross-Domain LearningZiqi Gao, Qiufu Li, Linlin ShenICCV 2025 · 2 citations
- FreqMAE: Frequency-Aware Masked Autoencoder for Multi-Modal IoT SensingDenizhan Kara, Tomoyoshi Kimura, Shengzhong Liu, Jinyang Li et al.WWW 2024 · 27 citations
- UP2ME: Univariate Pre-training to Multivariate Fine-tuning as a General-purpose Framework for Multivariate Time Series AnalysisYunhao Zhang, Minghao Liu, Shengyang Zhou, Junchi YanICML 2024 · 9 citations
- Task-customized Masked Autoencoder via Mixture of Cluster-conditional ExpertsZhili Liu, Kai Chen, Jianhua Han, Lanqing Hong et al.ICLR 2023 · 6 citations
- Mixed Autoencoder for Self-Supervised Visual Representation LearningKai Chen, Zhili Liu, Lanqing Hong, Hang Xu et al.CVPR 2023
