HiMAE: Hierarchical Masked Autoencoders Discover Resolution-Specific Structure in Wearable Time Series
Simon A. Lee, Cyrus Tanade, Hao Zhou, Juhyeon Lee, Megha Thukral, Md Sazzad Hissain Khan, Keum San Chun, Baiying Lu, Migyeong Gwak, Mehrab Bin Morshed, Viswam Nathan, Md. Mahbubur Rahman
Abstract
Wearable sensors provide abundant physiological time series observations, yet the resolution at which we should extract features for downstream tasks remain unclear. We hypothesize that temporal resolution is a fundamental axis of representation learning, with different clinical and behavioral outcomes relying on features at distinct scales. To test this resolution hypothesis, we introduce HiMAE (Hierarchical Masked Autoencoder), a self-supervised framework that combines masked autoencoding with a hierarchical convolutional encoder–decoder. HiMAE produces multi-resolution embeddings across its intermediate layers that enable systematic evaluation of which temporal scales carry predictive signal, transforming resolution from a hyperparameter into a probe for interpretability. Across classification and generative benchmarks, HiMAE consistently outperforms state-of-the-art foundation models that collapse scale, while being orders of magnitude smaller. Due to the convolution based design choices behind HiMAE, the model is also compact enough to run entirely on-device, achieving sub-millisecond inference on smartwatch-class CPUs for true edge inference. Together, these contributions position HiMAE as both an efficient self supervised learning method and a discovery tool for understanding how time resolution contributes to downstream task alignment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Pyraformer: Low-Complexity Pyramidal Attention for Long-Range Time Series Modeling and ForecastingShizhan Liu, Hang Yu, Cong Liao, Jianguo Li et al.ICLR 2022 · 975 citations
Related papers
- Spatial-Temporal Masked Autoencoder for Multi-Device Wearable Human Activity RecognitionShenghuan Miao, Ling Chen, Rong HuUbiComp 2024 · 23 citations
- Understanding Masked Autoencoders via Hierarchical Latent Variable ModelsLingjing Kong, Martin Q. Ma, Guangyi Chen, Eric P. Xing et al.CVPR 2023
- Physiology-Aware Masked Cross-Modal Reconstruction for Biosignal Representation LearningHao Zhou, Simon Lee, Cyrus Tanade, Keum San Chun et al.ICML 2026 · 3 citations
- Hi-GMAE: Hierarchical Graph Masked AutoencodersChuang Liu, Zelin Yao, Xueqi Ma, Mukun Chen et al.WWW 2026 · 3 citations
- Stochastic Optimal Control for Continuous-Time fMRI Representation LearningJoonhyeong Park, Byoungwoo Park, Chang-Bae Bang, Jungwon Choi et al.ICLR 2026 · 2 citations
