Fine-grained Control of Generative Data Augmentation in IoT Sensing
Tianshi Wang, Qikai Yang, Ruijie Wang, Dachun Sun, Jinyang Li, Yizhuo Chen, Yigong Hu, Chaoqi Yang, Tomoyoshi Kimura, Denizhan Kara, Tarek F. Abdelzaher
Abstract
Internet of Things (IoT) sensing models often suffer from overfitting due to data distribution shifts between training dataset and real-world scenarios. To address this, data augmentation techniques have been adopted to enhance model robustness by bolstering the diversity of synthetic samples within a defined vicinity of existing samples. This paper introduces a novel paradigm of data augmentation for IoT sensing signals by adding fine-grained control to generative models. We define a metric space with statistical metrics that capture the essential features of the short-time Fourier transformed (STFT) spectrograms of IoT sensing signals. These metrics serve as strong conditions for a generative model, enabling us to tailor the spectrogram characteristics in the time-frequency domain according to specific application needs. Furthermore, we propose a set of data augmentation techniques within this metric space to create new data samples. Our method is evaluated across various generative models, datasets, and downstream IoT sensing models. The results demonstrate that our approach surpasses the conventional transformation-based data augmentation techniques and prior generative data augmentation models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 03baf08c-27cd-4daf-bf22-b141a40f93b2Cited by top-tier papers2
- In-Context Compositional Q-Learning for Offline Reinforcement LearningQiushui Xu, Yuhao Huang, Yushu Jiang, Wenliang Zheng et al.ICLR 2026
- Unveiling Markov heads in Pretrained Language Models for Offline Reinforcement LearningWenhao Zhao, Qiushui Xu, Linjie Xu, Lei Song et al.ICML 2025
Builds on11
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao et al.ICLR 2021 · 1,902 citations
- NVAE: A Deep Hierarchical Variational AutoencoderArash Vahdat, Jan KautzNeurIPS 2020 · 1,141 citations
- DiffSinger: Singing Voice Synthesis via Shallow Diffusion MechanismJinglin Liu, Chengxi Li, Yi Ren, Feiyang Chen et al.AAAI 2022 · 348 citations
- FOCAL: Contrastive Learning for Multimodal Time-Series Sensing Signals in Factorized Orthogonal Latent SpaceShengzhong Liu, Tomoyoshi Kimura, Dongxin Liu, Ruijie Wang et al.NeurIPS 2023 · 72 citations
- Toward Understanding Generative Data AugmentationChenyu Zheng, Guoqiang Wu, Chongxuan LiNeurIPS 2023 · 51 citations
Related papers
- Dywave: Event-Aligned Dynamic Tokenization for Heterogeneous IoT Sensing SignalsTomoyoshi Kimura, Denizhan Kara, Jinyang Li, Hongjue Zhao et al.ICML 2026
- DyC-STG: Dynamic Causal Spatio-Temporal Graph Network for Real-time Data Credibility Analysis in IoTGuanjie Cheng, Boyi Li, Peihan Wu, Feiyi Chen et al.AAAI 2026
- FreqMAE: Frequency-Aware Masked Autoencoder for Multi-Modal IoT SensingDenizhan Kara, Tomoyoshi Kimura, Shengzhong Liu, Jinyang Li et al.WWW 2024 · 27 citations
- Pseudo-Non-Linear Data Augmentation: A Constrained Energy Minimization ViewpointPingbang Hu, Mahito SugiyamaICLR 2026
- Federated Generative Model on Multi-Source Heterogeneous Data in IoTZuobin Xiong, Wei Li, Zhipeng CaiAAAI 2023 · 40 citations
