Missing Value Imputation on Multidimensional Time Series
Parikshit Bansal, Prathamesh Deshpande, Sunita Sarawagi
Abstract
We present DeepMVI, a deep learning method for missing value imputation in multidimensional time-series datasets. Missing values are commonplace in decision support platforms that aggregate data over long time stretches from disparate sources, whereas reliable data analytics calls for careful handling of missing data. One strategy is imputing the missing values, and a wide variety of algorithms exist spanning simple interpolation, matrix factorization methods like SVD, statistical models like Kalman filters, and recent deep learning methods. We show that often these provide worse results on aggregate analytics compared to just excluding the missing data. DeepMVI expresses the distribution of each missing value conditioned on coarse and fine-grained signals along a time series, and signals from correlated series at the same time. Instead of resorting to linearity assumptions of conventional matrix factorization methods, DeepMVI harnesses a flexible deep network to extract and combine these signals in an end-to-end manner. To prevent over-fitting with high-capacity neural networks, we design a robust parameter training with labeled data created using synthetic missing blocks around available indices. Our neural network uses a modular design with a novel temporal transformer with convolutional features, and kernel regression with learned embeddings. Experiments across ten real datasets, five different missing scenarios, comparing seven conventional and three deep learning methods show that DeepMVI is significantly more accurate, reducing error by more than 50% in more than half the cases, compared to the best existing method. Although slower than simpler matrix factorization methods, we justify the increased time overheads by showing that DeepMVI provides significantly more accurate imputation that finally impacts quality of downstream analytics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 899f3c3f-0877-43ed-8772-339286167c5cCited by top-tier papers15
- A Multi-Scale Decomposition MLP-Mixer for Time Series AnalysisShuhan Zhong, Sizhe Song, Weipeng Zhuo, Guanyao Li et al.VLDB 2024 · 48 citations
- Conditional Information Bottleneck Approach for Time Series ImputationMinGyu Choi, Changhee LeeICLR 2024 · 24 citations
- Data Imputation for Sparse Radio Maps in Indoor PositioningXiao Li, Huan Li, Harry Kai-Ho Chan, Hua Lu et al.ICDE 2023 · 11 citations
- Scaling Up Multivariate Time Series Pre-Training with Decoupled Spatial-Temporal RepresentationsRui Zha, Le Zhang, Shuangli Li, Jingbo Zhou et al.ICDE 2024 · 9 citations
- Mining of Switching Sparse Networks for Missing Value Imputation in Multivariate Time SeriesKohei Obata, Koki Kawabata, Yasuko Matsubara, Yasushi SakuraiKDD 2024 · 6 citations
Builds on1
Related papers
- Probabilistic Imputation for Time-series Classification with Missing DataSeunghyun Kim, Hyunsu Kim, Eunggu Yun, Hwangrae Lee et al.ICML 2023 · 37 citations
- An Observed Value Consistent Diffusion Model for Imputing Missing Values in Multivariate Time SeriesXu Wang, Hongbo Zhang, Pengkun Wang, Yudong Zhang et al.KDD 2023 · 43 citations
- Factorized Inference in Deep Markov Models for Incomplete Multimodal Time SeriesZhi-Xuan Tan, Harold Soh, Desmond C. OngAAAI 2020 · 32 citations
- Time-Gated Multi-Scale Flow Matching for Time-Series ImputationHangtian Wang, Mahito SugiyamaICLR 2026
- A Transformer-based Framework for Multivariate Time Series Representation LearningGeorge Zerveas, Srideepika Jayaraman, Dhaval Patel, Anuradha Bhamidipaty et al.KDD 2021 · 66 citations
