Missing Value Imputation on Multidimensional Time Series
Parikshit Bansal, Prathamesh Deshpande, Sunita Sarawagi
摘要
We present DeepMVI, a deep learning method for missing value imputation in multidimensional time-series datasets. Missing values are commonplace in decision support platforms that aggregate data over long time stretches from disparate sources, whereas reliable data analytics calls for careful handling of missing data. One strategy is imputing the missing values, and a wide variety of algorithms exist spanning simple interpolation, matrix factorization methods like SVD, statistical models like Kalman filters, and recent deep learning methods. We show that often these provide worse results on aggregate analytics compared to just excluding the missing data. DeepMVI expresses the distribution of each missing value conditioned on coarse and fine-grained signals along a time series, and signals from correlated series at the same time. Instead of resorting to linearity assumptions of conventional matrix factorization methods, DeepMVI harnesses a flexible deep network to extract and combine these signals in an end-to-end manner. To prevent over-fitting with high-capacity neural networks, we design a robust parameter training with labeled data created using synthetic missing blocks around available indices. Our neural network uses a modular design with a novel temporal transformer with convolutional features, and kernel regression with learned embeddings. Experiments across ten real datasets, five different missing scenarios, comparing seven conventional and three deep learning methods show that DeepMVI is significantly more accurate, reducing error by more than 50% in more than half the cases, compared to the best existing method. Although slower than simpler matrix factorization methods, we justify the increased time overheads by showing that DeepMVI provides significantly more accurate imputation that finally impacts quality of downstream analytics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- A Multi-Scale Decomposition MLP-Mixer for Time Series AnalysisShuhan Zhong, Sizhe Song, Weipeng Zhuo, Guanyao Li 等VLDB 2024 · 被引用 48 次
- Conditional Information Bottleneck Approach for Time Series ImputationMinGyu Choi, Changhee LeeICLR 2024 · 被引用 24 次
- Data Imputation for Sparse Radio Maps in Indoor PositioningXiao Li, Huan Li, Harry Kai-Ho Chan, Hua Lu 等ICDE 2023 · 被引用 11 次
- Scaling Up Multivariate Time Series Pre-Training with Decoupled Spatial-Temporal RepresentationsRui Zha, Le Zhang, Shuangli Li, Jingbo Zhou 等ICDE 2024 · 被引用 9 次
- Mining of Switching Sparse Networks for Missing Value Imputation in Multivariate Time SeriesKohei Obata, Koki Kawabata, Yasuko Matsubara, Yasushi SakuraiKDD 2024 · 被引用 6 次
它引用的顶会 Paper1
相关 Paper
- Probabilistic Imputation for Time-series Classification with Missing DataSeunghyun Kim, Hyunsu Kim, Eunggu Yun, Hwangrae Lee 等ICML 2023 · 被引用 37 次
- An Observed Value Consistent Diffusion Model for Imputing Missing Values in Multivariate Time SeriesXu Wang, Hongbo Zhang, Pengkun Wang, Yudong Zhang 等KDD 2023 · 被引用 43 次
- Factorized Inference in Deep Markov Models for Incomplete Multimodal Time SeriesZhi-Xuan Tan, Harold Soh, Desmond C. OngAAAI 2020 · 被引用 32 次
- Time-Gated Multi-Scale Flow Matching for Time-Series ImputationHangtian Wang, Mahito SugiyamaICLR 2026
- A Transformer-based Framework for Multivariate Time Series Representation LearningGeorge Zerveas, Srideepika Jayaraman, Dhaval Patel, Anuradha Bhamidipaty 等KDD 2021 · 被引用 66 次
