Cross-Validation for Longitudinal Datasets with Unstable Correlations
Meera Krishnamoorthy, Michael W. Sjoding, Jenna Wiens
Abstract
Cross-validation (CV) approaches are widely used in machine learning model development. They select hyperparameters associated with models with high average performance across multiple folds sampled from training data. Popular CV approaches do this sampling either randomly (random CV) or temporally (block CV). Averaging performance across multiple folds is intended to result in robust models; however, it can mask a model's reliance on unstable correlations - transient relationships between inputs and outputs that do not hold over time - which can make a model fail over time. In this paper, we show, both theoretically and empirically, that random and block CV often select hyperparameters corresponding to models that rely on unstable correlations over stable correlations. In light of this shortcoming, we propose a novel CV approach that is more likely to result in models that rely on stable correlations over unstable correlations, compared to random and block CV. Our proposed approach leverages the fact that when a model relies on features whose relationship with the outcome is stable, random and block CV result in the same estimated model performance; it does not matter how you split the training data into CV folds. However, when a model relies on features whose relationship with the outcome is unstable, random and block CV result in different estimates of model performance. We use this difference as a signal that a model relies on an unstable correlation and propose a model selection technique that selects hyperparameters that minimize this difference. When combined with standard model performance criteria, our approach results in models with higher average and more stable performance over time compared to random and block CV. Applied to the task of predicting 5-year survival among individuals with lung cancer, our approach leads to models that perform better in the long term compared to the next best approach (0.850, IQR: [0.824, 0.853] vs. 0.804, IQR: [0.753, 0.850]). In high-stakes dynamic domains like healthcare, where good performance over time is crucial, our approach can help result in models that perform consistently well despite unstable correlations in the training set. Code to implement our approach and reproduce all experiments in the paper is available at https://github.com/mld3/cv_unstable_correlations.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get e928ce9b-b5e2-4534-958a-e8e062d4daadRelated papers
- Approximate Cross-Validation for Structured ModelsSoumya Ghosh, William T. Stephenson, Tin D. Nguyen, Sameer K. Deshpande et al.NeurIPS 2020 · 19 citations
- The Relative Instability of Model Comparison with Cross-validationAlexandre Bayle, Lucas Janson, Lester MackeyICML 2026 · 1 citation
- Cross-Validated Off-Policy EvaluationMatej Cief, Branislav Kveton, Michal KompanAAAI 2025 · 2 citations
- HyperTime: Hyperparameter Optimization for Combating Temporal Distribution ShiftsShaokun Zhang, Yiran Wu, Zhonghua Zheng, Qingyun Wu et al.ACM MM 2024 · 2 citations
- Reshuffling Resampling Splits Can Improve Generalization of Hyperparameter OptimizationThomas Nagler, Lennart Schneider, Bernd Bischl, Matthias FeurerNeurIPS 2024 · 8 citations
