Has the Deep Neural Network learned the Stochastic Process? An Evaluation Viewpoint
Harshit Kumar, Beomseok Kang, Biswadeep Chakraborty, Saibal Mukhopadhyay
Abstract
This paper presents the first systematic study of evaluating Deep Neural Networks (DNNs) designed to forecast the evolution of stochastic complex systems. We show that traditional evaluation methods like threshold-based classification metrics and error-based scoring rules assess a DNN's ability to replicate the observed ground truth but fail to measure the DNN's learning of the underlying stochastic process. To address this gap, we propose a new evaluation criterion called Fidelity to Stochastic Process (F2SP), representing the DNN's ability to predict the system property Statistic-GT-the ground truth of the stochastic process-and introduce an evaluation metric that exclusively assesses F2SP. We formalize F2SP within a stochastic framework and establish criteria for validly measuring it. We formally show that Expected Calibration Error (ECE) satisfies the necessary condition for testing F2SP, unlike traditional evaluation methods. Empirical experiments on synthetic datasets, including wildfire, host-pathogen, and stock market models, demonstrate that ECE uniquely captures F2SP. We further extend our study to real-world wildfire data, highlighting the limitations of conventional evaluation and discuss the practical utility of incorporating F2SP into model assessment. This work offers a new perspective on evaluating DNNs modeling complex systems by emphasizing the importance of capturing the underlying stochastic process 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bcdebaf2-e963-4d53-81a4-1714dde7a37aBuilds on6
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- SimVP: Simpler yet Better Video PredictionZhangyang Gao, Cheng Tan, Lirong Wu, Stan Z. LiCVPR 2022 · 313 citations
- Better Uncertainty Calibration via Proper Scores for Classification and BeyondSebastian G. Gruber, Florian BuettnerNeurIPS 2022 · 88 citations
- Regions of Reliability in the Evaluation of Multivariate Probabilistic ForecastsÉtienne Marcotte, Valentina Zantedeschi, Alexandre Drouin, Nicolas ChapadosICML 2023 · 10 citations
- Towards Improving the Trustworthiness of Hardware based Malware Detector using Online Uncertainty EstimationHarshit Kumar, Nikhil Chawla, Saibal MukhopadhyayDAC 2021 · 6 citations
Related papers
- Feature Likelihood Score: Evaluating the Generalization of Generative Models Using SamplesMarco Jiralerspong, Avishek Joey Bose, Ian Gemp, Chongli Qin et al.NeurIPS 2023 · 39 citations
- When in Doubt: Neural Non-Parametric Uncertainty Quantification for Epidemic ForecastingHarshavardhan Kamarthi, Lingkai Kong, Alexander Rodríguez, Chao Zhang et al.NeurIPS 2021 · 26 citations
- Evaluating Deep Neural Networks in Deployment: A Comparative Study (Replicability Study)Eduard Pinconschi, Divya Gopinath, Rui Abreu, Corina S. PasareanuISSTA 2024
- SNN-PDE: Learning Dynamic PDEs from Data with Simplicial Neural NetworksJae Choi, Yuzhou Chen, Huikyo Lee, Hyun Kim et al.AAAI 2024 · 3 citations
- Bayesian Oracle for bounding information gain in neural encoding modelsKonstantin-Klemens Lurz, Mohammad Bashiri, Edgar Y. Walker, Fabian H. SinzICLR 2023
