Has the Deep Neural Network learned the Stochastic Process? An Evaluation Viewpoint
Harshit Kumar, Beomseok Kang, Biswadeep Chakraborty, Saibal Mukhopadhyay
摘要
This paper presents the first systematic study of evaluating Deep Neural Networks (DNNs) designed to forecast the evolution of stochastic complex systems. We show that traditional evaluation methods like threshold-based classification metrics and error-based scoring rules assess a DNN's ability to replicate the observed ground truth but fail to measure the DNN's learning of the underlying stochastic process. To address this gap, we propose a new evaluation criterion called Fidelity to Stochastic Process (F2SP), representing the DNN's ability to predict the system property Statistic-GT-the ground truth of the stochastic process-and introduce an evaluation metric that exclusively assesses F2SP. We formalize F2SP within a stochastic framework and establish criteria for validly measuring it. We formally show that Expected Calibration Error (ECE) satisfies the necessary condition for testing F2SP, unlike traditional evaluation methods. Empirical experiments on synthetic datasets, including wildfire, host-pathogen, and stock market models, demonstrate that ECE uniquely captures F2SP. We further extend our study to real-world wildfire data, highlighting the limitations of conventional evaluation and discuss the practical utility of incorporating F2SP into model assessment. This work offers a new perspective on evaluating DNNs modeling complex systems by emphasizing the importance of capturing the underlying stochastic process 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- SimVP: Simpler yet Better Video PredictionZhangyang Gao, Cheng Tan, Lirong Wu, Stan Z. LiCVPR 2022 · 被引用 313 次
- Better Uncertainty Calibration via Proper Scores for Classification and BeyondSebastian G. Gruber, Florian BuettnerNeurIPS 2022 · 被引用 88 次
- Regions of Reliability in the Evaluation of Multivariate Probabilistic ForecastsÉtienne Marcotte, Valentina Zantedeschi, Alexandre Drouin, Nicolas ChapadosICML 2023 · 被引用 10 次
- Towards Improving the Trustworthiness of Hardware based Malware Detector using Online Uncertainty EstimationHarshit Kumar, Nikhil Chawla, Saibal MukhopadhyayDAC 2021 · 被引用 6 次
相关 Paper
- Feature Likelihood Score: Evaluating the Generalization of Generative Models Using SamplesMarco Jiralerspong, Avishek Joey Bose, Ian Gemp, Chongli Qin 等NeurIPS 2023 · 被引用 39 次
- When in Doubt: Neural Non-Parametric Uncertainty Quantification for Epidemic ForecastingHarshavardhan Kamarthi, Lingkai Kong, Alexander Rodríguez, Chao Zhang 等NeurIPS 2021 · 被引用 26 次
- Evaluating Deep Neural Networks in Deployment: A Comparative Study (Replicability Study)Eduard Pinconschi, Divya Gopinath, Rui Abreu, Corina S. PasareanuISSTA 2024
- SNN-PDE: Learning Dynamic PDEs from Data with Simplicial Neural NetworksJae Choi, Yuzhou Chen, Huikyo Lee, Hyun Kim 等AAAI 2024 · 被引用 3 次
- Bayesian Oracle for bounding information gain in neural encoding modelsKonstantin-Klemens Lurz, Mohammad Bashiri, Edgar Y. Walker, Fabian H. SinzICLR 2023
