A Path to Simpler Models Starts With Noise
Lesia Semenova, Harry Chen, Ronald Parr, Cynthia Rudin
Abstract
The Rashomon set is the set of models that perform approximately equally well on a given dataset, and the Rashomon ratio is the fraction of all models in a given hypothesis space that are in the Rashomon set. Rashomon ratios are often large for tabular datasets in criminal justice, healthcare, lending, education, and in other areas, which has practical implications about whether simpler models can attain the same level of accuracy as more complex models. An open question is why Rashomon ratios often tend to be large. In this work, we propose and study a mechanism of the data generation process, coupled with choices usually made by the analyst during the learning process, that determines the size of the Rashomon ratio. Specifically, we demonstrate that noisier datasets lead to larger Rashomon ratios through the way that practitioners train models. Additionally, we introduce a measure called pattern diversity, which captures the average difference in predictions between distinct classification patterns in the Rashomon set, and motivate why it tends to increase with label noise. Our results explain a key aspect of why simpler models often tend to perform as well as black box models on complex, noisier datasets. | F | , where | • | denotes cardinality. It is possible to weight the hypothesis space by a prior to define a weighted Rashomon ratio if desired.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 735b8dcd-8885-4d2d-a037-56e7d522f9b0Cited by top-tier papers8
- Using Noise to Infer Aspects of Simplicity Without LearningZachery Boner, Harry Chen, Lesia Semenova, Ronald Parr et al.NeurIPS 2024 · 10 citations
- ElliCE: Efficient and Provably Robust Algorithmic Recourse via the Rashomon SetsBohdan Turbal, Iryna Voitsitska, Lesia SemenovaNeurIPS 2025 · 6 citations
- When and Where to Reset Matters for Long-Term Test-Time AdaptationTaejun Lim, Joong-Won Hwang, Kibok LeeICLR 2026 · 3 citations
- DIVERSE: Disagreement-Inducing Vector Evolution for Rashomon Set ExplorationGilles Eerlings, Brent Zoomers, Jori Liesenborgs, Gustavo Alberto Rovelo Ruiz et al.ICLR 2026 · 2 citations
- The Double-Edged Nature of the Rashomon Set for Trustworthy Machine LearningEthan Hsu, Harry Chen, Chudi Zhong, Lesia SemenovaICML 2026 · 1 citation
Builds on15
- Predictive Multiplicity in ClassificationCharles T. Marx, Flávio P. Calmon, Berk UstunICML 2020 · 197 citations
- Generalized and Scalable Optimal Sparse Decision TreesJimmy Lin, Chudi Zhong, Diane Hu, Cynthia Rudin et al.ICML 2020 · 174 citations
- Label Noise SGD Provably Prefers Flat Global MinimizersAlex Damian, Tengyu Ma, Jason D. LeeNeurIPS 2021 · 155 citations
- Simple and Effective Regularization Methods for Training on Noisily Labeled Data with Generalization GuaranteeWei Hu, Zhiyuan Li, Dingli YuICLR 2020 · 140 citations
- Exploring the Whole Rashomon Set of Sparse Decision TreesRui Xin, Chudi Zhong, Zhi Chen, Takuya Takagi et al.NeurIPS 2022 · 117 citations
Related papers
- RashomonGB: Analyzing the Rashomon Effect and Mitigating Predictive Multiplicity in Gradient BoostingHsiang Hsu, Ivan Brugere, Shubham Sharma, Freddy Lécué et al.NeurIPS 2024 · 13 citations
- Rashomon Capacity: A Metric for Predictive Multiplicity in ClassificationHsiang Hsu, Flávio P. CalmonNeurIPS 2022 · 65 citations
- From Rashomon Theory to PRAXIS: Efficient Decision Tree Rashomon SetsZakk Heile, Hayden McTavish, Varun Babbar, Margo Seltzer et al.ICML 2026 · 1 citation
- Efficient Exploration of the Rashomon Set of Rule-Set ModelsMartino Ciaperoni, Han Xiao, Aristides GionisKDD 2024 · 3 citations
- Exploring the cloud of feature interaction scores in a Rashomon setSichao Li, Rong Wang, Quanling Deng, Amanda S. BarnardICLR 2024 · 9 citations
