A Path to Simpler Models Starts With Noise
Lesia Semenova, Harry Chen, Ronald Parr, Cynthia Rudin
摘要
The Rashomon set is the set of models that perform approximately equally well on a given dataset, and the Rashomon ratio is the fraction of all models in a given hypothesis space that are in the Rashomon set. Rashomon ratios are often large for tabular datasets in criminal justice, healthcare, lending, education, and in other areas, which has practical implications about whether simpler models can attain the same level of accuracy as more complex models. An open question is why Rashomon ratios often tend to be large. In this work, we propose and study a mechanism of the data generation process, coupled with choices usually made by the analyst during the learning process, that determines the size of the Rashomon ratio. Specifically, we demonstrate that noisier datasets lead to larger Rashomon ratios through the way that practitioners train models. Additionally, we introduce a measure called pattern diversity, which captures the average difference in predictions between distinct classification patterns in the Rashomon set, and motivate why it tends to increase with label noise. Our results explain a key aspect of why simpler models often tend to perform as well as black box models on complex, noisier datasets. | F | , where | • | denotes cardinality. It is possible to weight the hypothesis space by a prior to define a weighted Rashomon ratio if desired.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Using Noise to Infer Aspects of Simplicity Without LearningZachery Boner, Harry Chen, Lesia Semenova, Ronald Parr 等NeurIPS 2024 · 被引用 10 次
- ElliCE: Efficient and Provably Robust Algorithmic Recourse via the Rashomon SetsBohdan Turbal, Iryna Voitsitska, Lesia SemenovaNeurIPS 2025 · 被引用 6 次
- When and Where to Reset Matters for Long-Term Test-Time AdaptationTaejun Lim, Joong-Won Hwang, Kibok LeeICLR 2026 · 被引用 3 次
- DIVERSE: Disagreement-Inducing Vector Evolution for Rashomon Set ExplorationGilles Eerlings, Brent Zoomers, Jori Liesenborgs, Gustavo Alberto Rovelo Ruiz 等ICLR 2026 · 被引用 2 次
- The Double-Edged Nature of the Rashomon Set for Trustworthy Machine LearningEthan Hsu, Harry Chen, Chudi Zhong, Lesia SemenovaICML 2026 · 被引用 1 次
它引用的顶会 Paper15
- Predictive Multiplicity in ClassificationCharles T. Marx, Flávio P. Calmon, Berk UstunICML 2020 · 被引用 197 次
- Generalized and Scalable Optimal Sparse Decision TreesJimmy Lin, Chudi Zhong, Diane Hu, Cynthia Rudin 等ICML 2020 · 被引用 174 次
- Label Noise SGD Provably Prefers Flat Global MinimizersAlex Damian, Tengyu Ma, Jason D. LeeNeurIPS 2021 · 被引用 155 次
- Simple and Effective Regularization Methods for Training on Noisily Labeled Data with Generalization GuaranteeWei Hu, Zhiyuan Li, Dingli YuICLR 2020 · 被引用 140 次
- Exploring the Whole Rashomon Set of Sparse Decision TreesRui Xin, Chudi Zhong, Zhi Chen, Takuya Takagi 等NeurIPS 2022 · 被引用 117 次
相关 Paper
- RashomonGB: Analyzing the Rashomon Effect and Mitigating Predictive Multiplicity in Gradient BoostingHsiang Hsu, Ivan Brugere, Shubham Sharma, Freddy Lécué 等NeurIPS 2024 · 被引用 13 次
- Rashomon Capacity: A Metric for Predictive Multiplicity in ClassificationHsiang Hsu, Flávio P. CalmonNeurIPS 2022 · 被引用 65 次
- From Rashomon Theory to PRAXIS: Efficient Decision Tree Rashomon SetsZakk Heile, Hayden McTavish, Varun Babbar, Margo Seltzer 等ICML 2026 · 被引用 1 次
- Efficient Exploration of the Rashomon Set of Rule-Set ModelsMartino Ciaperoni, Han Xiao, Aristides GionisKDD 2024 · 被引用 3 次
- Exploring the cloud of feature interaction scores in a Rashomon setSichao Li, Rong Wang, Quanling Deng, Amanda S. BarnardICLR 2024 · 被引用 9 次
