How Much is Unseen Depends Chiefly on Information About the Seen
Seongmin Lee, Marcel Böhme
摘要
The missing mass refers to the proportion of data points in an unknown population of classifier inputs that belong to classes not present in the classifier's training data, which is assumed to be a random sample from that unknown population. We find that in expectation the missing mass is entirely determined by the number f k of classes that do appear in the training data the same number of times and an exponentially decaying error. While this is the first precise characterization of the expected missing mass in terms of the sample, the induced estimator suffers from an impractically high variance. However, our theory suggests a large search space of nearly unbiased estimators that can be searched effectively and efficiently. Hence, we cast distribution-free estimation as an optimization problem to find a distribution-specific estimator with a minimized mean-squared error (MSE), given only the sample. In our experiments, our search algorithm discovers estimators that have a substantially smaller MSE than the state-of-the-art Good-Turing estimator. This holds for over 93% of runs when there are at least as many samples as classes. Our estimators' MSE is roughly 80% of the Good-Turing estimator's.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Incoherence as Oracle-less Measure of Error in LLM-Based Code GenerationThomas Jean-Michel Valentin, Ardi Madadi, Gaetano Sapia, Marcel BöhmeAAAI 2026 · 被引用 4 次
- Accounting for Missing Events in Statistical Information Leakage AnalysisSeongmin Lee, Shreyas Minocha, Marcel BöhmeICSE 2025 · 被引用 2 次
- Dependency-aware Residual Risk AnalysisSeongmin Lee, Marcel BöhmeICSE 2026
它引用的顶会 Paper5
- Estimating residual risk in greybox fuzzingMarcel Böhme, Danushka Liyanage, Valentin WüstholzFSE 2021 · 被引用 27 次
- Statistical Reachability AnalysisSeongmin Lee, Marcel BöhmeFSE 2023 · 被引用 12 次
- Optimal Prediction of the Number of Unseen Species with MultiplicityYi Hao, Ping LiNeurIPS 2020 · 被引用 9 次
- Extrapolating Coverage Rate in Greybox FuzzingDanushka Liyanage, Seongmin Lee, Chakkrit Tantithamthavorn, Marcel BöhmeICSE 2024 · 被引用 6 次
- Accounting for Missing Events in Statistical Information Leakage AnalysisSeongmin Lee, Shreyas Minocha, Marcel BöhmeICSE 2025 · 被引用 2 次
相关 Paper
- Learning discrete distributions with infinite supportDoron Cohen, Aryeh Kontorovich, Geoffrey WolferNeurIPS 2020 · 被引用 21 次
- Conformalized matrix completionYu Gui, Rina Barber, Cong MaNeurIPS 2023 · 被引用 24 次
- Learning-based Support Estimation in Sublinear TimeTalya Eden, Piotr Indyk, Shyam Narayanan, Ronitt Rubinfeld 等ICLR 2021 · 被引用 8 次
- More Accurate Learning of k-DNF Reference ClassesBrendan Juba, Hengxuan LiAAAI 2020 · 被引用 3 次
- Kernel-Based Tests for Likelihood-Free Hypothesis TestingPatrik Róbert Gerber, Tianze Jiang, Yury Polyanskiy, Rui SunNeurIPS 2023 · 被引用 5 次
