How Much is Unseen Depends Chiefly on Information About the Seen
Seongmin Lee, Marcel Böhme
Abstract
The missing mass refers to the proportion of data points in an unknown population of classifier inputs that belong to classes not present in the classifier's training data, which is assumed to be a random sample from that unknown population. We find that in expectation the missing mass is entirely determined by the number f k of classes that do appear in the training data the same number of times and an exponentially decaying error. While this is the first precise characterization of the expected missing mass in terms of the sample, the induced estimator suffers from an impractically high variance. However, our theory suggests a large search space of nearly unbiased estimators that can be searched effectively and efficiently. Hence, we cast distribution-free estimation as an optimization problem to find a distribution-specific estimator with a minimized mean-squared error (MSE), given only the sample. In our experiments, our search algorithm discovers estimators that have a substantially smaller MSE than the state-of-the-art Good-Turing estimator. This holds for over 93% of runs when there are at least as many samples as classes. Our estimators' MSE is roughly 80% of the Good-Turing estimator's.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb76fd79-18b7-47b3-9d63-69f8ca32b739Cited by top-tier papers3
- Incoherence as Oracle-less Measure of Error in LLM-Based Code GenerationThomas Jean-Michel Valentin, Ardi Madadi, Gaetano Sapia, Marcel BöhmeAAAI 2026 · 4 citations
- Accounting for Missing Events in Statistical Information Leakage AnalysisSeongmin Lee, Shreyas Minocha, Marcel BöhmeICSE 2025 · 2 citations
- Dependency-aware Residual Risk AnalysisSeongmin Lee, Marcel BöhmeICSE 2026
Builds on5
- Estimating residual risk in greybox fuzzingMarcel Böhme, Danushka Liyanage, Valentin WüstholzFSE 2021 · 27 citations
- Statistical Reachability AnalysisSeongmin Lee, Marcel BöhmeFSE 2023 · 12 citations
- Optimal Prediction of the Number of Unseen Species with MultiplicityYi Hao, Ping LiNeurIPS 2020 · 9 citations
- Extrapolating Coverage Rate in Greybox FuzzingDanushka Liyanage, Seongmin Lee, Chakkrit Tantithamthavorn, Marcel BöhmeICSE 2024 · 6 citations
- Accounting for Missing Events in Statistical Information Leakage AnalysisSeongmin Lee, Shreyas Minocha, Marcel BöhmeICSE 2025 · 2 citations
Related papers
- Learning discrete distributions with infinite supportDoron Cohen, Aryeh Kontorovich, Geoffrey WolferNeurIPS 2020 · 21 citations
- Conformalized matrix completionYu Gui, Rina Barber, Cong MaNeurIPS 2023 · 24 citations
- Learning-based Support Estimation in Sublinear TimeTalya Eden, Piotr Indyk, Shyam Narayanan, Ronitt Rubinfeld et al.ICLR 2021 · 8 citations
- More Accurate Learning of k-DNF Reference ClassesBrendan Juba, Hengxuan LiAAAI 2020 · 3 citations
- Kernel-Based Tests for Likelihood-Free Hypothesis TestingPatrik Róbert Gerber, Tianze Jiang, Yury Polyanskiy, Rui SunNeurIPS 2023 · 5 citations
