Efficient Statistics With Unknown Truncation, Polynomial Time Algorithms, Beyond Gaussians
Jane H. Lee, Anay Mehrotra, Manolis Zampetakis
摘要
We study the estimation of distributional parameters when samples are shown only if they fall in some unknown set. Kontonis, Tzamos, and Zampetakis (FOCS'19) gave an algorithm for finding parameters for the special case of Gaussian distributions with diagonal covariance matrix. Recently, Diakonikolas, Kane, Pittas, and Zarifis (COLT'24) showed that an exponential dependence on the inverse of the accuracy parameter is necessary even when the set belongs to some well-behaved classes. These works leave the following open problems which we address in this work: Can we estimate the parameters of any Gaussian or even extend the results beyond Gaussians? Can we design polynomial-time algorithms when for simple sets such as a halfspace? Toward the first question, we provide an estimation algorithm for any exponential family that satisfies some structural assumptions and any unknown set that is approximable by polynomials. This result has two important applications: (a)The first algorithm for estimating arbitrary Gaussian distributions (even with non-diagonal covariance matrix) from samples truncated to unknown set; and (b)The first algorithm for linear regression with unknown truncation and Gaussian features. To address the second question, we provide an algorithm with polynomial sample and time complexity that works for a set of exponential families (that contains multivariate Gaussians) when the unknown survival set is a halfspace or an axis-aligned rectangle.11A preliminary version of this paper incorrectly claimed the result for finite unions of axis-aligned rectangles. The result only holds for a single axis-aligned rectangle. This is the first fully polynomial time algorithm for estimation with an unknown truncation set. Along the way, we develop new tools that may be of independent interest, including: (c)The first polynomial time algorithm for learning halfspaces using only positive examples when the samples have an unknown Gaussian distribution; and (d)A reduction from PAC learning with positive and unlabeled samples to PAC learning with positive and negative samples that is robust to certain covariate shifts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- On the Limits of Language Generation: Trade-Offs between Hallucination and Mode-CollapseAlkis Kalavasis, Anay Mehrotra, Grigoris VelegkasSTOC 2025 · 被引用 2 次
- Smoothed Analysis of Learning from Positive SamplesJane H. Lee, Anay Mehrotra, Manolis ZampetakisSTOC 2026 · 被引用 2 次
- Linear Regression with Unknown Truncation Beyond Gaussian FeaturesAlexandros Kouridakis, Anay Mehrotra, Alkis Kalavasis, Constantine CaramanisICML 2026 · 被引用 2 次
- Private Statistical Estimation via TruncationManolis Zampetakis, Felix ZhouNeurIPS 2025 · 被引用 1 次
- Towards a Holistic Understanding of Selection Bias for Causal Effect IdentificationYiwen (Evie) Qiu, Filip Kovačević, Shimeng Huang, Peter Spirtes 等ICML 2026
它引用的顶会 Paper12
- Provable benefits of score matchingChirag Pabbaraju, Dhruv Rohatgi, Anish Prasad Sevekari, Holden Lee 等NeurIPS 2023 · 被引用 19 次
- Truncated Linear Regression in High DimensionsConstantinos Daskalakis, Dhruv Rohatgi, Emmanouil ZampetakisNeurIPS 2020 · 被引用 19 次
- Efficient Truncated Linear Regression with Unknown Noise VarianceConstantinos Daskalakis, Patroklos Stefanou, Rui Yao, Emmanouil ZampetakisNeurIPS 2021 · 被引用 15 次
- A Computationally Efficient Method for Learning Exponential Family DistributionsAbhin Shah, Devavrat Shah, Gregory W. WornellNeurIPS 2021 · 被引用 15 次
- Learning Exponential Families from Truncated SamplesJane H. Lee, Andre Wibisono, Emmanouil ZampetakisNeurIPS 2023 · 被引用 7 次
相关 Paper
- Oracle efficient truncated statisticsKonstantinos Karatapanis, Vasilis Kontonis, Christos TzamosICLR 2025
- Algorithms for heavy-tailed statistics: regression, covariance estimation, and beyondYeshwanth Cherapanamjeri, Samuel B. Hopkins, Tarun Kathuria, Prasad Raghavendra 等STOC 2020 · 被引用 2 次
- Mean Estimation from Coarse Data: Characterizations and Efficient AlgorithmsAlkis Kalavasis, Anay Mehrotra, Manolis Zampetakis, Felix Zhou 等ICLR 2026
- Detecting Low-Degree TruncationAnindya De, Huan Li, Shivam Nadimpalli, Rocco A. ServedioSTOC 2024 · 被引用 2 次
- Efficiently learning halfspaces with Tsybakov noiseIlias Diakonikolas, Daniel M. Kane, Vasilis Kontonis, Christos Tzamos 等STOC 2021 · 被引用 2 次
