Efficiently learning structured distributions from untrusted batches
Sitan Chen, Jerry Li, Ankur Moitra
Abstract
We study the problem, introduced by Qiao and Valiant [QV17], of learning from untrusted batches. Here, we assume m users, all of whom have samples from some underlying distribution p over 1, . . . , n. Each user sends a batch of k i.i.d. samples from this distribution; however an ǫ-fraction of users are untrustworthy and can send adversarially chosen responses. The goal of the algorithm is then to learn p in total variation distance. When k = 1 this is the standard robust univariate density estimation setting and it is well-understood that Ω(ǫ) error is unavoidable. Suprisingly, [QV17] gave an estimator which improves upon this rate when k is large. Unfortunately, their algorithms run in time which is exponential in either n or k.
We first give a sequence of polynomial time algorithms whose estimation error approaches the information-theoretically optimal bound for this problem. Our approach is based on recent algorithms derived from the sum-of-squares hierarchy, in the context of high-dimensional robust estimation. We show that algorithms for learning from untrusted batches can also be cast in this framework, but by working with a more complicated set of test functions.
It turns out that this abstraction is quite powerful, and can be generalized to incorporate additional problem specific constraints. Our second and main result is to show that this technology can be leveraged to build in prior knowledge about the shape of the distribution. Crucially, this allows us to reduce the sample complexity of learning from untrusted batches to polylogarithmic in n for most natural classes of distributions, which is important in many applications. To do so, we demonstrate that these sum-of-squares algorithms for robust mean estimation can be made to handle complex combinatorial constraints (e.g. those arising from VC theory), which may be of independent technical interest.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a3d20af0-40f7-41ce-a737-93e143cc282aCited by top-tier papers7
- Learning Structured Distributions From Untrusted Batches: Faster and SimplerSitan Chen, Jerry Li, Ankur MoitraNeurIPS 2020 · 19 citations
- On the Sample Complexity of Adversarial Multi-Source PAC LearningNikola Konstantinov, Elias Frantar, Dan Alistarh, Christoph LampertICML 2020 · 18 citations
- A General Method for Robust Learning from BatchesAyush Jain, Alon OrlitskyNeurIPS 2020 · 17 citations
- Optimal Robust Learning of Discrete Distributions from BatchesAyush Jain, Alon OrlitskyICML 2020 · 16 citations
- Robust Density Estimation from Batches: The Best Things in Life are (Nearly) FreeAyush Jain, Alon OrlitskyICML 2021 · 10 citations
Builds on1
Related papers
- Robust Estimation Under Heterogeneous Corruption RatesSyomantak Chaudhuri, Jerry Li, Thomas A. CourtadeNeurIPS 2025
- Robustness Implies Privacy in Statistical EstimationSamuel B. Hopkins, Gautam Kamath, Mahbod Majid, Shyam NarayananSTOC 2023 · 16 citations
- Product Distribution Learning with Imperfect AdviceArnab Bhattacharyya, Davin Choo, Philips George John, Themis GouleakisNeurIPS 2025 · 3 citations
- Learning with User-Level PrivacyDaniel Levy, Ziteng Sun, Kareem Amin, Satyen Kale et al.NeurIPS 2021 · 113 citations
- Robust Learning of Mixtures of GaussiansDaniel M. KaneSODA 2021 · 12 citations
