Simultaneous Inference for Massive Data: Distributed Bootstrap
Yang Yu, Shih-Kang Chao, Guang Cheng
Abstract
In this paper, we propose a bootstrap method applied to massive data processed distributedly in a large number of machines. This new method is computationally efficient in that we bootstrap on the master machine without over-resampling, typically required by existing methods , while provably achieving optimal statistical efficiency with minimal communication. Our method does not require repeatedly re-fitting the model but only applies multiplier bootstrap in the master machine on the gradients received from the worker machines. Simulations validate our theory.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 18b560ba-5eb2-4540-8868-49c2a7d2bc2dCited by top-tier papers1
Ask how each one uses itRelated papers
- Bootstrap in High Dimension with Low ComputationHenry Lam, Zhenyuan LiuICML 2023 · 7 citations
- One-shot Distributed Ridge Regression in High DimensionsYue Sheng, Edgar DobribanICML 2020 · 49 citations
- Distributed Least Squares in Small Space via Sketching and Bias ReductionSachin Garg, Kevin Tan, Michal DerezinskiNeurIPS 2024 · 5 citations
- Communication-efficient Distributed Learning for Large Batch OptimizationRui Liu, Barzan MozafariICML 2022 · 9 citations
- Orthogonal Bootstrap: Efficient Simulation of Input UncertaintyKaizhao Liu, José H. Blanchet, Lexing Ying, Yiping LuICML 2024 · 2 citations
