Collaborative Learning of Discrete Distributions under Heterogeneity and Communication Constraints
Xinmeng Huang, Donghwan Lee, Edgar Dobriban, Hamed Hassani
Abstract
In modern machine learning, users often have to collaborate to learn the distribution of the data. Communication can be a significant bottleneck. Prior work has studied homogeneous users -- i.e., whose data follow the same discrete distribution -- and has provided optimal communication-efficient methods for estimating that distribution. However, these methods rely heavily on homogeneity, and are less applicable in the common case when users' discrete distributions are heterogeneous. Here we consider a natural and tractable model of heterogeneity, where users' discrete distributions only vary sparsely, on a small number of entries. We propose a novel two-stage method named SHIFT: First, the users collaborate by communicating with the server to learn a central distribution; relying on methods from robust statistics. Then, the learned central distribution is fine-tuned to estimate their respective individual distribution. We show that SHIFT is minimax optimal in our model of heterogeneity and under communication constraints. Further, we provide experimental results using both synthetic data and -gram frequency estimation in the text domain, which corroborate its efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 566d8ed3-d629-46ef-9f99-2103c82da4e8Builds on10
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang et al.ICLR 2020 · 2,930 citations
- Exploiting Shared Representations for Personalized Federated LearningLiam Collins, Hamed Hassani, Aryan Mokhtari, Sanjay ShakkottaiICML 2021 · 1,081 citations
- Personalized Federated Learning using HypernetworksAviv Shamsian, Aviv Navon, Ethan Fetaya, Gal ChechikICML 2021 · 452 citations
- Scaling and Benchmarking Self-Supervised Visual Representation LearningPriya Goyal, Dhruv Mahajan, Abhinav Gupta, Ishan MisraICCV 2019 · 429 citations
- Provable Meta-Learning of Linear RepresentationsNilesh Tripuraneni, Chi Jin, Michael I. JordanICML 2021 · 218 citations
Related papers
- Distributed Estimation with Multiple Samples per User: Sharp Rates and Phase TransitionJayadev Acharya, Clément L. Canonne, Yuhan Liu, Ziteng Sun et al.NeurIPS 2021 · 16 citations
- Mean Estimation with User-level Privacy under Data HeterogeneityRachel Cummings, Vitaly Feldman, Audra McMillan, Kunal TalwarNeurIPS 2022 · 35 citations
- Private and Personalized Frequency Estimation in a Federated SettingAmrith Setlur, Vitaly Feldman, Kunal TalwarNeurIPS 2024 · 1 citation
- Multiply Robust Estimation for Local Distribution Shifts with Multiple DomainsSteven Wilkins-Reeves, Xu Chen, Qi Ma, Christine Agarwal et al.ICML 2024 · 2 citations
- Robust Federated Learning: The Case of Affine Distribution ShiftsAmirhossein Reisizadeh, Farzan Farnia, Ramtin Pedarsani, Ali JadbabaieNeurIPS 2020 · 196 citations
