On the Power of Source Screening for Learning Shared Feature Extractors
Muxing Wang, Connor Mclaughlin, Lili Su
摘要
Learning with shared representation is widely recognized as an effective way to separate commonalities from heterogeneity across various heterogeneous sources. Most existing work includes all related data sources via simultaneously training a common feature extractor and source-specific heads. It is well understood that data sources with low relevance or poor quality may hinder representation learning. In this paper, we further dive into the question of which data sources should be learned jointly by focusing on the traditionally deemed "good" collection of sources, in which individual sources have similar relevance and qualities with respect to the true underlying common structure. Towards tractability, we focus on the linear setting where sources share a low-dimensional subspace. We find that source screening can play a central role in statistically optimal subspace estimation. We show that, for a broad class of problem instances, training on a carefully selected subset of sources suffices to achieve minimax optimality, even when a substantial portion of data is discarded. We formalize the notion of an informative subpopulation, develop algorithms and practical heuristics for identifying such subsets, and validate their effectiveness through both theoretical analysis and empirical evaluations on synthetic and real-world datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang 等ICLR 2020 · 被引用 2,930 次
- Personalized Federated Learning with Theoretical Guarantees: A Model-Agnostic Meta-Learning ApproachAlireza Fallah, Aryan Mokhtari, Asuman E. OzdaglarNeurIPS 2020 · 被引用 1,354 次
- Exploiting Shared Representations for Personalized Federated LearningLiam Collins, Hamed Hassani, Aryan Mokhtari, Sanjay ShakkottaiICML 2021 · 被引用 1,081 次
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 被引用 671 次
相关 Paper
- GIO: Gradient Information Optimization for Training Dataset SelectionDante Everaert, Christopher PottsICLR 2024 · 被引用 12 次
- Active Multi-Task Representation LearningYifang Chen, Kevin Jamieson, Simon S. DuICML 2022 · 被引用 18 次
- Information Retention via Learning Supplemental FeaturesZhipeng Xie, Yahe LiICLR 2024 · 被引用 1 次
- Few-Shot Learning via Learning the Representation, ProvablySimon Shaolei Du, Wei Hu, Sham M. Kakade, Jason D. Lee 等ICLR 2021 · 被引用 56 次
- Training Subset Selection for Weak SupervisionHunter Lang, Aravindan Vijayaraghavan, David A. SontagNeurIPS 2022 · 被引用 27 次
