From Generalist to Specialist Representation
Yujia Zheng, Fan Feng, Yuke Li, Shaoan Xie, Kevin Murphy, Kun Zhang
摘要
Given a generalist model, learning a task-relevant specialist representation is fundamental for downstream applications. Identifiability, the asymptotic guarantee of recovering the ground-truth representation, is critical because it sets the ultimate limit of any model, even with infinite data and computation. We study this problem in a completely nonparametric setting, without relying on interventions, parametric forms, or structural constraints. We first prove that the structure between time steps and tasks is identifiable in a fully unsupervised manner, even when sequences lack strict temporal dependence and may exhibit disconnections, and task assignments can follow arbitrarily complex and interleaving structures. We then prove that, within each time step, the task-relevant latent representation can be disentangled from the irrelevant part under a simple sparsity regularization, without any additional information or parametric constraints. Together, these results establish a hierarchical foundation: task structure is identifiable across time steps, and task-relevant latent representations are identifiable within each step. To our knowledge, each result provides a first general nonparametric identifiability guarantee, and together they mark a step toward provably moving from generalist to specialist models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Self-Supervised Learning with Data Augmentations Provably Isolates Content from StyleJulius von Kügelgen, Yash Sharma, Luigi Gresele, Wieland Brendel 等NeurIPS 2021 · 被引用 421 次
- Weakly supervised causal representation learningJohann Brehmer, Pim de Haan, Phillip Lippe, Taco S. CohenNeurIPS 2022 · 被引用 196 次
- Independent mechanism analysis, a new concept?Luigi Gresele, Julius von Kügelgen, Vincent Stimper, Bernhard Schölkopf 等NeurIPS 2021 · 被引用 133 次
相关 Paper
- Nonparametric Identifiability of Causal Representations from Unknown InterventionsJulius von Kügelgen, Michel Besserve, Wendong Liang, Luigi Gresele 等NeurIPS 2023 · 被引用 127 次
- Temporally Disentangled Representation LearningWeiran Yao, Guangyi Chen, Kun ZhangNeurIPS 2022 · 被引用 84 次
- Identification of Intermittent Temporal Latent ProcessYuke Li, Yujia Zheng, Guangyi Chen, Kun Zhang 等ICLR 2025
- Temporally Disentangled Representation Learning under Unknown NonstationarityXiangchen Song, Weiran Yao, Yewen Fan, Xinshuai Dong 等NeurIPS 2023 · 被引用 36 次
- Causal Representation Learning Made Identifiable by Grouping of Observational VariablesHiroshi Morioka, Aapo HyvärinenICML 2024 · 被引用 26 次
