Soft Task-Aware Routing of Experts for Equivariant Representation Learning
Jaebyeong Jeon, Hyunseo Jang, Jy-yong Sohn, Kibok Lee
摘要
Equivariant representation learning aims to capture variations induced by input transformations in the representation space, whereas invariant representation learning encodes semantic information by disregarding such transformations. Recent studies have shown that jointly learning both types of representations is often beneficial for downstream tasks, typically by employing separate projection heads. However, this design overlooks information shared between invariant and equivariant learning, which leads to redundant feature learning and inefficient use of model capacity. To address this, we introduce Soft Task-Aware Routing (STAR), a routing strategy for projection heads that models them as experts. STAR induces the experts to specialize in capturing either shared or task-specific information, thereby reducing redundant feature learning. We validate this effect by observing lower canonical correlations between invariant and equivariant embeddings. Experimental results show consistent improvements across diverse transfer learning tasks. The code is available at https://github.com/YonseiML/star.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 被引用 2,340 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
相关 Paper
- STAR: Rethinking MoE Routing as Structure-Aware Subspace LearningSumin Park, Noseong ParkICML 2026
- Tunable Soft Equivariance with GuaranteesMd Ashiqur Rahman, Lim Jun Hao, Jeremiah Jiang, Teck-Yian Lim 等CVPR 2026
- Soft Equivariance Regularization for Invariant Self-Supervised LearningJoohyung Lee, Changhun Kim, Hyunsu Kim, Kwanhyung Lee 等ICLR 2026 · 被引用 1 次
- Scalable Transfer Learning with Expert ModelsJoan Puigcerver, Carlos Riquelme Ruiz, Basil Mustafa, Cédric Renggli 等ICLR 2021 · 被引用 70 次
- Relaxing Equivariance Constraints with Non-stationary Continuous FiltersTycho F. A. van der Ouderaa, David W. Romero, Mark van der WilkNeurIPS 2022 · 被引用 51 次
