In-Context Symmetries: Self-Supervised Learning through Contextual World Models
Sharut Gupta, Chenyu Wang, Yifei Wang, Tommi S. Jaakkola, Stefanie Jegelka
Abstract
At the core of self-supervised learning for vision is the idea of learning invariant or equivariant representations with respect to a set of data transformations. This approach, however, introduces strong inductive biases, which can render the representations fragile in downstream tasks that do not conform to these symmetries. In this work, drawing insights from world models, we propose to instead learn a general representation that can adapt to be invariant or equivariant to different transformations by paying attention to context -- a memory module that tracks task-specific states, actions, and future states. Here, the action is the transformation, while the current and future states respectively represent the input's representation before and after the transformation. Our proposed algorithm, Contextual Self-Supervised Learning (ContextSSL), learns equivariance to all transformations (as opposed to invariance). In this way, the model can learn to encode all relevant features as general representations while having the versatility to tail down to task-wise symmetries when given a few examples as the context. Empirically, we demonstrate significant performance gains over existing methods on equivariance-related tasks, supported by both qualitative and quantitative evaluations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- seq-JEPA: Autoregressive Predictive Learning of Invariant-Equivariant World ModelsHafez Ghaemi, Eilif B. Muller, Shahab BakhtiariNeurIPS 2025 · 8 citations
- Context and Diversity Matter: The Emergence of In-Context Learning in World ModelsFan Wang, ZHIYUAN CHEN, YUXUAN ZHONG, Sunjian Zheng et al.ICLR 2026 · 5 citations
- Self-Supervised Learning from Structural InvarianceYipeng Zhang, Hafez Ghaemi, Jungyoon Lee, Shahab Bakhtiari et al.ICLR 2026 · 1 citation
- Any-Subgroup Equivariant Networks via Symmetry BreakingAbhinav Goel, Derek Lim, Hannah Lawrence, Stefanie Jegelka et al.ICLR 2026
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
Related papers
- Equivariant Self-Supervised Learning: Encouraging Equivariance in RepresentationsRumen Dangovski, Li Jing, Charlotte Loh, Seungwook Han et al.ICLR 2022 · 54 citations
- Contrastive-Equivariant Self-Supervised Learning Improves Alignment with Primate Visual Area ITThomas E. Yerxa, Jenelle Feather, Eero P. Simoncelli, SueYeon ChungNeurIPS 2024 · 12 citations
- Amortised Invariance Learning for Contrastive Self-SupervisionRuchika Chavhan, Jan Stuehmer, Calum Heggan, Mehrdad Yaghoobi et al.ICLR 2023 · 2 citations
- Learning from Memory: Non-Parametric Memory Augmented Self-Supervised Learning of Visual FeaturesThalles Silva, Hélio Pedrini, Adín Ramírez RiveraICML 2024 · 7 citations
- DUET: 2D Structured and Approximately Equivariant RepresentationsXavier Suau, Federico Danieli, T. Anderson Keller, Arno Blaas et al.ICML 2023 · 3 citations
