Learning Interpretable Low-dimensional Representation via Physical Symmetry
Xuanjie Liu, Daniel Chin, Yichen Huang, Gus Xia
Abstract
We have recently seen great progress in learning interpretable music representations, ranging from basic factors, such as pitch and timbre, to high-level concepts, such as chord and texture. However, most methods rely heavily on music domain knowledge. It remains an open question what general computational principles give rise to interpretable representations, especially low-dim factors that agree with human perception. In this study, we take inspiration from modern physics and use physical symmetry as a self consistency constraint for the latent space of time-series data. Specifically, it requires the prior model that characterises the dynamics of the latent states to be equivariant with respect to certain group transformations. We show that physical symmetry leads the model to learn a linear pitch factor from unlabelled monophonic music audio in a self-supervised fashion. In addition, the same methodology can be applied to computer vision, learning a 3D Cartesian space from videos of a simple moving object without labels. Furthermore, physical symmetry naturally leads to counterfactual representation augmentation, a new technique which improves sample efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4aedee4b-b9af-4b4d-86b6-69d91dbf8a97Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Spatio-temporal Self-Supervised Representation Learning for 3D Point CloudsSiyuan Huang, Yichen Xie, Song-Chun Zhu, Yixin ZhuICCV 2021 · 259 citations
- GRF: Learning a General Radiance Field for 3D Representation and RenderingAlex Trevithick, Bo YangICCV 2021 · 258 citations
- Towards Nonlinear Disentanglement in Natural Data with Temporal Sparse CodingDavid A. Klindt, Lukas Schott, Yash Sharma, Ivan Ustyuzhaninov et al.ICLR 2021 · 156 citations
- Equivariant Neural RenderingEmilien Dupont, Miguel Bautista Martin, Alex Colburn, Aditya Sankar et al.ICML 2020 · 69 citations
- Contrastively Disentangled Sequential Variational AutoencoderJunwen Bai, Weiran Wang, Carla P. GomesNeurIPS 2021 · 60 citations
Related papers
- Learning Disentangled Representations and Group Structure of Dynamical EnvironmentsRobin Quessard, Thomas D. Barrett, William R. ClementsNeurIPS 2020 · 53 citations
- Unsupervised Learning of Equivariant Structure from SequencesTakeru Miyato, Masanori Koyama, Kenji FukumizuNeurIPS 2022 · 17 citations
- Structuring Representations Using Group InvariantsMehran Shakerinava, Arnab Kumar Mondal, Siamak RavanbakhshNeurIPS 2022 · 23 citations
- Latent Mixture of Symmetries for Sample-Efficient Dynamic LearningHaoran Li, Chenhan Xiao, Muhao Guo, Yang WengNeurIPS 2025 · 7 citations
- Homomorphism AutoEncoder - Learning Group Structured Representations from Observed TransitionsHamza Keurti, Hsiao-Ru Pan, Michel Besserve, Benjamin F. Grewe et al.ICML 2023 · 22 citations
