To Augment or Not to Augment? Diagnosing Distributional Symmetry Breaking
Hannah Lawrence, Elyssa F. Hofgard, Vasco Portilheiro, Yuxuan Chen, Tess E. Smidt, Robin Walters
Abstract
Symmetry-aware methods for machine learning, such as data augmentation and equivariant architectures, encourage correct model behavior on all transformations (e.g. rotations or permutations) of the original dataset. These methods can improve generalization and sample efficiency, under the assumption that the transformed datapoints are highly probable, or"important", under the test distribution. In this work, we develop a method for critically evaluating this assumption. In particular, we propose a metric to quantify the amount of symmetry breaking in a dataset, via a two-sample classifier test that distinguishes between the original dataset and its randomly augmented equivalent. We validate our metric on synthetic datasets, and then use it to uncover surprisingly high degrees of symmetry-breaking in several benchmark point cloud datasets, constituting a severe form of dataset bias. We show theoretically that distributional symmetry-breaking can prevent invariant methods from performing optimally even when the underlying labels are truly invariant, for invariant ridge regression in the infinite feature limit. Empirically, the implication for symmetry-aware methods is dataset-dependent: equivariant methods still impart benefits on some symmetry-biased datasets, but not others, particularly when the symmetry bias is predictive of the labels. Overall, these findings suggest that understanding equivariance -- both when it works, and why -- may require rethinking symmetry biases in the data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e28075d2-7eec-434f-9816-c5c4306c5d18Builds on28
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng et al.NeurIPS 2021 · 1,632 citations
- Vector Neurons: A General Framework for SO(3)-Equivariant NetworksCongyue Deng, Or Litany, Yueqi Duan, Adrien Poulenard et al.ICCV 2021 · 411 citations
- EquiformerV2: Improved Equivariant Transformer for Scaling to Higher-Degree RepresentationsYi-Lun Liao, Brandon M. Wood, Abhishek Das, Tess E. SmidtICLR 2024 · 311 citations
- A Group-Theoretic Framework for Data AugmentationShuxiao Chen, Edgar Dobriban, Jane H. LeeNeurIPS 2020 · 254 citations
- Incorporating Symmetry into Deep Dynamics Models for Improved GeneralizationRui Wang, Robin Walters, Rose YuICLR 2021 · 201 citations
Related papers
- In What Ways Are Deep Neural Networks Invariant and How Should We Measure This?Henry Kvinge, Tegan Emerson, Grayson Jorgenson, Scott Vasquez et al.NeurIPS 2022 · 15 citations
- Smooth, exact rotational symmetrization for deep learning on point cloudsSergey Pozdnyakov, Michele CeriottiNeurIPS 2023 · 68 citations
- ART-Point: Improving Rotation Robustness of Point Cloud Classifiers via Adversarial RotationRuibin Wang, Yibo Yang, Dacheng TaoCVPR 2022 · 22 citations
- Probing Equivariance and Symmetry Breaking in Convolutional NetworksSharvaree Vadgama, Mohammad Mohaiminul Islam, Domas Buracas, Christian Shewmake et al.NeurIPS 2025 · 15 citations
- SymAttack: Symmetry-aware Imperceptible Adversarial Attacks on 3D Point CloudsKeke Tang, Zhensu Wang, Weilong Peng, Lujie Huang et al.ACM MM 2024 · 10 citations
