To Augment or Not to Augment? Diagnosing Distributional Symmetry Breaking
Hannah Lawrence, Elyssa F. Hofgard, Vasco Portilheiro, Yuxuan Chen, Tess E. Smidt, Robin Walters
摘要
Symmetry-aware methods for machine learning, such as data augmentation and equivariant architectures, encourage correct model behavior on all transformations (e.g. rotations or permutations) of the original dataset. These methods can improve generalization and sample efficiency, under the assumption that the transformed datapoints are highly probable, or"important", under the test distribution. In this work, we develop a method for critically evaluating this assumption. In particular, we propose a metric to quantify the amount of symmetry breaking in a dataset, via a two-sample classifier test that distinguishes between the original dataset and its randomly augmented equivalent. We validate our metric on synthetic datasets, and then use it to uncover surprisingly high degrees of symmetry-breaking in several benchmark point cloud datasets, constituting a severe form of dataset bias. We show theoretically that distributional symmetry-breaking can prevent invariant methods from performing optimally even when the underlying labels are truly invariant, for invariant ridge regression in the infinite feature limit. Empirically, the implication for symmetry-aware methods is dataset-dependent: equivariant methods still impart benefits on some symmetry-biased datasets, but not others, particularly when the symmetry bias is predictive of the labels. Overall, these findings suggest that understanding equivariance -- both when it works, and why -- may require rethinking symmetry biases in the data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper28
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng 等NeurIPS 2021 · 被引用 1,632 次
- Vector Neurons: A General Framework for SO(3)-Equivariant NetworksCongyue Deng, Or Litany, Yueqi Duan, Adrien Poulenard 等ICCV 2021 · 被引用 411 次
- EquiformerV2: Improved Equivariant Transformer for Scaling to Higher-Degree RepresentationsYi-Lun Liao, Brandon M. Wood, Abhishek Das, Tess E. SmidtICLR 2024 · 被引用 311 次
- A Group-Theoretic Framework for Data AugmentationShuxiao Chen, Edgar Dobriban, Jane H. LeeNeurIPS 2020 · 被引用 254 次
- Incorporating Symmetry into Deep Dynamics Models for Improved GeneralizationRui Wang, Robin Walters, Rose YuICLR 2021 · 被引用 201 次
相关 Paper
- In What Ways Are Deep Neural Networks Invariant and How Should We Measure This?Henry Kvinge, Tegan Emerson, Grayson Jorgenson, Scott Vasquez 等NeurIPS 2022 · 被引用 15 次
- Smooth, exact rotational symmetrization for deep learning on point cloudsSergey Pozdnyakov, Michele CeriottiNeurIPS 2023 · 被引用 68 次
- ART-Point: Improving Rotation Robustness of Point Cloud Classifiers via Adversarial RotationRuibin Wang, Yibo Yang, Dacheng TaoCVPR 2022 · 被引用 22 次
- Probing Equivariance and Symmetry Breaking in Convolutional NetworksSharvaree Vadgama, Mohammad Mohaiminul Islam, Domas Buracas, Christian Shewmake 等NeurIPS 2025 · 被引用 15 次
- SymAttack: Symmetry-aware Imperceptible Adversarial Attacks on 3D Point CloudsKeke Tang, Zhensu Wang, Weilong Peng, Lujie Huang 等ACM MM 2024 · 被引用 10 次
