Beyond the Permutation Symmetry of Transformers: The Role of Rotation for Model Fusion
Binchi Zhang, Zaiyi Zheng, Zhengzhang Chen, Jundong Li
摘要
Symmetry in the parameter space of deep neural networks (DNNs) has proven beneficial for various deep learning applications. A well-known example is the permutation symmetry in Multi-Layer Perceptrons (MLPs), where permuting the rows of weight matrices in one layer and applying the inverse permutation to adjacent layers yields a functionally equivalent model. While permutation symmetry fully characterizes the equivalence set for MLPs, its discrete nature limits its utility for transformers. In this paper, we introduce rotation symmetry, a novel form of parameter space symmetry for transformers that generalizes permutation symmetry by rotating parameter matrices in self-attention layers. Unlike permutation symmetry, rotation symmetry operates in a continuous domain, thereby significantly expanding the equivalence set for transformers. Based on this property, we propose a theoretically optimal parameter matching algorithm as a plug-and-play module to enhance model fusion. We evaluate our approach using pre-trained transformers across diverse natural language and vision tasks. Experimental results demonstrate that our rotation symmetrybased matching algorithm substantially improves model fusion, highlighting the potential of parameter space symmetry to facilitate model fusion. Our code is available on https://github. com/zhengzaiyi/RotationSymmetry .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Label-Free Cross-Task LoRA Merging with Null-Space CompressionWonyoung Lee, Wooseong Jeong, Kuk-Jin YoonCVPR 2026 · 被引用 3 次
- Quantifying LLM Attention-Head Stability: Implications for Circuit UniversalityKaran Bali, Jack Stanley, Praneet Suresh, Danilo BzdokICML 2026 · 被引用 2 次
- Preference-Aligned LoRA Merging: Preserving Subspace Coverage and Addressing Directional AnisotropyWooseong Jeong, Wonyoung Lee, Kuk-Jin YoonCVPR 2026 · 被引用 1 次
- MOMO: Mars Orbital MOdel Foundation Model for Mars Orbital ApplicationsMirali Purohit, Bimal Gajera, Irish Mehta, Bhanu Tokas 等CVPR 2026 · 被引用 1 次
- Learning from Diverse Reasoning Paths with Routing and CollaborationZhenyu Lei, Zhen Tan, Song Wang, Yaochen Zhu 等EMNLP 2025
它引用的顶会 Paper41
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang 等AAAI 2021 · 被引用 7,289 次
相关 Paper
- Transformer Fusion with Optimal TransportMoritz Imfeld, Jacopo Graldi, Marco Giordano, Thomas Hofmann 等ICLR 2024 · 被引用 35 次
- REViT: Roto-reflection Equivariant Convolutional Vision TransformerSheir A. Zaheer, Alexander Holston, Chan Youn ParkICML 2026
- Universal Neural FunctionalsAllan Zhou, Chelsea Finn, James HarrisonNeurIPS 2024 · 被引用 27 次
- Leveraging 3D Geometric Priors in 2D Rotation Symmetry DetectionAhyun Seo, Minsu ChoCVPR 2025
- Learning Symmetric Embeddings for Equivariant World ModelsJung Yeon Park, Ondrej Biza, Linfeng Zhao, Jan-Willem van de Meent 等ICML 2022 · 被引用 56 次
