REOrdering Patches Improves Vision Models
Declan Kutscher, David M. Chan, Yutong Bai, Trevor Darrell, Ritwik Gupta
Abstract
Sequence models such as transformers require inputs to be represented as one-dimensional sequences. In vision, this typically involves flattening images using a fixed row-major (raster-scan) order. While full self-attention is permutation-equivariant, modern long-sequence transformers increasingly rely on architectural approximations that break this invariance and introduce sensitivity to patch ordering. We show that patch order significantly affects model performance in such settings, with simple alternatives like column-major or Hilbert curves yielding notable accuracy shifts. Motivated by this, we propose REOrder, a two-stage framework for discovering task-optimal patch orderings. First, we derive an information-theoretic prior by evaluating the compressibility of various patch sequences. Then, we learn a policy over permutations by optimizing a Plackett-Luce policy using REINFORCE. This approach enables efficient learning in a combinatorial permutation space. REOrder improves top-1 accuracy over row-major ordering on ImageNet-1K by up to 3.01% and Functional Map of the World by 13.35%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on6
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Sample Efficient Reinforcement Learning with REINFORCEJunzi Zhang, Jongho Kim, Brendan O'Donoghue, Stephen P. BoydAAAI 2021 · 162 citations
- Understanding and Improving Robustness of Vision Transformers through Patch-based Negative AugmentationYao Qin, Chiyuan Zhang, Ting Chen, Balaji Lakshminarayanan et al.NeurIPS 2022 · 68 citations
- Permutation Equivariance of Transformers and its ApplicationsHengyuan Xu, Liyao Xiang, Hangyu Ye, Dixi Yao et al.CVPR 2024 · 11 citations
- xT: Nested Tokenization for Larger Context in Large ImagesRitwik Gupta, Shufan Li, Tyler Zhu, Jitendra Malik et al.ICML 2024 · 9 citations
Related papers
- Discovering Non-monotonic Autoregressive Orderings with Variational InferenceXuanlin Li, Brandon Trabucco, Dong Huk Park, Michael Luo et al.ICLR 2021 · 17 citations
- Sequencer: Deep LSTM for Image ClassificationYuki Tatsunami, Masato TakiNeurIPS 2022 · 124 citations
- Rethinking and Improving Relative Position Encoding for Vision TransformerKan Wu, Houwen Peng, Minghao Chen, Jianlong Fu et al.ICCV 2021 · 427 citations
- Scalable Vision Transformers with Hierarchical PoolingZizheng Pan, Bohan Zhuang, Jing Liu, Haoyu He et al.ICCV 2021 · 154 citations
- Learnable Fourier Features for Multi-dimensional Spatial Positional EncodingYang Li, Si Si, Gang Li, Cho-Jui Hsieh et al.NeurIPS 2021 · 171 citations
