REOrdering Patches Improves Vision Models
Declan Kutscher, David M. Chan, Yutong Bai, Trevor Darrell, Ritwik Gupta
摘要
Sequence models such as transformers require inputs to be represented as one-dimensional sequences. In vision, this typically involves flattening images using a fixed row-major (raster-scan) order. While full self-attention is permutation-equivariant, modern long-sequence transformers increasingly rely on architectural approximations that break this invariance and introduce sensitivity to patch ordering. We show that patch order significantly affects model performance in such settings, with simple alternatives like column-major or Hilbert curves yielding notable accuracy shifts. Motivated by this, we propose REOrder, a two-stage framework for discovering task-optimal patch orderings. First, we derive an information-theoretic prior by evaluating the compressibility of various patch sequences. Then, we learn a policy over permutations by optimizing a Plackett-Luce policy using REINFORCE. This approach enables efficient learning in a combinatorial permutation space. REOrder improves top-1 accuracy over row-major ordering on ImageNet-1K by up to 3.01% and Functional Map of the World by 13.35%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Sample Efficient Reinforcement Learning with REINFORCEJunzi Zhang, Jongho Kim, Brendan O'Donoghue, Stephen P. BoydAAAI 2021 · 被引用 162 次
- Understanding and Improving Robustness of Vision Transformers through Patch-based Negative AugmentationYao Qin, Chiyuan Zhang, Ting Chen, Balaji Lakshminarayanan 等NeurIPS 2022 · 被引用 68 次
- Permutation Equivariance of Transformers and its ApplicationsHengyuan Xu, Liyao Xiang, Hangyu Ye, Dixi Yao 等CVPR 2024 · 被引用 11 次
- xT: Nested Tokenization for Larger Context in Large ImagesRitwik Gupta, Shufan Li, Tyler Zhu, Jitendra Malik 等ICML 2024 · 被引用 9 次
相关 Paper
- Discovering Non-monotonic Autoregressive Orderings with Variational InferenceXuanlin Li, Brandon Trabucco, Dong Huk Park, Michael Luo 等ICLR 2021 · 被引用 17 次
- Sequencer: Deep LSTM for Image ClassificationYuki Tatsunami, Masato TakiNeurIPS 2022 · 被引用 124 次
- Rethinking and Improving Relative Position Encoding for Vision TransformerKan Wu, Houwen Peng, Minghao Chen, Jianlong Fu 等ICCV 2021 · 被引用 427 次
- Scalable Vision Transformers with Hierarchical PoolingZizheng Pan, Bohan Zhuang, Jing Liu, Haoyu He 等ICCV 2021 · 被引用 154 次
- Learnable Fourier Features for Multi-dimensional Spatial Positional EncodingYang Li, Si Si, Gang Li, Cho-Jui Hsieh 等NeurIPS 2021 · 被引用 171 次
