Channel Vision Transformers: An Image Is Worth 1 x 16 x 16 Words
Yujia Bao, Srinivasan Sivanandan, Theofanis Karaletsos
Abstract
Vision Transformer (ViT) has emerged as a powerful architecture in the realm of modern computer vision. However, its application in certain imaging fields, such as microscopy and satellite imaging, presents unique challenges. In these domains, images often contain multiple channels, each carrying semantically distinct and independent information. Furthermore, the model must demonstrate robustness to sparsity in input channels, as they may not be densely available during training or testing. In this paper, we propose a modification to the ViT architecture that enhances reasoning across the input channels and introduce Hierarchical Channel Sampling (HCS) as an additional regularization technique to ensure robustness when only partial channels are presented during test time. Our proposed model, ChannelViT, constructs patch tokens independently from each input channel and utilizes a learnable channel embedding that is added to the patch tokens, similar to positional embeddings. We evaluate the performance of ChannelViT on ImageNet, JUMP-CP (microscopy cell imaging), and So2Sat (satellite imaging). Our results show that ChannelViT outperforms ViT on classification tasks and generalizes well, even when a subset of input channels is used during testing. Across our experiments, HCS proves to be a powerful regularizer, independent of the architecture employed, suggesting itself as a straightforward technique for robust ViT training. Lastly, we find that ChannelViT generalizes effectively even when there is limited access to all channels during training, highlighting its potential for multi-channel imaging under real-world conditions with sparse sensors. Our code is available at https://github.com/insitro/ChannelViT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Enhancing Feature Diversity Boosts Channel-Adaptive Vision TransformersChau Pham, Bryan A. PlummerNeurIPS 2024 · 15 citations
- ChA-MAEViT: Unifying Channel-Aware Masked Autoencoders and Multi-Channel Vision Transformers for Improved Cross-Channel LearningChau Pham, Juan C. Caicedo, Bryan A. PlummerNeurIPS 2025 · 11 citations
- CellCLIP - Learning Perturbation Effects in Cell Painting via Text-Guided Contrastive LearningMingyu Lu, Ethan Weinberger, Chanwoo Kim, Su-In LeeNeurIPS 2025 · 9 citations
- CHAMMI-75: Pre-training multi-channel models with heterogeneous microscopy imagesVidit Agrawal, John Peters, Tyler N. Thompson, Mohammad V. Sanian et al.ICLR 2026 · 3 citations
- ViTally Consistent: Scaling Biological Representation Learning for Cell MicroscopyKian Kenyon-Dean, Zitong Jerry Wang, John Urbanik, Konstantin Donhauser et al.ICML 2025
Builds on15
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
Related papers
- Scale-space Tokenization for Improving the Robustness of Vision TransformersLei Xu, Rei Kawakami, Nakamasa InoueACM MM 2023 · 1 citation
- ChAda-ViT : Channel Adaptive Attention for Joint Representation Learning of Heterogeneous Microscopy ImageNicolas Bourriez, Ihab Bendidi, Ethan Cohen, Gabriel Watkinson et al.CVPR 2024
- Discrete Representations Strengthen Vision Transformer RobustnessChengzhi Mao, Lu Jiang, Mostafa Dehghani, Carl Vondrick et al.ICLR 2022 · 47 citations
- ViTs for SITS: Vision Transformers for Satellite Image Time SeriesMichail Tarasiou, Erik Chavez, Stefanos ZafeiriouCVPR 2023
- Scalable Vision Transformers with Hierarchical PoolingZizheng Pan, Bohan Zhuang, Jing Liu, Haoyu He et al.ICCV 2021 · 154 citations
