ChA-MAEViT: Unifying Channel-Aware Masked Autoencoders and Multi-Channel Vision Transformers for Improved Cross-Channel Learning
Chau Pham, Juan C. Caicedo, Bryan A. Plummer
Abstract
Prior work using Masked Autoencoders (MAEs) typically relies on random patch masking based on the assumption that images have significant redundancies across different channels, allowing for the reconstruction of masked content using cross-channel correlations. However, this assumption does not hold in Multi-Channel Imaging (MCI), where channels may provide complementary information with minimal feature overlap. Thus, these MAEs primarily learn local structures within individual channels from patch reconstruction, failing to fully leverage cross-channel interactions and limiting their MCI effectiveness. In this paper, we present ChA-MAEViT, an MAE-based method that enhances feature learning across MCI channels via four key strategies: (1) dynamic channel-patch masking, which compels the model to reconstruct missing channels in addition to masked patches, thereby enhancing cross-channel dependencies and improving robustness to varying channel configurations; (2) memory tokens, which serve as long-term memory aids to promote information sharing across channels, addressing the challenges of reconstructing structurally diverse channels; (3) hybrid token fusion module, which merges fine-grained patch tokens with a global class token to capture richer representations; and (4) Channel-Aware Decoder, a lightweight decoder utilizes channel tokens to effectively reconstruct image patches. Experiments on satellite and microscopy datasets, CHAMMI, JUMP-CP, and So2Sat, show that ChA-MAEViT significantly outperforms state-of-the-art MCI-ViTs by 3.0-21.5%, highlighting the importance of cross-channel interactions in MCI. Our code is publicly available at https://github.com/chaudatascience/cha_mae_vit.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 30893be6-41d5-423c-b30b-b157d1b8efe0Cited by top-tier papers2
- CHAMMI-75: Pre-training multi-channel models with heterogeneous microscopy imagesVidit Agrawal, John Peters, Tyler N. Thompson, Mohammad V. Sanian et al.ICLR 2026 · 3 citations
- PETRI: Learning Unified Cell Embeddings from Unpaired Modalities via Early-Fusion Joint ReconstructionRyan W Conrad, Ethan Weinberger, Saradha Venkatachalapathy, Yuwen Chen et al.ICLR 2026
Builds on23
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
Related papers
- Enhancing Feature Diversity Boosts Channel-Adaptive Vision TransformersChau Pham, Bryan A. PlummerNeurIPS 2024 · 15 citations
- Channel Vision Transformers: An Image Is Worth 1 x 16 x 16 WordsYujia Bao, Srinivasan Sivanandan, Theofanis KaraletsosICLR 2024 · 47 citations
- Masked Autoencoders for Microscopy are Scalable Learners of Cellular BiologyOren Kraus, Kian Kenyon-Dean, Saber Saberian, Maryam Fallah et al.CVPR 2024
- The Dynamic Duo of Collaborative Masking and Target for Advanced Masked Autoencoder LearningShentong MoAAAI 2025 · 1 citation
- IMTS is Worth Time × Channel Patches: Visual Masked Autoencoders for Irregular Multivariate Time Series PredictionZhangyi Hu, Jiemin Wu, Hua Xu, Mingqian Liao et al.ICML 2025
