BatchFormer: Learning to Explore Sample Relationships for Robust Representation Learning
Zhi Hou, Baosheng Yu, Dacheng Tao
Abstract
Despite the success of deep neural networks, there are still many challenges in deep representation learning due to the data scarcity issues such as data imbalance, unseen distribution, and domain shift. To address the above-mentioned issues, a variety of methods have been devised to explore the sample relationships in a vanilla way (i.e., from the perspectives of either the input or the loss function), failing to explore the internal structure of deep neural networks for learning with sample relationships. Inspired by this, we propose to enable deep neural networks themselves with the ability to learn the sample relationships from each mini-batch. Specifically, we introduce a batch transformer module or BatchFormer, which is then applied into the batch dimension of each mini-batch to implicitly explore sample relationships during training. By doing this, the proposed method enables the collaboration of different samples, e.g., the head-class samples can also contribute to the learning of the tail classes for long-tailed recognition. Furthermore, to mitigate the gap between training and testing, we share the classifier between with or without the BatchFormer during training, which can thus be removed during testing. We perform extensive experiments on over ten datasets and the proposed method achieves significant improvements on different data scarcity applications without any bells and whistles, including the tasks of long-tailed recognition, compositional zero-shot learning, domain generalization, and contrastive learning. Code is made publicly available at https://github.com/zhihou7/BatchFormer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers14
- SFC: Shared Feature Calibration in Weakly Supervised Semantic SegmentationXinqiao Zhao, Feilong Tang, Xiaoyang Wang, Jimin XiaoAAAI 2024 · 66 citations
- Enhancing Minority Classes by Mixing: An Adaptative Optimal Transport Approach for Long-tailed ClassificationJintong Gao, He Zhao, Zhuo Li, Dandan GuoNeurIPS 2023 · 64 citations
- Decoupled Contrastive Learning for Long-Tailed RecognitionShiyu Xuan, Shiliang ZhangAAAI 2024 · 29 citations
- Catalyst for Clustering-Based Unsupervised Object Re-identification: Feature CalibrationHuafeng Li, Qingsong Hu, Zhanxuan HuAAAI 2024 · 27 citations
- Class-level Structural Relation Modeling and Smoothing for Visual Representation LearningZitan Chen, Zhuang Qi, Xiao Cao, Xiangxian Li et al.ACM MM 2023 · 10 citations
Builds on32
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan et al.ICLR 2020 · 1,496 citations
Related papers
- Interpolation Normalization for Contrast Domain GeneralizationMengzhu Wang, Junyang Chen, Huan Wang, Huisi Wu et al.ACM MM 2023 · 5 citations
- RetFormer: Multimodal Retrieval for Enhancing Image RecognitionTianrui Yu, Xiubo Liang, Hongzhi WangCVPR 2026
- Gait Transformer: End-to-End Transformer Backbone for Gait RecognitionSaihui Hou, Wenpeng Lang, Jilong Wang, Yan Huang et al.AAAI 2026
- FCC: Feature Clusters Compression for Long-Tailed Visual RecognitionJian Li, Ziyao Meng, Daqian Shi, Rui Song et al.CVPR 2023
- MedSpaformer: A Transferable Transformer with Multi-Granularity Token Sparsification for Medical Time Series ClassificationJiexia Ye, Weiqi Zhang, Ziyue Li, Jia Li et al.AAAI 2026 · 1 citation
