Neural Clustering Based Visual Representation Learning
Guikun Chen, Xia Li, Yi Yang, Wenguan Wang
Abstract
We investigate a fundamental aspect of machine vision: the measurement of features, by revisiting clustering, one of the most classic approaches in machine learning and data analysis. Existing visual feature extractors, including Conv-Nets, ViTs, and MLPs, represent an image as rectangular regions. Though prevalent, such a grid-style paradigm is built upon engineering practice and lacks explicit modeling of data distribution. In this work, we propose feature extraction with clustering (FEC), a conceptually elegant yet surprisingly ad-hoc interpretable neural clustering framework, which views feature extraction as a process of selecting representatives from data and thus automatically captures the underlying data distribution. Given an image, FEC alternates between grouping pixels into individual clusters to abstract representatives and updating the deep features of pixels with current representatives. Such an iterative working mechanism is implemented in the form of several neural layers and the final representatives can be used for downstream tasks. The cluster assignments across layers, which can be viewed and inspected by humans, make the forward process of FEC fully transparent and empower it with promising ad-hoc interpretability. Extensive experiments on various visual recognition models and tasks verify the effectiveness, generality, and interpretability of FEC. We expect this work will provoke a rethink of the current de facto grid-style paradigm.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a4b71601-87e1-4220-a893-02f1e42d7beaCited by top-tier papers11
- Clustering for Protein Representation LearningRuijie Quan, Wenguan Wang, Fan Ma, Hehe Fan et al.CVPR 2024 · 10 citations
- Reconciling Visual Perception and Generation in Diffusion ModelsLiulei Li, Yi Yang, Wenguan WangICLR 2026
- DVHGNN: Multi-Scale Dilated Vision HGNN for Efficient Vision RecognitionCaoshuo Li, Tanzhe Li, Xiaobin Hu, Donghao Luo et al.CVPR 2025
- Enhancing Pre-trained Representation Classifiability can Boost its InterpretabilityShufan Shen, Zhaobo Qi, Junshu Sun, Qingming Huang et al.ICLR 2025
- Learning Clustering-based Prototypes for Compositional Zero-Shot LearningHongyu Qu, Jianan Wei, Xiangbo Shu, Wenguan WangICLR 2025
Builds on34
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
Related papers
- Image as Set of PointsXu Ma, Yuqian Zhou, Huan Wang, Can Qin et al.ICLR 2023 · 221 citations
- CLUENet: Cluster Attention Makes Neural Networks Have EyesXiangshuai Song, Jun-Jie Huang, Tianrui Liu, Ke Liang et al.AAAI 2026
- Deep Ensemble Clustering for Visual Representation LearningYuwei Wang, Guikun Chen, Xiruo Jiang, Yazhou Yao et al.ICML 2026
- Dynamic Clustering Convolutional Neural NetworkTanzhe Li, Baochang Zhang, Jiayi Lyu, Xiawu Zheng et al.AAAI 2025
- PICNN: A Pathway towards Interpretable Convolutional Neural NetworksWengang Guo, Jiayi Yang, Huilin Yin, Qijun Chen et al.AAAI 2024 · 6 citations
