Lune

ICML2026Top-tier venue

Deep Ensemble Clustering for Visual Representation Learning

Yuwei Wang, Guikun Chen, Xiruo Jiang, Yazhou Yao, Di Liu, Xiangbo Shu, Fumin Shen, Wenguan Wang

2026Year

Abstract

Recent advances in visual representation learning have seen the rise of clustering-based vision backbones, which adopt clustering as a core paradigm for feature extraction. However, existing clustering-based backbones typically rely on a single clustering algorithm, whose inherent inductive bias limits their representational capacity. To address this, we propose EnFormer, which embeds ensemble clustering as a core component of feature extraction. EnFormer structures feature extraction around two steps: (i) Ensemble Generation, where several differentiable base clustering methods are introduced to capture diverse semantic structures; and (ii) Consensus Aggregation, which employs a differentiable mechanism to fuse the results of all base clusterings to reconstruct refined visual features. Extensive experiments show that EnFormer consistently outperforms existing clustering-based backbones across core vision tasks, with higher performance and significantly improved throughput.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext fc022d5e-da70-4639-bdf1-300b8916af4f

Builds on30

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines