Scalable Set Encoding with Universal Mini-Batch Consistency and Unbiased Full Set Gradient Approximation
Jeffrey Willette, Seanie Lee, Bruno Andreis, Kenji Kawaguchi, Juho Lee, Sung Ju Hwang
Abstract
Recent work on mini-batch consistency (MBC) for set functions has brought attention to the need for sequentially processing and aggregating chunks of a partitioned set while guaranteeing the same output for all partitions. However, existing constraints on MBC architectures lead to models with limited expressive power. Additionally, prior work has not addressed how to deal with large sets during training when the full set gradient is required. To address these issues, we propose a Universally MBC (UMBC) class of set functions which can be used in conjunction with arbitrary non-MBC components while still satisfying MBC, enabling a wider range of function classes to be used in MBC settings. Furthermore, we propose an efficient MBC training algorithm which gives an unbiased approximation of the full set gradient and has a constant memory overhead for any set size for both train-and test-time. We conduct extensive experiments including image completion, text classification, unsupervised clustering, and cancer detection on high-resolution images to verify the efficiency and efficacy of our scalable set encoding framework. Our code is available at github.com/jeffwillette/umbc * Equal contribution 1 KAIST
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c9da5e10-82ea-49fe-b197-b247882d3bf4Cited by top-tier papers3
- Set-based Neural Network Encoding Without Weight TyingBruno Andreis, Bedionita Soro, Philip H. S. Torr, Sung Ju HwangNeurIPS 2024 · 8 citations
- Monotone and Separable Set Functions: Characterizations and Neural ModelsSoutrik Sarangi, Yonatan Sverdlov, Nadav Dym, Abir DeNeurIPS 2025 · 2 citations
- HORSE: Hierarchical Representation for Large-Scale Neural Subset SelectionBinghui Xie, Yixuan Wang, Yongqiang Chen, Kaiwen Zhou et al.NeurIPS 2024 · 1 citation
Builds on8
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- On the Almost Sure Convergence of Stochastic Gradient Descent in Non-Convex ProblemsPanayotis Mertikopoulos, Nadav Hallak, Ali Kavis, Volkan CevherNeurIPS 2020 · 120 citations
- FSPool: Learning Set Representations with Featurewise Sort PoolingYan Zhang, Jonathon S. Hare, Adam Prügel-BennettICLR 2020 · 92 citations
- A Trainable Optimal Transport Embedding for Feature Aggregation and its Relationship to AttentionGrégoire Mialon, Dexiong Chen, Alexandre d'Aspremont, Julien MairalICLR 2021 · 71 citations
Related papers
- Mini-Batch Consistent Slot Set Encoder for Scalable Set EncodingAndreis Bruno, Jeffrey Willette, Juho Lee, Sung Ju HwangNeurIPS 2021 · 10 citations
- Weakly Supervised Clustering by Exploiting Unique Class CountMustafa Umit Oner, Hwee Kuan Lee, Wing-Kin SungICLR 2020 · 8 citations
- Binary Classification from Multiple Unlabeled Datasets via Surrogate Set ClassificationNan Lu, Shida Lei, Gang Niu, Issei Sato et al.ICML 2021 · 17 citations
- ABC: Auxiliary Balanced Classifier for Class-imbalanced Semi-supervised LearningHyuck Lee, Seungjae Shin, Heeyoung KimNeurIPS 2021 · 131 citations
- Set2Graph: Learning Graphs From SetsHadar Serviansky, Nimrod Segol, Jonathan Shlomi, Kyle Cranmer et al.NeurIPS 2020 · 37 citations
