HORSE: Hierarchical Representation for Large-Scale Neural Subset Selection
Binghui Xie, Yixuan Wang, Yongqiang Chen, Kaiwen Zhou, Yu Li, Wei Meng, James Cheng
Abstract
Subset selection tasks, such as anomaly detection and compound selection in AI-assisted drug discovery, are crucial for a wide range of applications. Learning subset-valued functions with neural networks has achieved great success by incorporating permutation invariance symmetry into the architecture. However, existing neural set architectures often struggle to either capture comprehensive information from the superset or address complex interactions within the input. Additionally, they often fail to perform in scenarios where superset sizes surpass available memory capacity. To address these challenges, we introduce the novel concept of the Identity Property , which requires models to integrate information from the originating set, resulting in the development of neural networks that excel at performing effective subset selection from large supersets. Moreover, we present the Hierarchical Representation of Neural Subset Selection (HORSE), an attention-based method that learns complex interactions and retains information from both the input set and the optimal subset supervision signal. Specifically, HORSE enables the partitioning of the input ground set into manageable chunks that can be processed independently and then aggregated, ensuring consistent outcomes across different partitions. Through extensive experimentation, we demonstrate that HORSE significantly enhances neural subset selection performance by capturing more complex information and surpasses the state-of-the-art methods in handling large-scale inputs by a margin of up to 20%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 129eb4b1-4b0f-4c4a-9a7c-5c26c42ae6cdBuilds on17
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 494 citations
- Equivariant Subgraph Aggregation NetworksBeatrice Bevilacqua, Fabrizio Frasca, Derek Lim, Balasubramaniam Srinivasan et al.ICLR 2022 · 217 citations
- Set Functions for Time SeriesMax Horn, Michael Moor, Christian Bock, Bastian Rieck et al.ICML 2020 · 199 citations
- Structure-aware Interactive Graph Neural Networks for the Prediction of Protein-Ligand Binding AffinityShuangli Li, Jingbo Zhou, Tong Xu, Liang Huang et al.KDD 2021 · 184 citations
- On Learning Sets of Symmetric ElementsHaggai Maron, Or Litany, Gal Chechik, Ethan FetayaICML 2020 · 148 citations
Related papers
- Enhancing Neural Subset Selection: Integrating Background Information into Set RepresentationsBinghui Xie, Yatao Bian, Kaiwen Zhou, Yongqiang Chen et al.ICLR 2024 · 1 citation
- Learning Neural Set Functions Under the Optimal Subset OracleZijing Ou, Tingyang Xu, Qinliang Su, Yingzhen Li et al.NeurIPS 2022 · 13 citations
- Learning Set Functions with Implicit DifferentiationGözde Özcan, Chengzhi Shi, Stratis IoannidisAAAI 2025
- Efficient Data Subset Selection to Generalize Training Across Models: Transductive and Inductive NetworksEeshaan Jain, Tushar Nandy, Gaurav Aggarwal, Ashish Tendulkar et al.NeurIPS 2023 · 30 citations
- Graph Neural Networks with Adaptive ReadoutsDavid Buterez, Jon Paul Janet, Steven J. Kiddle, Dino Oglic et al.NeurIPS 2022 · 81 citations
