Learning Representations without Compositional Assumptions
Tennison Liu, Jeroen Berrevoets, Zhaozhi Qian, Mihaela van der Schaar
Abstract
This paper addresses unsupervised representation learning on tabular data containing multiple views generated by distinct sources of measurement. Traditional methods, which tackle this problem using the multi-view framework, are constrained by predefined assumptions that assume feature sets share the same information and representations should learn globally shared factors. However, this assumption is not always valid for real-world tabular datasets with complex dependencies between feature sets, resulting in localized information that is harder to learn. To overcome this limitation, we propose a data-driven approach that learns feature set dependencies by representing feature sets as graph nodes and their relationships as learnable edges. Furthermore, we introduce LEGATO, a novel hierarchical graph autoencoder that learns a smaller, latent graph to aggregate information from multiple views dynamically. This approach results in latent graph components that specialize in capturing localized information from different regions of the input, leading to superior downstream performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Learning Robust Representations via Multi-View Information BottleneckMarco Federici, Anjan Dutta, Patrick Forré, Nate Kushman et al.ICLR 2020 · 330 citations
- Graph Neural Networks with Adaptive ReadoutsDavid Buterez, Jon Paul Janet, Steven J. Kiddle, Dino Oglic et al.NeurIPS 2022 · 81 citations
- On the Limitations of Multimodal VAEsImant Daunhawer, Thomas M. Sutter, Kieran Chin-Cheong, Emanuele Palumbo et al.ICLR 2022 · 50 citations
- BayReL: Bayesian Relational Learning for Multi-omics Data IntegrationEhsan Hajiramezanali, Arman Hasanzadeh, Nick Duffield, Krishna Narayanan et al.NeurIPS 2020 · 14 citations
Related papers
- SubTab: Subsetting Features of Tabular Data for Self-Supervised Representation LearningTalip Ucar, Ehsan Hajiramezanali, Lindsay EdwardsNeurIPS 2021 · 189 citations
- C-iVAE: Causal Structure-Aware Identifiable VAE for Tabular Data with Limited EnvironmentsZejiang Wang, Feng ZhouKDD 2026
- XTab: Cross-table Pretraining for Tabular TransformersBingzhao Zhu, Xingjian Shi, Nick Erickson, Mu Li et al.ICML 2023 · 111 citations
- Improving Graph Learning-Based Fault Localization with Tailored Semi-supervised LearningChun Li, Hui Li, Zhong Li, Minxue Pan et al.FSE 2025 · 2 citations
- Aggregate to Adapt: Node-Centric Aggregation for Multi-Source-Free Graph Domain AdaptationZhen Zhang, Bingsheng HeWWW 2025 · 6 citations
