Metam: Goal-Oriented Data Discovery
Sainyam Galhotra, Yue Gong, Raul Castro Fernandez
Abstract
Data is a central component of machine learning and causal inference tasks. The availability of large amounts of data from sources such as open data repositories, data lakes and data marketplaces creates an opportunity to augment data and boost those tasks' performance. However, augmentation techniques rely on a user manually discovering and shortlisting useful candidate augmentations. Existing solutions do not leverage the synergy between discovery and augmentation, thus underexploiting data.
In this paper, we introduce METAM, a novel goal-oriented framework that queries the downstream task with a candidate dataset, forming a feedback loop that automatically steers the discovery and augmentation process. To select candidates efficiently, METAM leverages properties of the: i) data, ii) utility function, and iii) solution set size. We show METAM's theoretical guarantees and demonstrate those empirically on a broad set of tasks. All in all, we demonstrate the promise of goal-oriented data discovery to modern data science applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 353ffa00-6d25-4c7b-8aa0-e92a299785b7Cited by top-tier papers19
- Pneuma: Leveraging LLMs for Tabular Data Representation and Retrieval in an End-to-End SystemMuhammad Imam Luthfi Balaka, David Alexander, Qiming Wang, Yue Gong et al.SIGMOD 2025 · 12 citations
- Nexus: Correlation Discovery over Collections of Spatio-Temporal Tabular DataYue Gong, Sainyam Galhotra, Raul Castro FernandezSIGMOD 2024 · 10 citations
- Fainder: A Fast and Accurate Index for Distribution-Aware Dataset SearchLennart Behme, Sainyam Galhotra, Kaustubh Beedkar, Volker MarklVLDB 2024 · 9 citations
- Qualitative Join Discovery in Data Lakes using ExamplesMir Mahathir Mohammad, El Kindi RezigSIGMOD 2026 · 6 citations
- Gen-T: Table Reclamation in Data LakesGrace Fan, Roee Shraga, Renée J. MillerICDE 2024 · 5 citations
Builds on8
- Finding Related Tables in Data Lakes for Interactive Data ScienceYi Zhang, Zachary G. IvesSIGMOD 2020 · 98 citations
- Selective Data Acquisition in the Wild for Model ChargingChengliang Chai, Jiabin Liu, Nan Tang, Guoliang Li et al.VLDB 2022 · 62 citations
- Correlation Sketches for Approximate Join-Correlation QueriesAécio S. R. Santos, Aline Bessa, Fernando Chirigati, Christopher Musco et al.SIGMOD 2021 · 45 citations
- Automated Feature Engineering for Algorithmic FairnessRicardo Salazar, Felix Neutatz, Ziawasch AbedjanVLDB 2021 · 42 citations
- Leva: Boosting Machine Learning Performance with Relational Embedding Data AugmentationZixuan Zhao, Raul Castro FernandezSIGMOD 2022 · 19 citations
Related papers
- Meta Learning for Causal DirectionJean-François Ton, Dino Sejdinovic, Kenji FukumizuAAAI 2021 · 26 citations
- Matryoshka: Uncovering Relevant Features in Data Lakes to Enhance Machine Learning ApplicationsFedor Turchenko, Runjie Zhang, Binger Chen, Matthias Boehm et al.VLDB 2026
- Trust Your 𝛁: Gradient-based Intervention Targeting for Causal DiscoveryMateusz Olko, Michal Zajac, Aleksandra Nowak, Nino Scherrer et al.NeurIPS 2023
- An Analysis of Causal Effect Estimation using Outcome Invariant Data AugmentationUzair Akbar, Niki Kilbertus, Hao Shen, Krikamol Muandet et al.NeurIPS 2025 · 3 citations
- Automatic Auxiliary Task Selection and Adaptive Weighting Boost Molecular Property PredictionZhiqiang Zhong, Davide MottinNeurIPS 2025 · 3 citations
