Lune

SIGMOD2020Top-tier venue

Finding Related Tables in Data Lakes for Interactive Data Science

Yi Zhang, Zachary G. Ives

2020Year
98Citations
32Top-tier citations

Abstract

Many modern data science applications build on data lakes, schema-agnostic repositories of data files and data products that offer limited organization and management capabilities. There is a need to build data lake search capabilities into data science environments, so scientists and analysts can find tables, schemas, workflows, and datasets useful to their task at hand. We develop search and management solutions for the Jupyter Notebook data science platform, to enable scientists to augment training data, find potential features to extract, clean data, and find joinable or linkable tables. Our core methods also generalize to other settings where computational tasks involve execution of programs or scripts.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext b0cd7e4b-860a-4997-a6ab-dbed1199a415

Cited by top-tier papers32

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines