Learn to Explore: on Bootstrapping Interactive Data Exploration with Meta-learning
Yukun Cao, Xike Xie, Kexin Huang
Abstract
Interactive data exploration (IDE) is an effective way of comprehending big data, whose volume and complexity are beyond human abilities. The main goal of IDE is to discover user interest regions from a database through multi-rounds of user labelling. Existing IDEs adopt active-learning framework, where users iteratively discriminate or label the interestingness of selected tuples. The process of data exploration can be viewed as the process of training of a classifier, which determines whether a database tuple is interesting to a user. An efficient exploration thus takes very few iterations of user labelling to reach the data region of interest. In this work, we consider the data exploration as the process of few-shot learning, where the classifier is learned with only a few training examples, or exploration iterations. To this end, we propose a learning-to-explore framework, based on meta-learning, which learns how to learn a classifier with automatically generated meta-tasks, so that the exploration process can be much shortened. Extensive experiments on real datasets show that our proposal outperforms existing explore-by-example solutions in terms of accuracy and efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc0a763b-985b-4f2c-b236-348c7c0ff506Builds on5
- MAMO: Memory-Augmented Meta-Optimization for Cold-start RecommendationManqing Dong, Feng Yuan, Lina Yao, Xiwei Xu et al.KDD 2020 · 161 citations
- MetaInsight: Automatic Discovery of Structured Knowledge for Exploratory Data AnalysisPingchuan Ma, Rui Ding, Shi Han, Dongmei ZhangSIGMOD 2021 · 35 citations
- Efficient Exploration of Interesting Aggregates in RDF GraphsYanlei Diao, Pawel Guzewicz, Ioana Manolescu, Mirjana MazuranSIGMOD 2021 · 5 citations
- Multi-dimensional Probabilistic Regression over Imprecise Data StreamsRan Gao, Xike Xie, Kai Zou, Torben Bach PedersenWWW 2022 · 5 citations
- Relational Data Synthesis using Generative Adversarial Networks: A Design Space ExplorationJu Fan, Tongyu Liu, Guoliang Li, Junyou Chen et al.VLDB 2020
Related papers
- Fast Search-By-Classification for Large-Scale Databases Using Index-Aware Decision Trees and Random ForestsChristian Lülf, Denis Mayr Lima Martins, Marcos Antonio Vaz Salles, Yongluan Zhou et al.VLDB 2023 · 5 citations
- Meta-RCNN: Meta Learning for Few-Shot Object DetectionXiongwei Wu, Doyen Sahoo, Steven C. H. HoiACM MM 2020 · 94 citations
- SeLeP: Learning Based Semantic Prefetching for Exploratory Database WorkloadsFarzaneh Zirak, Farhana Murtaza Choudhury, Renata Borovica-GajicVLDB 2024 · 4 citations
- Exploring Task Difficulty for Few-Shot Relation ExtractionJiale Han, Bo Cheng, Wei LuEMNLP 2021 · 74 citations
- Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few ExamplesEleni Triantafillou, Tyler Zhu, Vincent Dumoulin, Pascal Lamblin et al.ICLR 2020 · 692 citations
