Rethinking Dataset Discovery with DataScout
Rachel Lin, Bhavya Chopra, Wenjing Lin, Shreya Shankar, Madelon Hulsebos, Aditya G. Parameswaran
Abstract
Dataset Search-the process of finding appropriate datasets for a given task-remains a critical yet under-explored challenge in data science workflows. Assessing dataset suitability for a task (e.g., training a classification model) is a multi-pronged affair that involves understanding: data characteristics (e.g. granularity, attributes, size), semantics (e.g., data semantics, creation goals), and relevance to the task at hand. Present-day dataset search interfaces are restrictiveusers struggle to convey implicit preferences and lack visibility into the search space and result inclusion criteria-making query iteration challenging. To bridge these gaps, we introduce DataScout to proactively steer users through the process of dataset discovery via-(i) AI-assisted query reformulations informed by the underlying search space, (ii) semantic search and filtering based on dataset content, including attributes (columns) and granularity (rows), and (iii) dataset relevance indicators, generated dynamically based on the user-specified task. A within-subjects study with 12 participants comparing DataScout to keyword and semantic dataset search reveals that users uniquely employ DataScout's features not only for structured explorations, but also to glean feedback on their search queries and build conceptual models of the search space.
• Human-centered computing → Systems and tools for interaction design; • Information systems → Search interfaces; Collaborative search.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 242f9d15-15b7-4065-a206-aa97f875efc8Builds on11
- Sensecape: Enabling Multilevel Exploration and Sensemaking with Large Language ModelsSangho Suh, Bryan Min, Srishti Palani, Haijun XiaUIST 2023 · 147 citations
- Luminate: Structured Generation and Exploration of Design Space with Large Language Models for Human-AI Co-CreationSangho Suh, Meng Chen, Bryan Min, Toby Jia-Jun Li et al.CHI 2024 · 143 citations
- Semantics-aware Dataset Discovery from Data Lakes with Contextualized Column-based Representation LearningGrace Fan, Jin Wang, Yuliang Li, Dan Zhang et al.VLDB 2023 · 139 citations
- Leading Conversational Search by Suggesting Useful QuestionsCorbin Rosset, Chenyan Xiong, Xia Song, Daniel Campos et al.WWW 2020 · 85 citations
- SANTOS: Relationship-based Semantic Table Union SearchAamod Khatiwada, Grace Fan, Roee Shraga, Zixuan Chen et al.SIGMOD 2023 · 61 citations
Related papers
- DataFinder: Scientific Dataset Recommendation from Natural Language DescriptionsVijay Viswanathan, Luyu Gao, Tongshuang Wu, Pengfei Liu et al.ACL 2023 · 9 citations
- DataPilot: Utilizing Quality and Usage Information for Subset Selection during Visual Data PreparationArpit Narechania, Fan Du, Atanu R. Sinha, Ryan A. Rossi et al.CHI 2023 · 13 citations
- PhotoScout: Synthesis-Powered Multi-Modal Image SearchCeleste Barnaby, Qiaochu Chen, Chenglong Wang, Isil DilligCHI 2024 · 9 citations
- Revisiting Task-Oriented Dataset Search in the Era of Large Language Models: Challenges, Benchmark, and SolutionZixin Wei, Yucan Guo, Jinyang Li, Xiaolin Han et al.VLDB 2026 · 1 citation
- Enhancing Dataset Search with Compact Data SnippetsQiaosheng Chen, Jiageng Chen, Xiao Zhou, Gong ChengSIGIR 2024 · 5 citations
