Rethinking Dataset Discovery with DataScout
Rachel Lin, Bhavya Chopra, Wenjing Lin, Shreya Shankar, Madelon Hulsebos, Aditya G. Parameswaran
摘要
Dataset Search-the process of finding appropriate datasets for a given task-remains a critical yet under-explored challenge in data science workflows. Assessing dataset suitability for a task (e.g., training a classification model) is a multi-pronged affair that involves understanding: data characteristics (e.g. granularity, attributes, size), semantics (e.g., data semantics, creation goals), and relevance to the task at hand. Present-day dataset search interfaces are restrictiveusers struggle to convey implicit preferences and lack visibility into the search space and result inclusion criteria-making query iteration challenging. To bridge these gaps, we introduce DataScout to proactively steer users through the process of dataset discovery via-(i) AI-assisted query reformulations informed by the underlying search space, (ii) semantic search and filtering based on dataset content, including attributes (columns) and granularity (rows), and (iii) dataset relevance indicators, generated dynamically based on the user-specified task. A within-subjects study with 12 participants comparing DataScout to keyword and semantic dataset search reveals that users uniquely employ DataScout's features not only for structured explorations, but also to glean feedback on their search queries and build conceptual models of the search space.
• Human-centered computing → Systems and tools for interaction design; • Information systems → Search interfaces; Collaborative search.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Sensecape: Enabling Multilevel Exploration and Sensemaking with Large Language ModelsSangho Suh, Bryan Min, Srishti Palani, Haijun XiaUIST 2023 · 被引用 147 次
- Luminate: Structured Generation and Exploration of Design Space with Large Language Models for Human-AI Co-CreationSangho Suh, Meng Chen, Bryan Min, Toby Jia-Jun Li 等CHI 2024 · 被引用 143 次
- Semantics-aware Dataset Discovery from Data Lakes with Contextualized Column-based Representation LearningGrace Fan, Jin Wang, Yuliang Li, Dan Zhang 等VLDB 2023 · 被引用 139 次
- Leading Conversational Search by Suggesting Useful QuestionsCorbin Rosset, Chenyan Xiong, Xia Song, Daniel Campos 等WWW 2020 · 被引用 85 次
- SANTOS: Relationship-based Semantic Table Union SearchAamod Khatiwada, Grace Fan, Roee Shraga, Zixuan Chen 等SIGMOD 2023 · 被引用 61 次
相关 Paper
- DataFinder: Scientific Dataset Recommendation from Natural Language DescriptionsVijay Viswanathan, Luyu Gao, Tongshuang Wu, Pengfei Liu 等ACL 2023 · 被引用 9 次
- DataPilot: Utilizing Quality and Usage Information for Subset Selection during Visual Data PreparationArpit Narechania, Fan Du, Atanu R. Sinha, Ryan A. Rossi 等CHI 2023 · 被引用 13 次
- PhotoScout: Synthesis-Powered Multi-Modal Image SearchCeleste Barnaby, Qiaochu Chen, Chenglong Wang, Isil DilligCHI 2024 · 被引用 9 次
- Revisiting Task-Oriented Dataset Search in the Era of Large Language Models: Challenges, Benchmark, and SolutionZixin Wei, Yucan Guo, Jinyang Li, Xiaolin Han 等VLDB 2026 · 被引用 1 次
- Enhancing Dataset Search with Compact Data SnippetsQiaosheng Chen, Jiageng Chen, Xiao Zhou, Gong ChengSIGIR 2024 · 被引用 5 次
