Solo: Data Discovery Using Natural Language Questions Via A Self-Supervised Approach
Qiming Wang, Raul Castro Fernandez
摘要
Most deployed data discovery systems, such as Google Datasets, and open data portals only support keyword search. Keyword search is geared towards general audiences but limits the types of queries the systems can answer. We propose a new system that lets users write natural language questions directly. A major barrier to using this learned data discovery system is it needs expensive-to-collect training data, thus limiting its utility. In this paper, we introduce a self-supervised approach to assemble training datasets and train learned discovery systems without human intervention. It requires addressing several challenges, including the design of self-supervised strategies for data discovery, table representation strategies to feed to the models, and relevance models that work well with the synthetically generated questions. We combine all the above contributions into a system, Solo, that solves the problem end to end. The evaluation results demonstrate the new techniques outperform state-of-the-art approaches on well-known benchmarks. All in all, the technique is a stepping stone towards building learned discovery systems. CCS Concepts: • Information systems → Structured text search.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Pneuma: Leveraging LLMs for Tabular Data Representation and Retrieval in an End-to-End SystemMuhammad Imam Luthfi Balaka, David Alexander, Qiming Wang, Yue Gong 等SIGMOD 2025 · 被引用 12 次
- Featurized-Decomposition Join: Low-Cost Semantic Joins with GuaranteesSepanta Zeighami, Shreya Shankar, Aditya G. ParameswaranVLDB 2026 · 被引用 11 次
- Fainder: A Fast and Accurate Index for Distribution-Aware Dataset SearchLennart Behme, Sainyam Galhotra, Kaustubh Beedkar, Volker MarklVLDB 2024 · 被引用 9 次
- BIRDIE: Natural Language-Driven Table Discovery Using Differentiable Search IndexYuxiang Guo, Zhonghao Hu, Yuren Mao, Baihua Zheng 等VLDB 2025 · 被引用 6 次
- Qualitative Join Discovery in Data Lakes using ExamplesMir Mahathir Mohammad, El Kindi RezigSIGMOD 2026 · 被引用 6 次
它引用的顶会 Paper21
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai 等ICML 2022 · 被引用 1,629 次
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 被引用 845 次
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 被引用 417 次
- Lessons Learned from the Chameleon TestbedKate Keahey, Jason Anderson, Zhuo Zhen, Pierre Riteau 等USENIX ATC 2020 · 被引用 398 次
相关 Paper
- AutoDDG: Automated Dataset Description Generation using Large Language ModelsHaoxiang Zhang, Yurong Liu, Aécio S. R. Santos, Wei-Lun Hung 等SIGMOD 2026 · 被引用 17 次
- DiscoveryBench: Towards Data-Driven Discovery with Large Language ModelsBodhisattwa Prasad Majumder, Harshit Surana, Dhruv Agarwal, Bhavana Dalvi Mishra 等ICLR 2025
- Language-Guided Audio-Visual Source Separation via Trimodal ConsistencyReuben Tan, Arijit Ray, Andrea Burns, Bryan A. Plummer 等CVPR 2023
- AutoData: A Multi-Agent System for Open Web Data CollectionTianyi Ma, Yiyue Qian, Zheyuan Zhang, Zehong Wang 等NeurIPS 2025 · 被引用 28 次
- Large-Scale Unsupervised Object DiscoveryHuy V. Vo, Elena Sizikova, Cordelia Schmid, Patrick Pérez 等NeurIPS 2021 · 被引用 63 次
