AutoDDG: Automated Dataset Description Generation using Large Language Models
Haoxiang Zhang, Yurong Liu, Aécio S. R. Santos, Wei-Lun Hung, Juliana Freire
Abstract
The proliferation of datasets across open data portals and enterprise data lakes presents an opportunity for deriving data-driven insights. Widely-used dataset search systems rely on keyword search over dataset metadata, including descriptions, to support discovery. Therefore, when these descriptions are incomplete, missing, or inconsistent with dataset contents, findability is severely compromised. To improve findability, we introduce AutoDDG, a framework that automatically generates descriptions of tabular data. By adopting a data-driven approach to summarize dataset contents and leveraging large language models (LLMs) to enrich summaries with semantic information and produce human-readable text, AutoDDG derives descriptions that are comprehensive, accurate, readable, and concise. A critical challenge in this problem is evaluating the effectiveness of description generation methods and assessing the quality of the generated descriptions. We propose a comprehensive evaluation methodology that combines retrieval, reference-based, and reference-free assessment, with human validation. Our experimental results using new benchmarks demonstrate that AutoDDG generates high-quality, accurate descriptions at scale, significantly improving dataset retrieval performance across diverse use cases. AutoDDG is publicly available at https://github.com/VIDA-NYU/AutoDDG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- MosaicJoin: Compact Semantic Sketches for Value-Level Join DiscoveryGrace Fan, Eden Wu, Majid Daliri, Juliana FreireVLDB 2026 · 1 citation
- LakeQA: An Exploratory QA Benchmark over a Million-Scale Data LakeHaonan Wang, Jiaxiang Liu, Yurong Liu, Austin Wijaya et al.ICML 2026
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu et al.VLDB 2021 · 2,406 citations
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 865 citations
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
Related papers
- Pneuma: Leveraging LLMs for Tabular Data Representation and Retrieval in an End-to-End SystemMuhammad Imam Luthfi Balaka, David Alexander, Qiming Wang, Yue Gong et al.SIGMOD 2025 · 12 citations
- TabReX: Tabular Referenceless eXplainable EvaluationTejas Anvekar, Junha Park, Aparna Garimella, Vivek GuptaACL 2026
- NameGuess: Column Name Expansion for Tabular DataJiani Zhang, Zhengyuan Shen, Balasubramaniam Srinivasan, Shen Wang et al.EMNLP 2023 · 6 citations
- DiscoveryBench: Towards Data-Driven Discovery with Large Language ModelsBodhisattwa Prasad Majumder, Harshit Surana, Dhruv Agarwal, Bhavana Dalvi Mishra et al.ICLR 2025
- Autocompose: Automatic Generation of Pose Transition Descriptions for Composed Pose Retrieval Using Multimodal LLMsYi-Ting Shen, Sungmin Eum, Doheon Lee, Rohit Shete et al.ICCV 2025
