Pneuma: Leveraging LLMs for Tabular Data Representation and Retrieval in an End-to-End System
Muhammad Imam Luthfi Balaka, David Alexander, Qiming Wang, Yue Gong, Adila Krisnadhi, Raul Castro Fernandez
Abstract
Finding relevant tables among databases, lakes, and repositories is the first step in extracting value from data. Such a task remains difficult because assessing whether a table is relevant to a problem does not always depend only on its content but also on the context, which is usually tribal knowledge known to the individual or team. While tools like data catalogs and academic data discovery systems target this problem, they rely on keyword search or more complex interfaces, limiting non-technical users' ability to find relevant data. The advent of large language models (LLMs) offers a unique opportunity for users to ask questions directly in natural language, making dataset discovery more intuitive, accessible, and efficient. In this paper, we introduce Pneuma , a retrieval-augmented generation (RAG) system designed to efficiently and effectively discover tabular data. Pneuma leverages large language models (LLMs) for both table representation and table retrieval. For table representation, Pneuma preserves schema and row-level information to ensure comprehensive data understanding. For table retrieval, Pneuma augments LLMs with traditional information retrieval techniques, such as full-text and vector search, harnessing the strengths of both to improve retrieval performance. To evaluate Pneuma , we generate comprehensive benchmarks that simulate table discovery workload on six real-world datasets including enterprise data, scientific databases, warehousing data, and open data. Our results demonstrate that Pneuma outperforms widely used table search systems (such as full-text search and state-of-the-art RAG systems) in accuracy and resource efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 48e864ef-7c05-482e-aaa4-d6886eee38eeCited by top-tier papers6
- AutoDDG: Automated Dataset Description Generation using Large Language ModelsHaoxiang Zhang, Yurong Liu, Aécio S. R. Santos, Wei-Lun Hung et al.SIGMOD 2026 · 17 citations
- Task Cascades for Efficient Unstructured Data ProcessingShreya Shankar, Sepanta Zeighami, Aditya G. ParameswaranSIGMOD 2026 · 7 citations
- Relational Deep Dive: Error-Aware Queries Over Unstructured DataDaren Chao, Kaiwen Chen, Naiqing Guan, Nick KoudasVLDB 2026 · 3 citations
- Revisiting Task-Oriented Dataset Search in the Era of Large Language Models: Challenges, Benchmark, and SolutionZixin Wei, Yucan Guo, Jinyang Li, Xiaolin Han et al.VLDB 2026 · 1 citation
- Decomposition-Driven Multi-Table Retrieval and Reasoning for Numerical Question AnsweringFeng Luo, Hai Lan, Hui Luo, Zhifeng Bao et al.ICDE 2026 · 1 citation
Builds on19
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-ServeAmey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan et al.OSDI 2024 · 537 citations
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 417 citations
- Matryoshka Representation LearningAditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford et al.NeurIPS 2022 · 364 citations
- TAPEX: Table Pre-training via Learning a Neural SQL ExecutorQian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi et al.ICLR 2022 · 347 citations
Related papers
- TableRAG: Million-Token Table Understanding with Language ModelsSi-An Chen, Lesly Miculicich, Julian Eisenschlos, Zifeng Wang et al.NeurIPS 2024 · 86 citations
- SchemaRAG: A Schema-aware Retrieval-Augmented Generation Framework for Text-to-SQLDi Wu, Zetong Tang, Yi He, Xin LuoSIGMOD 2026 · 9 citations
- SARA: Selective and Adaptive Retrieval-augmented Generation with Context CompressionYiqiao Jin, Kartik Sharma, Vineeth Rakesh, Yingtong Dou et al.ACL 2026 · 7 citations
- Large Language Models are Versatile Decomposers: Decomposing Evidence and Questions for Table-based ReasoningYunhu Ye, Binyuan Hui, Min Yang, Binhua Li et al.SIGIR 2023 · 75 citations
- Querying Templatized Document Collections with Large Language ModelsYiming Lin, Madelon Hulsebos, Ruiying Ma, Shreya Shankar et al.ICDE 2025 · 4 citations
