Dense Representation Learning and Retrieval for Tabular Data Prediction
Lei Zheng, Ning Li, Xianyu Chen, Quan Gan, Weinan Zhang
Abstract
Data science is concerned with mining data patterns from a database, which is assembled by tabular data. As the routine of machine learning, most of the previous work mining the tabular data's pattern based on a single instance. However, they neglect the similar tabular data instances that could help make the label prediction of the target data instance. Recently, some retrieval-based methods for tabular data label prediction have been proposed, which, however, treat the data as sparse vectors to perform the retrieval, which fails to make use of the semantic information of the tabular data. To address such a problem, in this paper, we propose a novel framework of dense retrieval on tabular data (DERT) to support flexible data representation learning and effective label prediction on tabular data. DERT consists of two major components: (i) the encoder that makes the tabular data as embeddings, which could be trained by flexible neural networks and auxiliary loss functions; (ii) the retrieval and prediction component, which makes use of similar rows in the table to make label prediction of the target row. We test DERT on two tasks based on five real-world datasets and experimental results show that DERT achieves consistent improvements over the state-of-the-art and various baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 85baefa2-ed7f-44d3-9979-22c8a03a36a1Cited by top-tier papers1
Ask how each one uses itBuilds on16
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu et al.VLDB 2021 · 2,406 citations
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 1,847 citations
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer et al.ICLR 2020 · 1,038 citations
Related papers
- Retrieval & Interaction Machine for Tabular Data PredictionJiarui Qin, Weinan Zhang, Rong Su, Zhirong Liu et al.KDD 2021 · 33 citations
- TabR: Tabular Deep Learning Meets Nearest NeighborsYury Gorishniy, Ivan Rubachev, Nikolay Kartashev, Daniil Shlenskii et al.ICLR 2024 · 78 citations
- DANets: Deep Abstract Networks for Tabular Data Classification and RegressionJintai Chen, Kuanlun Liao, Yao Wan, Danny Z. Chen et al.AAAI 2022 · 82 citations
- Learning Enhanced Representation for Tabular Data via Neighborhood PropagationKounianhua Du, Weinan Zhang, Ruiwen Zhou, Yangkun Wang et al.NeurIPS 2022 · 23 citations
- PTaRL: Prototype-based Tabular Representation Learning via Space CalibrationHangting Ye, Wei Fan, Xiaozhuang Song, Shun Zheng et al.ICLR 2024 · 37 citations
