RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data Preparation
Nan Tang, Ju Fan, Fangyi Li, Jianhong Tu, Xiaoyong Du, Guoliang Li, Samuel Madden, Mourad Ouzzani
摘要
Can AI help automate human-easy but computer-hard data preparation tasks that burden data scientists, practitioners, and crowd workers? We answer this question by presenting RPT, a denoising autoencoder for tuple-to-X models ("𝑋 " could be tuple, token, label, JSON, and so on). RPT is pre-trained for a tuple-to-tuple model by corrupting the input tuple and then learning a model to reconstruct the original tuple. It adopts a Transformer-based neural translation architecture that consists of a bidirectional encoder (similar to BERT) and a left-to-right autoregressive decoder (similar to GPT), leading to a generalization of both BERT and GPT. The pre-trained RPT can already support several common data preparation tasks such as data cleaning, auto-completion and schema matching. Better still, RPT can be fine-tuned on a wide range of data preparation tasks, such as value normalization, data transformation, data annotation, etc. To complement RPT, we also discuss several appealing techniques such as collaborative training and few-shot learning for entity resolution, and few-shot learning and NLP question-answering for information extraction. In addition, we identify a series of research opportunities to advance the field of data preparation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Annotating Columns with Pre-trained Language ModelsYoshihiko Suhara, Jinfeng Li, Yuliang Li, Dan Zhang 等SIGMOD 2022 · 被引用 81 次
- CodexDB: Synthesizing Code for Query Processing from Natural Language Instructions using GPT-3 CodexImmanuel TrummerVLDB 2022 · 被引用 77 次
- DeepJoin: Joinable Table Discovery with Pre-trained Language ModelsYuyang Dong, Chuan Xiao, Takuma Nozawa, Masafumi Enomoto 等VLDB 2023 · 被引用 53 次
- Domain Adaptation for Deep Entity ResolutionJianhong Tu, Ju Fan, Nan Tang, Peng Wang 等SIGMOD 2022 · 被引用 46 次
- Sudowoodo: Contrastive Self-supervised Learning for Multi-purpose Data Integration and PreparationRunhui Wang, Yuliang Li, Jin WangICDE 2023 · 被引用 32 次
它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu 等VLDB 2021 · 被引用 2,406 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang 等ICLR 2020 · 被引用 674 次
相关 Paper
- Pre-Trained Image Processing TransformerHanting Chen, Yunhe Wang, Tianyu Guo, Chang Xu 等CVPR 2021
- BrainBERT: Self-supervised representation learning for intracranial recordingsChristopher Wang, Vighnesh Subramaniam, Adam Uri Yaari, Gabriel Kreiman 等ICLR 2023 · 被引用 13 次
- Improving the Efficiency and Effectiveness for BERT-based Entity ResolutionBing Li, Yukai Miao, Yaoshu Wang, Yifang Sun 等AAAI 2021 · 被引用 45 次
- No more fine-tuning? an experimental evaluation of prompt tuning in code intelligenceChaozheng Wang, Yuanhang Yang, Cuiyun Gao, Yun Peng 等FSE 2022 · 被引用 148 次
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan 等VLDB 2021 · 被引用 484 次
