Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and Beyond
Zhengjie Miao, Yuliang Li, Xiaolan Wang
Abstract
Deep Learning revolutionizes almost all fields of computer science including data management. However, the demand for high-quality training data is slowing down deep neural nets' wider adoption. To this end, data augmentation (DA), which generates more labeled examples from existing ones, becomes a common technique. Meanwhile, the risk of creating noisy examples and the large space of hyper-parameters make DA less attractive in practice. We introduce Rotom, a multi-purpose data augmentation framework for a range of data management and mining tasks including entity matching, data cleaning, and text classification. Rotom features InvDA, a new DA operator that generates natural yet diverse augmented examples by formulating DA as a seq2seq task. The key technical novelty of Rotom is a meta-learning framework that automatically learns a policy for combining examples from different DA operators, whereby combinatorially reduces the hyper-parameters space. Our experimental results show that Rotom effectively improves a model's performance by combining multiple DA operators, even when applying them individually does not yield performance improvement. With this strength, Rotom outperforms the state-of-the-art entity matching and data cleaning systems in the low-resource settings as well as two recently proposed DA techniques for text classification.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- PromptEM: Prompt-tuning for Low-resource Generalized Entity MatchingPengfei Wang, Xiaocan Zeng, Lu Chen, Fan Ye et al.VLDB 2023 · 39 citations
- Sudowoodo: Contrastive Self-supervised Learning for Multi-purpose Data Integration and PreparationRunhui Wang, Yuliang Li, Jin WangICDE 2023 · 32 citations
- Text AutoAugment: Learning Compositional Augmentation Policy for Text ClassificationShuhuai Ren, Jinchao Zhang, Lei Li, Xu Sun et al.EMNLP 2021 · 22 citations
- FlexER: Flexible Entity Resolution for Multiple IntentsBar Genossar, Roee Shraga, Avigdor GalSIGMOD 2023 · 15 citations
- Watchog: A Light-weight Contrastive Learning based Framework for Column AnnotationZhengjie Miao, Jin WangSIGMOD 2024 · 14 citations
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
Related papers
- RoPDA: Robust Prompt-Based Data Augmentation for Low-Resource Named Entity RecognitionSihan Song, Furao Shen, Jian ZhaoAAAI 2024 · 7 citations
- What Makes Better Augmentation Strategies? Augment Difficult but Not too DifferentJaehyung Kim, Dongyeop Kang, Sungsoo Ahn, Jinwoo ShinICLR 2022 · 14 citations
- Robust and Informative Text Augmentation (RITA) via Constrained Worst-Case Transformations for Low-Resource Named Entity RecognitionHyunwoo Sohn, Baekkwan ParkKDD 2022 · 3 citations
- KnowDA: All-in-One Knowledge Mixture Model for Data Augmentation in Low-Resource NLPYufei Wang, Jiayi Zheng, Can Xu, Xiubo Geng et al.ICLR 2023 · 2 citations
- Data Boost: Text Data Augmentation Through Reinforcement Learning Guided Conditional GenerationRuibo Liu, Guangxuan Xu, Chenyan Jia, Weicheng Ma et al.EMNLP 2020 · 62 citations
