Transplant Then Regenerate: A New Paradigm for Text Data Augmentation
Guangzhan Wang, Hongyu Zhang, Beijun Shen, Xiaodong Gu
摘要
Data augmentation is a critical technique in deep learning. Traditional methods like Backtranslation typically focus on lexical-level rephrasing, which primarily produces variations with the same semantics. While large language models (LLMs) have enhanced text augmentation by their "knowledge emergence" capability, controlling the style and structure of these outputs remains challenging and requires meticulous prompt engineering. In this paper, we propose LMTransplant, a novel text augmentation paradigm leveraging LLMs. The core idea of LMTransplant is transplant-thenregenerate: incorporating seed text into a context expanded by LLM, and asking the LLM to regenerate a variant based on the expanded context. This strategy allows the model to create more diverse and creative content-level variants by fully leveraging the knowledge embedded in LLMs, while preserving the core attributes of the original text. We evaluate LMTransplant across various text-related tasks, demonstrating its superior performance over existing text augmentation methods. Moreover, LMTransplant demonstrates exceptional scalability as the size of augmented data grows.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and InferenceBenjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller 等ACL 2025 · 被引用 552 次
- LLM-powered Data Augmentation for Enhanced Cross-lingual PerformanceChenxi Whitehouse, Monojit Choudhury, Alham Fikri AjiEMNLP 2023 · 被引用 53 次
- MLQA: Evaluating Cross-lingual Extractive Question AnsweringPatrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel 等ACL 2020 · 被引用 52 次
- ABEX: Data Augmentation for Low-Resource NLU via Expanding Abstract DescriptionsSreyan Ghosh, Utkarsh Tyagi, Sonal Kumar, Chandra Kiran Reddy Evuru 等ACL 2024 · 被引用 3 次
相关 Paper
- Do Not Have Enough Data? Deep Learning to the Rescue!Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor 等AAAI 2020 · 被引用 398 次
- Paraphrasing as Zero-shot Translation with Feature-guided Diversity EnhancementZiyue Yan, Hongying Zan, Xinglin Lyu, Hongfei XuACL 2026
- Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text ClassificationHsun-Yu Kuo, Yin-Hsiang Liao, Yu-Chieh Chao, Wei-Yun Ma 等ICLR 2025
- Learning to Generalize to More: Continuous Semantic Augmentation for Neural Machine TranslationXiangpeng Wei, Heng Yu, Yue Hu, Rongxiang Weng 等ACL 2022 · 被引用 26 次
- CultureLLM: Incorporating Cultural Differences into Large Language ModelsCheng Li, Mengzhuo Chen, Jindong Wang, Sunayana Sitaram 等NeurIPS 2024 · 被引用 101 次
