Exploring Quality and Diversity in Synthetic Data Generation for Argument Mining
Jianzhu Bao, Yuqi Huang, Yang Sun, Wenya Wang, Yice Zhang, Bojun Jin, Ruifeng Xu
Abstract
The advancement of Argument Mining (AM) is hindered by a critical bottleneck: the scarcity of structure-annotated datasets, which are expensive to create manually. Inspired by recent successes in synthetic data generation across various NLP tasks, this paper explores methodologies for LLMs to generate synthetic data for AM. We investigate two complementary synthesis perspectives: a quality-oriented synthesis approach, which employs structure-aware paraphrasing to preserve annotation quality, and a diversity-oriented synthesis approach, which generates novel argumentative texts with diverse topics and argument structures. Experiments on three datasets show that augmenting original training data with our synthetic data, particularly when combining both quality-and diversity-oriented instances, significantly enhances the performance of existing AM models, both in full-data and low-resource settings. Moreover, the positive correlation between synthetic data volume and model performance highlights the scalability of our methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cfe51872-37d5-4e98-a556-b6ce4fc6c0e2Builds on16
- Diversity-Driven Synthesis: Enhancing Dataset Distillation through Directed Weight AdjustmentJiawei Du, Xin Zhang, Juncheng Hu, Wenxin Huang et al.NeurIPS 2024 · 43 citations
- Exploiting Asymmetry for Synthetic Training Data Generation: SynthIE and the Case of Information ExtractionMartin Josifoski, Marija Sakota, Maxime Peyrard, Robert WestEMNLP 2023 · 43 citations
- Progressive Self-Training with Discriminator for Aspect Term ExtractionQianlong Wang, Zhiyuan Wen, Qin Zhao, Min Yang et al.EMNLP 2021 · 39 citations
- A Generative Model for End-to-End Argument Mining with Reconstructed Positional Encoding and Constrained Pointer MechanismJianzhu Bao, Yuhang He, Yang Sun, Bin Liang et al.EMNLP 2022 · 15 citations
- A Synthetic Data Generation Framework for Grounded DialoguesJianzhu Bao, Rui Wang, Yasheng Wang, Aixin Sun et al.ACL 2023 · 11 citations
Related papers
- Diversity-oriented Data Augmentation with Large Language ModelsZaitian Wang, Jinghan Zhang, Xinhao Zhang, Kunpeng Liu et al.ACL 2025 · 11 citations
- Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text ClassificationHsun-Yu Kuo, Yin-Hsiang Liao, Yu-Chieh Chao, Wei-Yun Ma et al.ICLR 2025
- Synthetic Data Generation for Training Diversified Commonsense Reasoning ModelsTianhui Zhang, Bei Peng, Danushka BollegalaACL 2026 · 1 citation
- Paraphrase Augmented Task-Oriented Dialog GenerationSilin Gao, Yichi Zhang, Zhijian Ou, Zhou YuACL 2020 · 78 citations
- Scaling Low-Resource MT via Synthetic Data Generation with LLMsOna de Gibert, Joseph Attieh, Teemu Vahtola, Mikko Aulamo et al.EMNLP 2025 · 2 citations
