Exploring Quality and Diversity in Synthetic Data Generation for Argument Mining
Jianzhu Bao, Yuqi Huang, Yang Sun, Wenya Wang, Yice Zhang, Bojun Jin, Ruifeng Xu
摘要
The advancement of Argument Mining (AM) is hindered by a critical bottleneck: the scarcity of structure-annotated datasets, which are expensive to create manually. Inspired by recent successes in synthetic data generation across various NLP tasks, this paper explores methodologies for LLMs to generate synthetic data for AM. We investigate two complementary synthesis perspectives: a quality-oriented synthesis approach, which employs structure-aware paraphrasing to preserve annotation quality, and a diversity-oriented synthesis approach, which generates novel argumentative texts with diverse topics and argument structures. Experiments on three datasets show that augmenting original training data with our synthetic data, particularly when combining both quality-and diversity-oriented instances, significantly enhances the performance of existing AM models, both in full-data and low-resource settings. Moreover, the positive correlation between synthetic data volume and model performance highlights the scalability of our methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Diversity-Driven Synthesis: Enhancing Dataset Distillation through Directed Weight AdjustmentJiawei Du, Xin Zhang, Juncheng Hu, Wenxin Huang 等NeurIPS 2024 · 被引用 43 次
- Exploiting Asymmetry for Synthetic Training Data Generation: SynthIE and the Case of Information ExtractionMartin Josifoski, Marija Sakota, Maxime Peyrard, Robert WestEMNLP 2023 · 被引用 43 次
- Progressive Self-Training with Discriminator for Aspect Term ExtractionQianlong Wang, Zhiyuan Wen, Qin Zhao, Min Yang 等EMNLP 2021 · 被引用 39 次
- A Generative Model for End-to-End Argument Mining with Reconstructed Positional Encoding and Constrained Pointer MechanismJianzhu Bao, Yuhang He, Yang Sun, Bin Liang 等EMNLP 2022 · 被引用 15 次
- A Synthetic Data Generation Framework for Grounded DialoguesJianzhu Bao, Rui Wang, Yasheng Wang, Aixin Sun 等ACL 2023 · 被引用 11 次
相关 Paper
- Diversity-oriented Data Augmentation with Large Language ModelsZaitian Wang, Jinghan Zhang, Xinhao Zhang, Kunpeng Liu 等ACL 2025 · 被引用 11 次
- Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text ClassificationHsun-Yu Kuo, Yin-Hsiang Liao, Yu-Chieh Chao, Wei-Yun Ma 等ICLR 2025
- Synthetic Data Generation for Training Diversified Commonsense Reasoning ModelsTianhui Zhang, Bei Peng, Danushka BollegalaACL 2026 · 被引用 1 次
- Paraphrase Augmented Task-Oriented Dialog GenerationSilin Gao, Yichi Zhang, Zhijian Ou, Zhou YuACL 2020 · 被引用 78 次
- Scaling Low-Resource MT via Synthetic Data Generation with LLMsOna de Gibert, Joseph Attieh, Teemu Vahtola, Mikko Aulamo 等EMNLP 2025 · 被引用 2 次
