Distribution-Aligned Synthetic Text Generation via Tail-Aware Enhancement
Yuan Fan, Xiaoyuan Liu, Bo Liu, Wubing Wang, Jia Sun, Wenzhi Chen, Huaikang Fang, Lifeng Tao, Fan Mo
摘要
Recent advances in generative AI have popularized synthetic content for training, offering a practical alternative to costly data curation while addressing privacy concerns. However, accumulating evidence shows that the indiscriminate reuse of synthetic data can induce model collapse—a degenerative process that contracts the learned distribution and erodes rare features. For instance, when models are iteratively trained on their own synthetic outputs, the upper tail of the perplexity distribution substantially compresses, with high-percentile values dropping by nearly half—a clear indicator of severe diversity loss. To counter this, we introduce DASGen, a Distribution-Aligned Synthetic Text Generation framework via tail-aware enhancement. Our method first identifies underrepresented regions via embedding-space mining, then steers a frozen, hosted LLM using semantically-structured prompts and a discriminative diversity objective to enrich tail features. This training-free approach enables direct deployment in existing data pipelines. Extensive evaluations on Yelp and ICLR'25 review benchmarks show that DASGen significantly outperforms competitive baselines, achieving tail coverage (98.54% on Yelp; 92.00% on ICLR'25) along with improved downstream accuracy. Overall, DASGen provides a practical path to synthesizing distribution-aligned text by explicitly enhancing tail regions, producing synthetic corpora with enhanced coverage and diversity for more reliable long-tailed applications.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- DEALT: LLM-driven Diversity-Enhanced Data Augmentation for Long-Tail Text ClassificationWayne Lu, Xiaoxi CuiAAAI 2026 · 被引用 2 次
- Breaking the Generator Barrier: Disentangled Representation for Generalizable AI-Text DetectionXiao Pu, Zepeng Cheng, Lin Yuan, Yu Wu 等ACL 2026 · 被引用 1 次
- Synthetic Text Generation with Differential Privacy: A Simple and Practical RecipeXiang Yue, Huseyin A. Inan, Xuechen Li, Girish Kumar 等ACL 2023 · 被引用 24 次
- Language Generation with Replay: A Learning-Theoretic View of Model CollapseGiorgio Racca, Michal Valko, Amartya SanyalICML 2026 · 被引用 4 次
- Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM DiversityJiayi Zhang, Simon Yu, Derek Chong, Anthony Sicilia 等ICML 2026 · 被引用 102 次
