On the Role of Anticausal Direction in LLM-based Data Synthesis
Bohan Jiang, Pingchuan Ma, Zhen Tan, Zhuoyu Shi, Fred Morstatter, Adrienne Raglin, Huan Liu
摘要
Large Language Models (LLMs) are increasingly used to generate synthetic data. Most LLM-based data synthesis workflows are anticausal : the user injects Y into the prompt to enforce targeted generation of X (YX). This anticausal direction seems contradictory to the natural direction of data synthesis in machine learning and data mining: raw content X is produced first and the supervision signal Y is assigned afterward (XY), i.e., the causal direction. This contrast impels us to investigate if a synthesis direction can impact synthetic data quality and downstream utility. We therefore construct paired causal and anticausal synthetic datasets for five tasks, sentiment analysis, causal language detection, ideology detection, math reasoning, and summarization. We then evaluate them via supervised fine-tuning and post-training models across these tasks. Our study finds that models trained on causal synthetic data consistently outperform those trained on anticausal data. We then compare them with those models trained on real-world data. The research motivates us to propose XSyn, a generative framework that transforms anticausal and causal synthetic data to better match real-world distributions. XSyn is jointly optimized with four distinct reward models, steering transformations toward outputs that are closer to real-world data. Further experiments on BERT, Llama, and Qwen show that training on XSyn-transformed data yields substantial gains over both causal and anticausal counterparts. We conclude this study with recommendations and future work.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- SIPDO: Closed-Loop Prompt Optimization via Synthetic Data FeedbackYaoning Yu, Ye Yu, Peiyan Zhang, Kai Wei 等ICLR 2026 · 被引用 7 次
- A Rigorous Evaluation of LLM Data Generation Strategies for Low-Resource LanguagesTatiana Anikina, Ján Cegin, Jakub Simko, Simon OstermannEMNLP 2025
- LLMs Are Prone to Fallacies in Causal InferenceNitish Joshi, Abulhair Saparov, Yixin Wang, He HeEMNLP 2024 · 被引用 10 次
- Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text ClassificationHsun-Yu Kuo, Yin-Hsiang Liao, Yu-Chieh Chao, Wei-Yun Ma 等ICLR 2025
- OptimSyn: Influence-Guided Rubrics Optimization for Synthetic Data GenerationZhiting Fan, Ruizhe Chen, Tianxiang Hu, Ru Peng 等ICLR 2026 · 被引用 3 次
