Lune

KDD2026顶会

On the Role of Anticausal Direction in LLM-based Data Synthesis

Bohan Jiang, Pingchuan Ma, Zhen Tan, Zhuoyu Shi, Fred Morstatter, Adrienne Raglin, Huan Liu

2026年份

摘要

Large Language Models (LLMs) are increasingly used to generate synthetic data. Most LLM-based data synthesis workflows are anticausal : the user injects Y into the prompt to enforce targeted generation of X (YX). This anticausal direction seems contradictory to the natural direction of data synthesis in machine learning and data mining: raw content X is produced first and the supervision signal Y is assigned afterward (XY), i.e., the causal direction. This contrast impels us to investigate if a synthesis direction can impact synthetic data quality and downstream utility. We therefore construct paired causal and anticausal synthetic datasets for five tasks, sentiment analysis, causal language detection, ideology detection, math reasoning, and summarization. We then evaluate them via supervised fine-tuning and post-training models across these tasks. Our study finds that models trained on causal synthetic data consistently outperform those trained on anticausal data. We then compare them with those models trained on real-world data. The research motivates us to propose XSyn, a generative framework that transforms anticausal and causal synthetic data to better match real-world distributions. XSyn is jointly optimized with four distinct reward models, steering transformations toward outputs that are closer to real-world data. Further experiments on BERT, Llama, and Qwen show that training on XSyn-transformed data yields substantial gains over both causal and anticausal counterparts. We conclude this study with recommendations and future work.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

lune papers get d995e7d3-4228-44af-bc3c-e6b89765b489

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖