Explain the Synth: Interpretable Evaluation of LLM Data Synthesis
Yue Yang, Fan Yang, Yu Bai, Hao Wang
Abstract
Large language models (LLMs) are increasingly used to generate synthetic data, in which tabular data constitute a fundamental data modality across a wide range of domains. Yet, current evaluation practices often provide limited insights into whether the synthetic data preserve real data-generating relationships or introduce plausible-looking artifacts. We present a conceptually simple, interpretable auditing framework that compares the explanatory structure induced by real versus synthetic data. The key idea is to use a transparent rule-based model as a shared explanatory language: we extract rules from real data to summarize how features relate to labels, then examine how this rule structure changes when explained using LLM-generated data. Importantly, these rules are derived by an independent rule auditor rather than by the generator itself. The resulting "explanation shift" reveals which relationships are preserved, weakened, removed, or newly introduced by the generator, offering actionable diagnostics beyond aggregate fidelity scores. We further provide a theoretical perspective that links explanation shift and cross-domain predictive gaps to distribution mismatch within an interpretable hypothesis class. Overall, our approach turns synthetic data evaluation into a human-auditable comparison of explanations, improving transparency for LLM-based tabular synthesis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fac340a7-27ab-4014-a54d-5105d3ebb768Builds on5
- TabDDPM: Modelling Tabular Data with Diffusion ModelsAkim Kotelnikov, Dmitry Baranchuk, Ivan Rubachev, Artem BabenkoICML 2023 · 518 citations
- Mixed-Type Tabular Data Synthesis with Score-based Diffusion in Latent SpaceHengrui Zhang, Jiani Zhang, Zhengyuan Shen, Balasubramaniam Srinivasan et al.ICLR 2024 · 233 citations
- Language Models are Realistic Tabular Data GeneratorsVadim Borisov, Kathrin Seßler, Tobias Leemann, Martin Pawelczyk et al.ICLR 2023 · 45 citations
- Towards Explaining Distribution ShiftsSean Kulinski, David I. InouyeICML 2023 · 38 citations
- Neural+Symbolic Approaches for Interpretable Actor-Critic Reinforcement LearningYue Yang, Fan Yang, Yu Bai, Hao WangICLR 2026
Related papers
- TabReX: Tabular Referenceless eXplainable EvaluationTejas Anvekar, Junha Park, Aparna Garimella, Vivek GuptaACL 2026
- Synthetic Artifact Auditing: Tracing LLM-Generated Synthetic Data Usage in Downstream ApplicationsYixin Wu, Ziqing Yang, Yun Shen, Michael Backes et al.USENIX Security 2025
- ConSim: Measuring Concept-Based Explanations' Effectiveness with Automated SimulatabilityAntonin Poché, Alon Jacovi, Agustin Martin Picard, Victor Boutin et al.ACL 2025 · 8 citations
- Systematic Assessment of Tabular Data SynthesisYuntao Du, Ninghui LiCCS 2025 · 2 citations
- Making Sense of LLM Decisions: A Prototype-based Framework for Explainable ClassificationBowen Wei, Mehrdad Fazli, Ziwei ZhuAAAI 2026
