Lune

NeurIPS2025Top-tier venue

TreeSynth: Synthesizing Diverse Data from Scratch via Tree-Guided Subspace Partitioning

Sheng Wang, Pengan Chen, Jingqi Zhou, Qintong Li, Jingwei Dong, Jiahui Gao, Boyang Xue, Jiyue Jiang, Lingpeng Kong, Chuan Wu

2025Year
9Citations
1Top-tier citations

Abstract

Model customization necessitates high-quality and diverse datasets, but acquiring such data remains time-consuming and labor-intensive. Despite the great potential of large language models (LLMs) for data synthesis, current approaches are constrained by limited seed data, model biases, and low-variation prompts, resulting in limited diversity and biased distributions with the increase of data scales. To tackle this challenge, we introduce TREESYNTH, a tree-guided subspace-based data synthesis approach inspired by decision trees. It constructs a spatial partitioning tree to recursively divide a task-specific full data space (i.e., root node) into numerous atomic subspaces (i.e., leaf nodes) with mutually exclusive and exhaustive attributes to ensure both distinctiveness and comprehensiveness before synthesizing samples within each atomic subspace. This globally dividing-and-synthesizing method finally collects subspace samples into a comprehensive dataset, effectively circumventing repetition and space collapse to ensure the diversity of large-scale data synthesis. Furthermore, the spatial partitioning tree enables sample allocation into atomic subspaces, allowing the rebalancing of existing datasets for more balanced and comprehensive distributions. Empirically, extensive experiments across diverse benchmarks consistently demonstrate the superior data diversity, model performance, and robust scalability of TREESYNTH compared to both humancrafted datasets and peer data synthesis methods, with an average performance gain reaching 10%. Besides, the consistent improvements of TREESYNTH-balanced datasets highlight its efficacious application to redistribute existing datasets for more comprehensive coverage and the induced performance enhancement. The code is available at https://github.com/cpa2001/TreeSynth. scale data synthesis. Furthermore, the improved results achieved by applying TREESYNTH to the synthetic datasets demonstrate its efficacious application to redistribute existing datasets for more comprehensive coverage and the induced performance enhancement.

The main contributions are summarized as follows:

• We propose TREESYNTH, a tree-guided subspace-based data synthesis approach, which features mutually exclusive and exhaustive subspace partitioning, effectively circumventing repetition and space collapse to ensure the diversity of large-scale data synthesis.

• Extensive experiments consistently highlight TREESYNTH's superior data diversity, model performance and robust scalability over human-crafted datasets and peer synthesis methods.

• The sample allocation of TREESYNTH allows re-balancing existing datasets for more comprehensive coverage, leading to empirically verified downstream performance enhancement.

A fruit seller has 250 apples and sells them in bags of 5. If he sells 30 bags of apples, how many apples does he have left? Emily is reading a book that has 240 pages. She reads 20 pages every day. After reading for 8 days, how many pages does she still have left to read? … In a school, there are 3 classes with 24 students each and 2 classes with 28 students each. How many students are there in total in the school?

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers1

Ask how each one uses it

Builds on27

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines