Lune

NeurIPS2025顶会

TreeSynth: Synthesizing Diverse Data from Scratch via Tree-Guided Subspace Partitioning

Sheng Wang, Pengan Chen, Jingqi Zhou, Qintong Li, Jingwei Dong, Jiahui Gao, Boyang Xue, Jiyue Jiang, Lingpeng Kong, Chuan Wu

2025年份
9被引次数
1顶会引用

摘要

Model customization necessitates high-quality and diverse datasets, but acquiring such data remains time-consuming and labor-intensive. Despite the great potential of large language models (LLMs) for data synthesis, current approaches are constrained by limited seed data, model biases, and low-variation prompts, resulting in limited diversity and biased distributions with the increase of data scales. To tackle this challenge, we introduce TREESYNTH, a tree-guided subspace-based data synthesis approach inspired by decision trees. It constructs a spatial partitioning tree to recursively divide a task-specific full data space (i.e., root node) into numerous atomic subspaces (i.e., leaf nodes) with mutually exclusive and exhaustive attributes to ensure both distinctiveness and comprehensiveness before synthesizing samples within each atomic subspace. This globally dividing-and-synthesizing method finally collects subspace samples into a comprehensive dataset, effectively circumventing repetition and space collapse to ensure the diversity of large-scale data synthesis. Furthermore, the spatial partitioning tree enables sample allocation into atomic subspaces, allowing the rebalancing of existing datasets for more balanced and comprehensive distributions. Empirically, extensive experiments across diverse benchmarks consistently demonstrate the superior data diversity, model performance, and robust scalability of TREESYNTH compared to both humancrafted datasets and peer data synthesis methods, with an average performance gain reaching 10%. Besides, the consistent improvements of TREESYNTH-balanced datasets highlight its efficacious application to redistribute existing datasets for more comprehensive coverage and the induced performance enhancement. The code is available at https://github.com/cpa2001/TreeSynth. scale data synthesis. Furthermore, the improved results achieved by applying TREESYNTH to the synthetic datasets demonstrate its efficacious application to redistribute existing datasets for more comprehensive coverage and the induced performance enhancement.

The main contributions are summarized as follows:

• We propose TREESYNTH, a tree-guided subspace-based data synthesis approach, which features mutually exclusive and exhaustive subspace partitioning, effectively circumventing repetition and space collapse to ensure the diversity of large-scale data synthesis.

• Extensive experiments consistently highlight TREESYNTH's superior data diversity, model performance and robust scalability over human-crafted datasets and peer synthesis methods.

• The sample allocation of TREESYNTH allows re-balancing existing datasets for more comprehensive coverage, leading to empirically verified downstream performance enhancement.

A fruit seller has 250 apples and sells them in bags of 5. If he sells 30 bags of apples, how many apples does he have left? Emily is reading a book that has 240 pages. She reads 20 pages every day. After reading for 8 days, how many pages does she still have left to read? … In a school, there are 3 classes with 24 students each and 2 classes with 28 students each. How many students are there in total in the school?

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext ca054c24-1227-4a04-b606-d543d1678453

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper27

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖