TreeSynth: Synthesizing Diverse Data from Scratch via Tree-Guided Subspace Partitioning
Sheng Wang, Pengan Chen, Jingqi Zhou, Qintong Li, Jingwei Dong, Jiahui Gao, Boyang Xue, Jiyue Jiang, Lingpeng Kong, Chuan Wu
摘要
Model customization necessitates high-quality and diverse datasets, but acquiring such data remains time-consuming and labor-intensive. Despite the great potential of large language models (LLMs) for data synthesis, current approaches are constrained by limited seed data, model biases, and low-variation prompts, resulting in limited diversity and biased distributions with the increase of data scales. To tackle this challenge, we introduce TREESYNTH, a tree-guided subspace-based data synthesis approach inspired by decision trees. It constructs a spatial partitioning tree to recursively divide a task-specific full data space (i.e., root node) into numerous atomic subspaces (i.e., leaf nodes) with mutually exclusive and exhaustive attributes to ensure both distinctiveness and comprehensiveness before synthesizing samples within each atomic subspace. This globally dividing-and-synthesizing method finally collects subspace samples into a comprehensive dataset, effectively circumventing repetition and space collapse to ensure the diversity of large-scale data synthesis. Furthermore, the spatial partitioning tree enables sample allocation into atomic subspaces, allowing the rebalancing of existing datasets for more balanced and comprehensive distributions. Empirically, extensive experiments across diverse benchmarks consistently demonstrate the superior data diversity, model performance, and robust scalability of TREESYNTH compared to both humancrafted datasets and peer data synthesis methods, with an average performance gain reaching 10%. Besides, the consistent improvements of TREESYNTH-balanced datasets highlight its efficacious application to redistribute existing datasets for more comprehensive coverage and the induced performance enhancement. The code is available at https://github.com/cpa2001/TreeSynth. scale data synthesis. Furthermore, the improved results achieved by applying TREESYNTH to the synthetic datasets demonstrate its efficacious application to redistribute existing datasets for more comprehensive coverage and the induced performance enhancement.
The main contributions are summarized as follows:
• We propose TREESYNTH, a tree-guided subspace-based data synthesis approach, which features mutually exclusive and exhaustive subspace partitioning, effectively circumventing repetition and space collapse to ensure the diversity of large-scale data synthesis.
• Extensive experiments consistently highlight TREESYNTH's superior data diversity, model performance and robust scalability over human-crafted datasets and peer synthesis methods.
• The sample allocation of TREESYNTH allows re-balancing existing datasets for more comprehensive coverage, leading to empirically verified downstream performance enhancement.
A fruit seller has 250 apples and sells them in bags of 5. If he sells 30 bags of apples, how many apples does he have left? Emily is reading a book that has 240 pages. She reads 20 pages every day. After reading for 8 days, how many pages does she still have left to read? … In a school, there are 3 classes with 24 students each and 2 classes with 28 students each. How many students are there in total in the school?
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper27
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex InstructionsCan Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng 等ICLR 2024 · 被引用 1,206 次
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun 等ICLR 2024 · 被引用 945 次
相关 Paper
- CorrSynth - A Correlated Sampling Method for Diverse Dataset Generation from LLMsSuhas S. Kowshik, Abhishek Divekar, Vijit MalikEMNLP 2024
- SPA: A Simple but Tough-to-Beat Baseline for Knowledge InjectionKexian Tang, Jiani Wang, Shaowen Wang, Kaifeng LyuICML 2026
- VOYAGER: A Training Free Approach for Generating Diverse Datasets using LLMsAvinash Amballa, Yashas Malur Saidutta, Chi-Heng Lin, Vivek Kulkarni 等ACL 2026
- SynthesizRR: Generating Diverse Datasets with Retrieval AugmentationAbhishek Divekar, Greg DurrettEMNLP 2024 · 被引用 4 次
- Synthesize, Partition, then Adapt: Eliciting Diverse Samples from Foundation ModelsYeming Wen, Swarat ChaudhuriNeurIPS 2024 · 被引用 1 次
