'Oh LLM, I'm Asking Thee, Please Give Me a Decision Tree': Zero-Shot Decision Tree Induction and Embedding with Large Language Models
Ricardo Knauer, Mario Koddenbrock, Raphael Wallsberger, Nicholas M. Brisson, Georg N. Duda, Deborah Falla, David W. Evans, Erik Rodner
Abstract
Large language models (LLMs) provide powerful means to leverage prior knowledge for predictive modeling when data is limited. In this work, we demonstrate how LLMs can use their compressed world knowledge to generate intrinsically interpretable machine learning models, i.e., decision trees, without any training data. We find that these zero-shot decision trees can even surpass data-driven trees on some small-sized tabular datasets and that embeddings derived from these trees perform better than data-driven tree-based embeddings on average. Our decision tree induction and embedding approaches can therefore serve as new knowledge-driven baselines for data-driven machine learning methods in the low-data regime. Furthermore, they offer ways to harness the rich world knowledge within LLMs for tabular machine learning tasks. Our code and results are available at https://github.com/ml-lab-htw/ llm-trees.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5ffc6a7f-d70f-4c5f-9daf-71f616c73c66Cited by top-tier papers1
Ask how each one uses itBuilds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 1,030 citations
Related papers
- Language Models are Weak LearnersHariharan Manikandan, Yiding Jiang, J. Zico KolterNeurIPS 2023 · 32 citations
- Optimized Feature Generation for Tabular Data via LLMs with Decision Tree ReasoningJaehyun Nam, Kyuyoung Kim, Seunghyuk Oh, Jihoon Tack et al.NeurIPS 2024 · 78 citations
- Large Language Models Can Automatically Engineer Features for Few-Shot Tabular LearningSungwon Han, Jinsung Yoon, Sercan Ö. Arik, Tomas PfisterICML 2024 · 81 citations
- LLM Meeting Decision Trees on Tabular DataHangting Ye, Jinmeng Li, He Zhao, Dandan Guo et al.NeurIPS 2025 · 8 citations
- Efficient Reinforcement Learning with Large Language Model PriorsXue Yan, Yan Song, Xidong Feng, Mengyue Yang et al.ICLR 2025
