ICML2026
Theoretical Investigation on Inductive Bias of Isolation Forest
Qin-Cheng Zheng, Shao-Qun Zhang, Shen-Huan Lyu, Yuan Jiang, Zhi-Hua Zhou
1 citation
Abstract
Isolation Forest (iForest) is one of the most widely used unsupervised anomaly detectors, owing to its efficiency and performance on large-scale tasks. Despite its broad applications, there is still a lack of theoretical understanding of iForest's empirical success. In this work, we study the inductive bias of iForest and examine when and to what extent it performs well. The main idea is to characterize the random growth process of iForest, in which both split dimensions and split values are selected randomly. We model the growth process of iForest as a random walk and derive the expected path length function, the outcome of iForest that determines the anomaly score, by analyzing the hitting time of the absorbing state. The infinite-sample size analysis reveals that, unlike -Nearest Neighbor (-NN), whose score reflects only the local density, the iForest path length combines the density and the centrality. Since central points naturally have larger path lengths, iForest is therefore less sensitive to central anomalies. Analyses of fixed datasets corroborate this finding and further show that iForest is more parameter-adaptive than -NN. Our study provides a theoretical understanding of the effectiveness of iForest and establishes a foundation for further exploration.