Positive-Unlabeled Learning using Random Forests via Recursive Greedy Risk Minimization
Jonathan Wilton, Abigail M. Y. Koay, Ryan K. L. Ko, Miao Xu, Nan Ye
Abstract
The need to learn from positive and unlabeled data, or PU learning, arises in many applications and has attracted increasing interest. While random forests are known to perform well on many tasks with positive and negative data, recent PU algorithms are generally based on deep neural networks, and the potential of tree-based PU learning is under-explored. In this paper, we propose new random forest algorithms for PU-learning. Key to our approach is a new interpretation of decision tree algorithms for positive and negative data as recursive greedy risk minimization algorithms. We extend this perspective to the PU setting to develop new decision tree learning algorithms that directly minimizes PU-data based estimators for the expected risk. This allows us to develop an efficient PU random forest algorithm, PU extra trees. Our approach features three desirable properties: it is robust to the choice of the loss function in the sense that various loss functions lead to the same decision trees; it requires little hyperparameter tuning as compared to neural network based PU learning; it supports a feature importance that directly measures a feature's contribution to risk minimization. Our algorithms demonstrate strong performance on several datasets. Our code is available at https://github.com/puetpaper/PUExtraTrees.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Beyond Myopia: Learning from Positive and Unlabeled Data through Holistic Predictive TrendsXinrui Wang, Wenhai Wan, Chuanxing Geng, Shaoyuan Li et al.NeurIPS 2023 · 24 citations
- Class Prior-Free Positive-Unlabeled Learning with Taylor Variational Loss for Hyperspectral Remote Sensing ImageryHengwei Zhao, Xinyu Wang, Jingtao Li, Yanfei ZhongICCV 2023 · 15 citations
- Robust Loss Functions for Training Decision Trees with Noisy LabelsJonathan Wilton, Nan YeAAAI 2024 · 8 citations
- Regression with Sensor Data Containing Incomplete ObservationsTakayuki Katsuki, Takayuki OsogamiICML 2023 · 1 citation
- Noisy-Pair Robust Representation Alignment for Positive-Unlabeled LearningHengwei Zhao, Zhengzhong Tu, Zhuo Zheng, Wei Wang et al.ICLR 2026
Builds on1
Related papers
- Balancing Positive and Negative Classification Error Rates in Positive-Unlabeled LearningXiming Li, Yuanchao Dai, Bing Wang, Changchun Li et al.NeurIPS 2025 · 3 citations
- A Variational Approach for Learning from Positive and Unlabeled DataHui Chen, Fangqing Liu, Yin Wang, Liyue Zhao et al.NeurIPS 2020 · 76 citations
- Learning from Positive and Unlabeled Data with Arbitrary Positive ShiftZayd Hammoudeh, Daniel LowdNeurIPS 2020 · 53 citations
- A Closer Look to Positive-Unlabeled Learning from Fine-grained Perspectives: An Empirical StudyYuanchao Dai, Zhengzhang Hou, Changchun Li, Yuanbo Xu et al.NeurIPS 2025 · 2 citations
- Improving Neural Relation Extraction with Positive and Unlabeled LearningZhengqiu He, Wenliang Chen, Yuyi Wang, Wei Zhang et al.AAAI 2020 · 18 citations
