AutoCT: Automating Interpretable Clinical Trial Prediction with LLM Agents
Fengze Liu, Haoyu Wang, Joonhyuk Cho, Dan Roth, Andrew Lo
Abstract
Clinical trials are critical for advancing medical treatments but remain prohibitively expensive and time-consuming. Accurate prediction of clinical trial outcomes can significantly reduce research and development costs and accelerate drug discovery. While recent deep learning models have shown promise by leveraging unstructured data, their black-box nature, lack of interpretability, and vulnerability to label leakage limit their practical use in high-stakes biomedical contexts. In this work, we propose AutoCT, a novel framework that combines the reasoning capabilities of large language models with the explainability of classical machine learning. AutoCT autonomously generates, evaluates, and refines tabular features based on public information without human input. Our method uses Monte Carlo Tree Search to iteratively optimize predictive performance. Experimental results show that AutoCT performs on par with or better than SOTA methods on clinical trial prediction tasks within only a limited number of self-refinement iterations, establishing a new paradigm for scalable, interpretable, and cost-efficient clinical trial prediction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li et al.ICLR 2024 · 867 citations
- MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-MakingYubin Kim, Chanwoo Park, Hyewon Jeong, Yik Siu Chan et al.NeurIPS 2024 · 291 citations
- Large Language Models Can Automatically Engineer Features for Few-Shot Tabular LearningSungwon Han, Jinsung Yoon, Sercan Ö. Arik, Tomas PfisterICML 2024 · 81 citations
- COMPOSE: Cross-Modal Pseudo-Siamese Network for Patient Trial MatchingJunyi Gao, Cao Xiao, Lucas M. Glass, Jimeng SunKDD 2020 · 51 citations
- AutoM3L: An Automated Multimodal Machine Learning Framework with Large Language ModelsDaqin Luo, Chengjian Feng, Yuxuan Nong, Yiqing ShenACM MM 2024 · 16 citations
Related papers
- Optimized Feature Generation for Tabular Data via LLMs with Decision Tree ReasoningJaehyun Nam, Kyuyoung Kim, Seunghyuk Oh, Jihoon Tack et al.NeurIPS 2024 · 78 citations
- AutoTrial: Prompting Language Models for Clinical Trial DesignZifeng Wang, Cao Xiao, Jimeng SunEMNLP 2023 · 14 citations
- CoFE: Collaborative Feature Engineering via Semantically-Guided Exploration and Diagnostic-Driven RefinementWeihao Jiang, Ziang Nan, Zhihui Shi, Ya Cong et al.KDD 2026
- Timely Clinical Diagnosis through Active Test SelectionSilas Ruhrberg Estévez, Nicolás Astorga, Mihaela van der SchaarNeurIPS 2025 · 4 citations
- Autoformulation of Mathematical Optimization Models Using LLMsNicolás Astorga, Tennison Liu, Yuanzhang Xiao, Mihaela van der SchaarICML 2025
