CLaDMoP: Learning Transferrable Models from Successful Clinical Trials via LLMs
Yiqing Zhang, Xiaozhong Liu, Fabricio Murai
摘要
Many existing models for clinical trial outcome prediction are optimized using task-specific loss functions on trial phase-specific data. While this scheme may boost prediction for common diseases and drugs, it can hinder learning of generalizable representations, leading to more false positives/negatives. To address this limitation, we introduce CLaDMoP, a new pre-training approach for clinical trial outcome prediction, alongside the Successful Clinical Trials dataset (SCT), specifically designed for this task. CLaDMoP leverages a Large Language Model-to encode trials' eligibility criteria-linked to a lightweight Drug-Molecule branch through a novel multi-level fusion technique. To efficiently fuse long embeddings across levels, we incorporate a grouping block, drastically reducing computational overhead. CLaDMoP avoids reliance on task-specific objectives by pre-training on a "pair matching" proxy task. Compared to established zero-shot and few-shot baselines, our method significantly improves both PR-AUC and ROC-AUC, especially for phase I and phase II trials. We further evaluate and perform ablation on CLaDMoP after Parameter-Efficient Fine-Tuning, comparing it to state-of-the-art supervised baselines, including MEXA-CTP, on the Trial Outcome Prediction (TOP) benchmark. CLaDMoP achieves up to 10.5% improvement in PR-AUC and 3.6% in ROC-AUC, while attaining comparable F1 score to MEXA-CTP, highlighting its potential for clinical trial outcome prediction. Code and SCT dataset can be downloaded from https://github.com/murai-lab/CLaDMoP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Rapid Learning or Feature Reuse? Towards Understanding the Effectiveness of MAMLAniruddh Raghu, Maithra Raghu, Samy Bengio, Oriol VinyalsICLR 2020 · 被引用 736 次
- Improving CLIP Training with Language RewritesLijie Fan, Dilip Krishnan, Phillip Isola, Dina Katabi 等NeurIPS 2023 · 被引用 308 次
- Why do We Need Large Batchsizes in Contrastive Learning? A Gradient-Bias PerspectiveChangyou Chen, Jianyi Zhang, Yi Xu, Liqun Chen 等NeurIPS 2022 · 被引用 61 次
相关 Paper
- AutoCT: Automating Interpretable Clinical Trial Prediction with LLM AgentsFengze Liu, Haoyu Wang, Joonhyuk Cho, Dan Roth 等EMNLP 2025 · 被引用 1 次
- Predicting Clinical Trial Results by Implicit Evidence IntegrationQiao Jin, Chuanqi Tan, Mosha Chen, Xiaozhong Liu 等EMNLP 2020 · 被引用 6 次
- Enhancing Activity Prediction Models in Drug Discovery with the Ability to Understand Human LanguagePhilipp Seidl, Andreu Vall, Sepp Hochreiter, Günter KlambauerICML 2023 · 被引用 69 次
- Generalizable Drug-Target Interaction Prediction via ESM-2 Representations and Progressive Contrastive Curriculum LearningQianyang Wu, Jingwei Lv, Zilong Zhang, Feifei CuiAAAI 2026
- An Empirical Investigation Towards Efficient Multi-Domain Language Model Pre-trainingKristjan Arumae, Qing Sun, Parminder BhatiaEMNLP 2020
