Contextualizing biological perturbation experiments through language
Menghua Wu, Russell Littman, Jacob Levine, Lin Qiu, Tommaso Biancalani, David Richmond, Jan-Christian Huetter
摘要
High-content genetic perturbation experiments provide insights into biomolecular pathways at unprecedented resolution, yet experimental and analysis costs pose barriers to their widespread adoption. In-silico modeling of unseen perturbations has the potential to alleviate this burden by leveraging prior knowledge to enable more efficient exploration of the perturbation space. However, current knowledgegraph approaches neglect the semantic richness of the relevant biology, beyond simple adjacency graphs. To enable holistic modeling, we hypothesize that natural language is an appropriate medium for interrogating experimental outcomes and representing biological relationships. We propose PERTURBQA as a set of real-world tasks for benchmarking large language model (LLM) reasoning over structured, biological data. PERTURBQA is comprised of three tasks: prediction of differential expression and change of direction for unseen perturbations, and gene set enrichment. As a proof of concept, we present SUMMER (SUMMarize, retrievE, and answeR), a simple LLM-based framework that matches or exceeds the current state-of-the-art on this benchmark. We evaluated graph and language-based models on differential expression and direction of change tasks, finding that SUMMER performed best overall. Notably, SUMMER's outputs, unlike models that solely rely on knowledge graphs, are easily interpretable by domain experts, aiding in understanding model limitations and contextualizing experimental outcomes. Additionally, SUMMER excels in gene set enrichment, surpassing over-representation analysis baselines in most cases and effectively summarizing clusters lacking a manual annotation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- VCWorld: A Biological World Model for Virtual Cell SimulationZhijian Wei, Runze Ma, Zichen Wang, Zhongmin Li 等ICLR 2026 · 被引用 18 次
- Learning Adaptive Perturbation-Conditioned Contexts for Robust Transcriptional Response PredictionYinhua Piao, Hyomin Kim, SEONGHWAN KIM, Yunhak Oh 等ICML 2026 · 被引用 1 次
- DC-W2S: Dual-Consensus Weak-to-Strong Training for Reliable Process Reward Modeling in Biological ReasoningChi-Min Chan, Ehsan Hajiramezanali, Xiner Li, Edward De Brouwer 等ICML 2026 · 被引用 1 次
- Judge and Improve: Towards a Better Reasoning of Knowledge Graphs with Large Language ModelsMo Zhiqiang, Yang Hua, Jiahui Li, Yuan Liu 等EMNLP 2025
它引用的顶会 Paper6
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger 等AAAI 2024 · 被引用 1,292 次
- Caduceus: Bi-Directional Equivariant Long-Range DNA Sequence ModelingYair Schiff, Chia-Hsiang Kao, Aaron Gokaslan, Tri Dao 等ICML 2024 · 被引用 195 次
- BioBridge: Bridging Biomedical Foundation Models via Knowledge GraphsZifeng Wang, Zichen Wang, Balasubramaniam Srinivasan, Vassilis N. Ioannidis 等ICLR 2024 · 被引用 29 次
相关 Paper
- Assessing LLMs for Serendipity Discovery in Knowledge Graphs: A Case for Drug RepurposingMengying Wang, Chenhui Ma, Ao Jiao, Tuo Liang 等AAAI 2026
- BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation ExperimentsYusuf H. Roohani, Andrew H. Lee, Qian Huang, Jian Vora 等ICLR 2025
- GenomeQA: Benchmarking General Large Language Models for Genome Sequence UnderstandingWeicai Long, Yusen Hou, Junning Feng, Houcheng Su 等ACL 2026
- MolecularIQ: Characterizing Chemical Reasoning Capabilities Through Symbolic Verification on Molecular GraphsChristoph Bartmann, Johannes Schimunek, Mykyta Ielanskyi, Philipp Seidl 等ICLR 2026 · 被引用 5 次
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation ExplainersAdam Karvonen, James Chua, Clément Dumas, Kit Fraser-Taliente 等ICML 2026 · 被引用 42 次
