Controllable Molecule Generation via Sparse Representation Editing: An Interpretability-Driven Perspective
Zhuoran Li, Xu Sun, Chang Chen, Wanyu LIN
摘要
Controllable molecule generation is crucial for diverse scientific applications, such as drug discovery and materials design. While large language models (LLMs) show great promise, their dense and entangled representations impede precise control over the generation of molecules with bespoke substructures or properties. To address this, we propose Sparse Representation Editing (SpaRE), an interpretability-driven framework for fine-grained and precise control in LLM-based molecule generation. The crux of SpaRE is to learn an overcomplete sparse feature space that disentangles LLM representations into a compact set of latent features corresponding to chemically meaningful concepts. Within this space, we can directly manipulate these concept-aligned latent features to achieve (1) local control, by generating target atoms and functional groups at specified positions; and (2) global control, by customizing the overall structural and physicochemical properties within defined ranges. In this way, our framework advances interpretability from post-hoc analysis to actionable generative control. Experiments show that SpaRE can generate chemically desirable molecules under complex constraints in real-world scenarios, while offering mechanistic insights for quantitative structure-property analysis. The code and demo are available at https://github. com/WanyuGroup/ICML2026_SpaRE.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Multi-Objective Molecule Generation using Interpretable SubstructuresWengong Jin, Regina Barzilay, Tommi S. JaakkolaICML 2020 · 被引用 238 次
- MARS: Markov Molecular Sampling for Multi-objective Drug DiscoveryYutong Xie, Chence Shi, Hao Zhou, Yuwei Yang 等ICLR 2021 · 被引用 186 次
- Translation between Molecules and Natural LanguageCarl Edwards, Tuan Manh Lai, Kevin Ros, Garrett Honke 等EMNLP 2022 · 被引用 112 次
相关 Paper
- Adversarial Representation Engineering: A General Model Editing Framework for Large Language ModelsYihao Zhang, Zeming Wei, Jun Sun, Meng SunNeurIPS 2024 · 被引用 16 次
- MolEditRL: Structure-Preserving Molecular Editing via Discrete Diffusion and Reinforcement LearningYuanxin Zhuang, Dazhong Shen, Ying SunICLR 2026 · 被引用 2 次
- Sparse Fine-Tuning of Transformers for Generative TasksWei Chen, Jingxi Yu, Zichen Miao, Qiang QiuICCV 2025 · 被引用 1 次
- Towards a Unified Paradigm of Concept Editing in Large Language ModelsZhuowen Han, Xinwei Wu, Dan Shi, Renren Jin 等EMNLP 2025
- Learning Retrieval Models with Sparse AutoencodersThibault Formal, Maxime Louis, Hervé Déjean, Stéphane ClinchantICLR 2026 · 被引用 9 次
