Tanimoto Random Features for Scalable Molecular Machine Learning
Austin Tripp, Sergio Bacallado, Sukriti Singh, José Miguel Hernández-Lobato
Abstract
The Tanimoto coefficient is commonly used to measure the similarity between molecules represented as discrete fingerprints, either as a distance metric or a positive definite kernel. While many kernel methods can be accelerated using random feature approximations, at present there is a lack of such approximations for the Tanimoto kernel. In this paper we propose two kinds of novel random features to allow this kernel to scale to large datasets, and in the process discover a novel extension of the kernel to real-valued vectors. We theoretically characterize these random features, and provide error bounds on the spectral norm of the Gram matrix. Experimentally, we show that these random features are effective at approximating the Tanimoto coefficient of real-world datasets and are useful for molecular property prediction and optimization tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Stochastic Gradient Descent for Gaussian Processes Done RightJihao Andreas Lin, Shreyas Padhy, Javier Antorán, Austin Tripp et al.ICLR 2024 · 17 citations
- Retro-fallback: retrosynthetic planning in an uncertain worldAustin Tripp, Krzysztof Maziarz, Sarah Lewis, Marwin H. S. Segler et al.ICLR 2024 · 14 citations
- Flexible Kernels for Protein Property PredictionMartin Jankowiak, Yerdos Ordabayev, Rudraksh Tuwani, Henry Ward et al.ICML 2026
- Variance-Reducing Couplings for Random FeaturesIsaac Reid, Stratis Markou, Krzysztof Marcin Choromanski, Richard E. Turner et al.ICLR 2025
- Return of the Latent Space COWBOYS: Re-thinking the use of VAEs for Bayesian Optimisation of Structured SpacesHenry B. Moss, Sebastian W. Ober, Tom DietheICML 2025
Builds on4
- Efficiently sampling functions from Gaussian process posteriorsJames T. Wilson, Viacheslav Borovitskiy, Alexander Terenin, Peter Mostowsky et al.ICML 2020 · 186 citations
- Oblivious Sketching of High-Degree Polynomial KernelsThomas D. Ahle, Michael Kapralov, Jakob Bæk Tejs Knudsen, Rasmus Pagh et al.SODA 2020 · 42 citations
- Scaling Neural Tangent Kernels via Sketching and Random FeaturesAmir Zandieh, Insu Han, Haim Avron, Neta Shoham et al.NeurIPS 2021 · 42 citations
- Bayesian Algorithm Execution: Estimating Computable Properties of Black-box Functions Using Mutual InformationWillie Neiswanger, Ke Alexander Wang, Stefano ErmonICML 2021 · 40 citations
Related papers
- TRF: Learning Kernels with Tuned Random FeaturesAlistair Shilton, Sunil Gupta, Santu Rana, Arun Kumar Anjanapura Venkatesh et al.AAAI 2022
- Taming graph kernels with random featuresKrzysztof Marcin ChoromanskiICML 2023 · 21 citations
- Weisfeiler and Leman Go Walking: Random Walk Kernels RevisitedNils M. KriegeNeurIPS 2022 · 22 citations
- Learning with Optimized Random Features: Exponential Speedup by Quantum Machine Learning without Sparsity and Low-Rank AssumptionsHayata Yamasaki, Sathyawageeswar Subramanian, Sho Sonoda, Masato KoashiNeurIPS 2020 · 23 citations
- Measuring dissimilarity with diffeomorphism invarianceThéophile Cantelobre, Carlo Ciliberto, Benjamin Guedj, Alessandro RudiICML 2022 · 1 citation
