Thompson Sampling via Fine-Tuning of LLMs
Nicolas Menet, Aleksandar Terzic, Michael Hersche, Andreas Krause, Abbas Rahimi
摘要
Bayesian optimization in large unstructured discrete spaces is often hindered by the computational cost of maximizing acquisition functions due to the absence of gradients. We propose a scalable alternative based on Thompson sampling that eliminates the need for acquisition function maximization by directly parameterizing the probability that a candidate yields the maximum reward. Our approach, Thompson Sampling via Fine-Tuning (TOSFIT), leverages the prior knowledge embedded in prompt-conditioned large language models, and incrementally adapts them toward the posterior. Theoretically, we derive a novel regret bound for a variational formulation of Thompson Sampling that matches the strong guarantees of its standard counterpart. Our analysis reveals the critical role of careful adaptation to the posterior probability of maximality-a principle that underpins our TOSFIT algorithm. Empirically, we validate our method on three diverse tasks: FAQ response refinement, thermally stable protein search, and quantum circuit design. Within a collection of methods covering in-context Bayesian optimization, reinforcement learning, and evolutionary search, TOSFIT exhibits both state-of-the-art sample efficiency and computational efficiency. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- More Than a Toy: Random Matrix Models Predict How Real-World Neural Representations GeneralizeAlexander Wei, Wei Hu, Jacob SteinhardtICML 2022 · 被引用 90 次
- Improving black-box optimization in VAE latent space using decoder uncertaintyPascal Notin, José Miguel Hernández-Lobato, Yarin GalNeurIPS 2021 · 被引用 76 次
- Back to Basics: Revisiting REINFORCE-Style Optimization for Learning from Human Feedback in LLMsArash Ahmadian, Chris Cremer, Matthias Gallé, Marzieh Fadaee 等ACL 2024 · 被引用 20 次
- Probabilistic Inference in Reinforcement Learning Done RightJean Tarbouriech, Tor Lattimore, Brendan O'DonoghueNeurIPS 2023 · 被引用 15 次
- Variational Bayesian Optimistic SamplingBrendan O'Donoghue, Tor LattimoreNeurIPS 2021 · 被引用 8 次
相关 Paper
- FunBO: Discovering Acquisition Functions for Bayesian Optimization with FunSearchVirginia Aglietti, Ira Ktena, Jessica Schrouff, Eleni Sgouritsa 等ICML 2025
- A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules?Agustinus Kristiadi, Felix Strieth-Kalthoff, Marta Skreta, Pascal Poupart 等ICML 2024 · 被引用 55 次
- Steering Generative Models with Experimental Data for Protein Fitness OptimizationJason Yang, Wenda Chu, Daniel Khalil, Raul Astudillo 等NeurIPS 2025 · 被引用 13 次
- BOTS: A Unified Framework for Bayesian Online Task Selection in LLM Reinforcement FinetuningQianli Shen, Daoyuan Chen, Yilun Huang, Zhenqing Ling 等ICLR 2026 · 被引用 15 次
- Scaling Multi-Task Bayesian Optimization with Large Language ModelsYimeng Zeng, Natalie Maus, Haydn Thomas Jones, Jeffrey Tao 等ICLR 2026 · 被引用 1 次
