Hypermodels for Exploration
Vikranth Dwaracherla, Xiuyuan Lu, Morteza Ibrahimi, Ian Osband, Zheng Wen, Benjamin Van Roy
摘要
We study the use of hypermodels to represent epistemic uncertainty and guide exploration. This generalizes and extends the use of ensembles to approximate Thompson sampling. The computational cost of training an ensemble grows with its size, and as such, prior work has typically been limited to ensembles with tens of elements. We show that alternative hypermodels can enjoy dramatic efficiency gains, enabling behavior that would otherwise require hundreds or thousands of elements, and even succeed in situations where ensemble methods fail to learn regardless of size. This allows more accurate approximation of Thompson sampling as well as use of more sophisticated exploration schemes. In particular, we consider an approximate form of information-directed sampling and demonstrate performance gains relative to Thompson sampling. As alternatives to ensembles, we consider linear and neural network hypermodels, also known as hypernetworks. We prove that, with neural network base models, a linear hypermodel can represent essentially any distribution over functions, and as such, hypernetworks do not extend what can be represented.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Epistemic Neural NetworksIan Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla 等NeurIPS 2023 · 被引用 142 次
- Posterior Meta-Replay for Continual LearningChristian Henning, Maria R. Cervera, Francesco D'Angelo, Johannes von Oswald 等NeurIPS 2021 · 被引用 78 次
- Efficient Exploration for LLMsVikranth Dwaracherla, Seyed Mohammad Asghari, Botao Hao, Benjamin Van RoyICML 2024 · 被引用 45 次
- An Analysis of Ensemble SamplingChao Qin, Zheng Wen, Xiuyuan Lu, Benjamin Van RoyNeurIPS 2022 · 被引用 30 次
- Deciding What to Learn: A Rate-Distortion ApproachDilip Arumugam, Benjamin Van RoyICML 2021 · 被引用 29 次
相关 Paper
- Parameterized Indexed Value Function for Efficient Exploration in Reinforcement LearningTian Tan, Zhihan Xiong, Vikranth R. DwaracherlaAAAI 2020 · 被引用 5 次
- Estimating Epistemic and Aleatoric Uncertainty with a Single ModelMatthew Chan, Maria Molina, Chris MetzlerNeurIPS 2024 · 被引用 76 次
- Scalable Exploration via Ensemble++Yingru Li, Jiawei Xu, Baoxiang Wang, Zhi-Quan Tom LuoNeurIPS 2025
- Uncertainty Quantification with the Empirical Neural Tangent KernelJoseph Wilson, Chris van der Heide, Liam Hodgkinson, Fred RoostaNeurIPS 2025 · 被引用 11 次
- HyperDQN: A Randomized Exploration Method for Deep Reinforcement LearningZiniu Li, Yingru Li, Yushun Zhang, Tong Zhang 等ICLR 2022 · 被引用 14 次
