Exploring Large Action Sets with Hyperspherical Embeddings using von Mises-Fisher Sampling
Walid Bendada, Guillaume Salha-Galvan, Romain Hennequin, Théo Bontempelli, Thomas Bouabça, Tristan Cazenave
Abstract
This paper introduces von Mises-Fisher exploration (vMF-exp), a scalable method for exploring large action sets in reinforcement learning problems where hyperspherical embedding vectors represent these actions. vMF-exp involves initially sampling a state embedding representation using a von Mises-Fisher distribution, then exploring this representation's nearest neighbors, which scales to virtually unlimited numbers of candidate actions. We show that, under theoretical assumptions, vMF-exp asymptotically maintains the same probability of exploring each action as Boltzmann Exploration (B-exp), a popular alternative that, nonetheless, suffers from scalability issues as it requires computing softmax values for each action. Consequently, vMF-exp serves as a scalable alternative to B-exp for exploring large action sets with hyperspherical embeddings. Experiments on simulated data, real-world public data, and the successful large-scale deployment of vMF-exp on the recommender system of a global music streaming service empirically validate the key properties of the proposed method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Beyond Distributions: Geometric Action Control for Continuous Reinforcement LearningZhihao LinICLR 2026 · 1 citation
- Polaris: Coupled Orbital Polar Embeddings for Hierarchical Concept LearningSahil Mishra, Srinitish Srinivasan, Sourish Dasgupta, Tanmoy ChakrabortyICML 2026
Builds on6
- Accelerating Large-Scale Inference with Anisotropic Vector QuantizationRuiqi Guo, Philip Sun, Erik Lindgren, Quan Geng et al.ICML 2020 · 539 citations
- Reward-Free Exploration for Reinforcement LearningChi Jin, Akshay Krishnamurthy, Max Simchowitz, Tiancheng YuICML 2020 · 226 citations
- Guarantees for Epsilon-Greedy Reinforcement Learning with Function ApproximationChristoph Dann, Yishay Mansour, Mehryar Mohri, Ayush Sekhari et al.ICML 2022 · 76 citations
- Adap-τ : Adaptively Modulating Embedding Magnitude for RecommendationJiawei Chen, Junkang Wu, Jiancan Wu, Xuezhi Cao et al.WWW 2023 · 48 citations
- Latent exploration for Reinforcement LearningAlberto Silvio Chiappa, Alessandro Marin Vargas, Ann Zixiang Huang, Alexander MathisNeurIPS 2023 · 41 citations
Related papers
- Scalable Exploration for High-Dimensional Continuous Control via Value-Guided FlowYunyue Wei, Chenhui Zuo, Yanan SuiICLR 2026 · 8 citations
- Dynamic Deep Clustering of High-Dimensional Directional Data via Hyperspherical Embeddings with Bayesian Nonparametric MixturesZhiwen Luo, Wentao Fan, Manar Amayri, Nizar BouguilaKDD 2025 · 4 citations
- Exploration and Regularization of the Latent Action Space in RecommendationShuchang Liu, Qingpeng Cai, Bowen Sun, Yuhao Wang et al.WWW 2023 · 54 citations
- VibeSpace: Automatic Generation of Data and Vector Embeddings for Arbitrary Domains and Cross-domain Mappings using LLMsKipp Freud, Daniel E. Collins, Delmiro D. Sampaio Neto, Grant StevensACM MM 2025
- Latent Variable Representation for Reinforcement LearningTongzheng Ren, Chenjun Xiao, Tianjun Zhang, Na Li et al.ICLR 2023 · 1 citation
