Exploring Large Action Sets with Hyperspherical Embeddings using von Mises-Fisher Sampling
Walid Bendada, Guillaume Salha-Galvan, Romain Hennequin, Théo Bontempelli, Thomas Bouabça, Tristan Cazenave
摘要
This paper introduces von Mises-Fisher exploration (vMF-exp), a scalable method for exploring large action sets in reinforcement learning problems where hyperspherical embedding vectors represent these actions. vMF-exp involves initially sampling a state embedding representation using a von Mises-Fisher distribution, then exploring this representation's nearest neighbors, which scales to virtually unlimited numbers of candidate actions. We show that, under theoretical assumptions, vMF-exp asymptotically maintains the same probability of exploring each action as Boltzmann Exploration (B-exp), a popular alternative that, nonetheless, suffers from scalability issues as it requires computing softmax values for each action. Consequently, vMF-exp serves as a scalable alternative to B-exp for exploring large action sets with hyperspherical embeddings. Experiments on simulated data, real-world public data, and the successful large-scale deployment of vMF-exp on the recommender system of a global music streaming service empirically validate the key properties of the proposed method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Beyond Distributions: Geometric Action Control for Continuous Reinforcement LearningZhihao LinICLR 2026 · 被引用 1 次
- Polaris: Coupled Orbital Polar Embeddings for Hierarchical Concept LearningSahil Mishra, Srinitish Srinivasan, Sourish Dasgupta, Tanmoy ChakrabortyICML 2026
它引用的顶会 Paper6
- Accelerating Large-Scale Inference with Anisotropic Vector QuantizationRuiqi Guo, Philip Sun, Erik Lindgren, Quan Geng 等ICML 2020 · 被引用 539 次
- Reward-Free Exploration for Reinforcement LearningChi Jin, Akshay Krishnamurthy, Max Simchowitz, Tiancheng YuICML 2020 · 被引用 226 次
- Guarantees for Epsilon-Greedy Reinforcement Learning with Function ApproximationChristoph Dann, Yishay Mansour, Mehryar Mohri, Ayush Sekhari 等ICML 2022 · 被引用 76 次
- Adap-τ : Adaptively Modulating Embedding Magnitude for RecommendationJiawei Chen, Junkang Wu, Jiancan Wu, Xuezhi Cao 等WWW 2023 · 被引用 48 次
- Latent exploration for Reinforcement LearningAlberto Silvio Chiappa, Alessandro Marin Vargas, Ann Zixiang Huang, Alexander MathisNeurIPS 2023 · 被引用 41 次
相关 Paper
- Scalable Exploration for High-Dimensional Continuous Control via Value-Guided FlowYunyue Wei, Chenhui Zuo, Yanan SuiICLR 2026 · 被引用 8 次
- Dynamic Deep Clustering of High-Dimensional Directional Data via Hyperspherical Embeddings with Bayesian Nonparametric MixturesZhiwen Luo, Wentao Fan, Manar Amayri, Nizar BouguilaKDD 2025 · 被引用 4 次
- Exploration and Regularization of the Latent Action Space in RecommendationShuchang Liu, Qingpeng Cai, Bowen Sun, Yuhao Wang 等WWW 2023 · 被引用 54 次
- VibeSpace: Automatic Generation of Data and Vector Embeddings for Arbitrary Domains and Cross-domain Mappings using LLMsKipp Freud, Daniel E. Collins, Delmiro D. Sampaio Neto, Grant StevensACM MM 2025
- Latent Variable Representation for Reinforcement LearningTongzheng Ren, Chenjun Xiao, Tianjun Zhang, Na Li 等ICLR 2023 · 被引用 1 次
