Hexaïssa: Standing on Giants' Shoulders - Routing the Best Chess Engines with Mixture-of-Experts and Latent Reward Learning
Bach Ngo, Nguyen Hoang Khoi Do
Abstract
We present Hexaïssa, a novel framework for adaptive chess engine routing that formulates expert selection as a Mixtureof-Experts (MoE) problem. Hexaïssa learns a gating policy that dynamically selects among heterogeneous state-of-the-art engines such as Stockfish, LCZero, and Obsidian, depending on the tactical and strategic complexity of each board state. This adaptive mechanism enables stronger performance and more efficient computation than any fixed engine or static configuration. However, training such a gating policy is fundamentally challenging due to sparse optimization signals and long-horizon credit assignment in chess games. To address these, we introduce a score-based inverse reinforcement learning (IRL) method that models expert engine trajectories as samples from a latent distribution over optimal behaviors. By recovering the Stein score function of this distribution via stochastic differential equations (SDEs), we infer dense, per-move reward signals consistent with potential-based IRL. These latent rewards allow efficient training of the gating network without requiring additional environment interaction or human supervision. Empirical results on standard chess benchmarks demonstrate that Hexaïssa significantly outperforms individual engines, conventional MoE, and IRL baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 745b5ea0-dd05-42bb-b88b-30ce27d6092dBuilds on13
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann et al.NeurIPS 2021 · 1,213 citations
- IQ-Learn: Inverse soft-Q Learning for ImitationDivyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song et al.NeurIPS 2021 · 271 citations
- Online and Offline Reinforcement Learning by Planning with a Learned ModelJulian Schrittwieser, Thomas Hubert, Amol Mandhane, Mohammadamin Barekatain et al.NeurIPS 2021 · 149 citations
- Policy improvement by planning with GumbelIvo Danihelka, Arthur Guez, Julian Schrittwieser, David SilverICLR 2022 · 84 citations
- Denoising Likelihood Score Matching for Conditional Score-based Data GenerationChen-Hao Chao, Wei-Fang Sun, Bo-Wun Cheng, Yi-Chen Lo et al.ICLR 2022 · 56 citations
Related papers
- Chessformer: A Unified Architecture for Chess ModelingDaniel Monroe, George Eilender, Philip Chalmers, Zhenwei Tang et al.ICLR 2026 · 8 citations
- Evidence of Learned Look-Ahead in a Chess-Playing Neural NetworkErik Jenner, Shreyas Kapur, Vasil Georgiev, Cameron Allen et al.NeurIPS 2024 · 42 citations
- Inherently Explainable Reinforcement Learning in Natural LanguageXiangyu Peng, Mark O. Riedl, Prithviraj AmmanabroluNeurIPS 2022 · 29 citations
- Out-of-Distribution Evaluation of Rule-Based and Strategic Reasoning in Chess TransformersAnna Mészáros, Patrik Reizinger, Ferenc HuszárICML 2026
- RouteMoA: Dynamic Routing without Pre-Inference Boosts Efficient Mixture-of-AgentsJize Wang, Han Wu, Zhiyuan You, Yiming Song et al.ACL 2026 · 2 citations
