Hexaïssa: Standing on Giants' Shoulders - Routing the Best Chess Engines with Mixture-of-Experts and Latent Reward Learning
Bach Ngo, Nguyen Hoang Khoi Do
摘要
We present Hexaïssa, a novel framework for adaptive chess engine routing that formulates expert selection as a Mixtureof-Experts (MoE) problem. Hexaïssa learns a gating policy that dynamically selects among heterogeneous state-of-the-art engines such as Stockfish, LCZero, and Obsidian, depending on the tactical and strategic complexity of each board state. This adaptive mechanism enables stronger performance and more efficient computation than any fixed engine or static configuration. However, training such a gating policy is fundamentally challenging due to sparse optimization signals and long-horizon credit assignment in chess games. To address these, we introduce a score-based inverse reinforcement learning (IRL) method that models expert engine trajectories as samples from a latent distribution over optimal behaviors. By recovering the Stein score function of this distribution via stochastic differential equations (SDEs), we infer dense, per-move reward signals consistent with potential-based IRL. These latent rewards allow efficient training of the gating network without requiring additional environment interaction or human supervision. Empirical results on standard chess benchmarks demonstrate that Hexaïssa significantly outperforms individual engines, conventional MoE, and IRL baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann 等NeurIPS 2021 · 被引用 1,213 次
- IQ-Learn: Inverse soft-Q Learning for ImitationDivyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song 等NeurIPS 2021 · 被引用 271 次
- Online and Offline Reinforcement Learning by Planning with a Learned ModelJulian Schrittwieser, Thomas Hubert, Amol Mandhane, Mohammadamin Barekatain 等NeurIPS 2021 · 被引用 149 次
- Policy improvement by planning with GumbelIvo Danihelka, Arthur Guez, Julian Schrittwieser, David SilverICLR 2022 · 被引用 84 次
- Denoising Likelihood Score Matching for Conditional Score-based Data GenerationChen-Hao Chao, Wei-Fang Sun, Bo-Wun Cheng, Yi-Chen Lo 等ICLR 2022 · 被引用 56 次
相关 Paper
- Chessformer: A Unified Architecture for Chess ModelingDaniel Monroe, George Eilender, Philip Chalmers, Zhenwei Tang 等ICLR 2026 · 被引用 8 次
- Evidence of Learned Look-Ahead in a Chess-Playing Neural NetworkErik Jenner, Shreyas Kapur, Vasil Georgiev, Cameron Allen 等NeurIPS 2024 · 被引用 42 次
- Inherently Explainable Reinforcement Learning in Natural LanguageXiangyu Peng, Mark O. Riedl, Prithviraj AmmanabroluNeurIPS 2022 · 被引用 29 次
- Out-of-Distribution Evaluation of Rule-Based and Strategic Reasoning in Chess TransformersAnna Mészáros, Patrik Reizinger, Ferenc HuszárICML 2026
- RouteMoA: Dynamic Routing without Pre-Inference Boosts Efficient Mixture-of-AgentsJize Wang, Han Wu, Zhiyuan You, Yiming Song 等ACL 2026 · 被引用 2 次
