Lune

AAAI2026顶会

Hexaïssa: Standing on Giants' Shoulders - Routing the Best Chess Engines with Mixture-of-Experts and Latent Reward Learning

Bach Ngo, Nguyen Hoang Khoi Do

2026年份

摘要

We present Hexaïssa, a novel framework for adaptive chess engine routing that formulates expert selection as a Mixtureof-Experts (MoE) problem. Hexaïssa learns a gating policy that dynamically selects among heterogeneous state-of-the-art engines such as Stockfish, LCZero, and Obsidian, depending on the tactical and strategic complexity of each board state. This adaptive mechanism enables stronger performance and more efficient computation than any fixed engine or static configuration. However, training such a gating policy is fundamentally challenging due to sparse optimization signals and long-horizon credit assignment in chess games. To address these, we introduce a score-based inverse reinforcement learning (IRL) method that models expert engine trajectories as samples from a latent distribution over optimal behaviors. By recovering the Stein score function of this distribution via stochastic differential equations (SDEs), we infer dense, per-move reward signals consistent with potential-based IRL. These latent rewards allow efficient training of the gating network without requiring additional environment interaction or human supervision. Empirical results on standard chess benchmarks demonstrate that Hexaïssa significantly outperforms individual engines, conventional MoE, and IRL baselines.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper13

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖