Lune

ICML2026Top-tier venue

Neural Logistic Bandits

Seoungbin Bae, Dabeen Lee

2026Year
1Top-tier citations

Abstract

We study the problem of neural logistic bandits, where the main task is to learn an unknown reward function within a logistic link function using a neural network. Existing approaches either exhibit unfavorable dependencies on κ\kappa, where 1/κ1/\kappa represents the minimum variance of reward distributions, or suffer from direct dependence on the feature dimension dd, which can be huge in neural network–based settings. In this work, we introduce a novel Bernstein-type inequality for self-normalized vector-valued martingales that is designed to bypass a direct dependence on the ambient dimension. This lets us deduce a regret upper bound that grows with the effective dimension d~\widetilde{d}, not the feature dimension, while keeping a minimal dependence on κ\kappa. Based on the concentration inequality, we propose two algorithms, NeuralLog-UCB-1 and NeuralLog-UCB-2, that guarantee regret upper bounds of order O~(d~κT)\widetilde{O}(\widetilde{d}\sqrt{\kappa T}) and O~(d~T/κ)\widetilde{O}(\widetilde{d}\sqrt{T/\kappa}), respectively, improving on the existing results. Lastly, we report numerical results on both synthetic and real datasets to validate our theoretical findings.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 7132fac6-5cba-4901-9fca-16c413095007

Cited by top-tier papers1

Ask how each one uses it

Builds on9

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines