In-Context Learning for Pure Exploration
Alessio Russo, Ryan Welch, Aldo Pacchiano
Abstract
We study the active sequential hypothesis testing problem, also known as pure exploration: given a new task, the learner adaptively collects data from the environment to efficiently determine an underlying correct hypothesis. A classical instance of this problem is the task of identifying the best arm in a multi-armed bandit problem (a.k.a. BAI, Best-Arm Identification), where actions index hypotheses. Another important case is generalized search, a problem of determining the correct label through a sequence of strategically selected queries that indirectly reveal information about the label. In this work, we introduce In-Context Pure Explorer (ICPE), which meta-trains Transformers to map observation histories to query actions and a predicted hypothesis, yielding a model that transfers in-context. At inference time, ICPE actively gathers evidence on new tasks and infers the true hypothesis without parameter updates. Across deterministic, stochastic, and structured benchmarks, including BAI and generalized search, ICPE is competitive with adaptive baselines while requiring no explicit modeling of information structure. Our results support Transformers as practical architectures for general sequential testing. Code repository https://github.com/rssalessio/icpe . * Equal contribution (alphabetical order). 1 Note that, in this particular case, the hypothesis space coincides with the query space of the agent.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 996005a0-e811-48b1-b03e-a4babc680ae1Cited by top-tier papers3
- Large Language Models Think Too Fast To Explore EffectivelyLan Pan, Hanbo Xie, Robert C. WilsonNeurIPS 2025 · 10 citations
- Robust In-Context Reinforcement Learning Under Reward Poisoning AttacksPaulius Sasnauskas, Yiğit Yalın, Goran RadanovicICML 2026
- Asymptotically Optimal Sequential Testing with Markovian DataAlhad Sethi, SOFIA SAGAR KAVALI, Shubhada Agrawal, Debabrota Basu et al.ICML 2026
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
- RvS: What is Essential for Offline RL via Supervised Learning?Scott Emmons, Benjamin Eysenbach, Ilya Kostrikov, Sergey LevineICLR 2022 · 225 citations
- Supervised Pretraining Can Learn In-Context Reinforcement LearningJonathan Lee, Annie Xie, Aldo Pacchiano, Yash Chandak et al.NeurIPS 2023 · 170 citations
Related papers
- In-Context Learning through the Bayesian PrismMadhur Panwar, Kabir Ahuja, Navin GoyalICLR 2024 · 79 citations
- Transformers as Algorithms: Generalization and Stability in In-context LearningYingcong Li, Muhammed Emrullah Ildiz, Dimitris Papailiopoulos, Samet OymakICML 2023 · 242 citations
- Unsupervised Meta-Learning via In-Context LearningAnna Vettoruzzo, Lorenzo Braccaioli, Joaquin Vanschoren, Marlena NowaczykICLR 2025
- Technical Debt in In-Context Learning: Diminishing Efficiency in Long ContextTaejong Joo, Diego KlabjanNeurIPS 2025
- Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm SelectionYu Bai, Fan Chen, Huan Wang, Caiming Xiong et al.NeurIPS 2023 · 356 citations
