Planning from Pixels in Atari with Learned Symbolic Representations
Andrea Dittadi, Frederik K. Drachmann, Thomas Bolander
Abstract
Width-based planning methods have been shown to yield state-of-the-art performance in the Atari 2600 domain using pixel input. One successful approach, RolloutIW, represents states with the B-PROST boolean feature set. An augmented version of RolloutIW, pi-IW, shows that learned features can be competitive with handcrafted ones for width-based search. In this paper, we leverage variational autoencoders (VAEs) to learn features directly from pixels in a principled manner, and without supervision. The inference model of the trained VAEs extracts boolean features from pixels, and RolloutIW plans with these features. The resulting combination outperforms the original RolloutIW and human professional play on Atari 2600 and drastically reduces the size of the feature set.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- The Role of Pretrained Representations for the OOD Generalization of RL AgentsFrederik Träuble, Andrea Dittadi, Manuel Wuthrich, Felix Widmaier et al.ICLR 2022 · 19 citations
- Symbolic Distillation for Learned TCP Congestion ControlS. P. Sharan, Wenqing Zheng, Kuo-Feng Hsu, Jiarong Xing et al.NeurIPS 2022 · 9 citations
- Width-based Lookaheads with Learnt Base Policies and Heuristics Over the Atari-2600 BenchmarkStefan O'Toole, Nir Lipovetzky, Miquel Ramírez, Adrian R. PearceNeurIPS 2021
Builds on3
- NVAE: A Deep Hierarchical Variational AutoencoderArash Vahdat, Jan KautzNeurIPS 2020 · 1,141 citations
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski et al.ICLR 2020 · 969 citations
- Optimal Variance Control of the Score-Function Gradient Estimator for Importance-Weighted BoundsValentin Liévin, Andrea Dittadi, Anders Christensen, Ole WintherNeurIPS 2020 · 9 citations
Related papers
- Discovering diverse athletic jumping strategiesZhiqi Yin, Zeshi Yang, Michiel van de Panne, KangKang YinSIGGRAPH 2021 · 46 citations
- Scalable Decision-Making in Stochastic Environments through Learned Temporal AbstractionBaiting Luo, Ava Pettet, Aron Laszka, Abhishek Dubey et al.ICLR 2025
- Efficient Planning in a Compact Latent Action SpaceZhengyao Jiang, Tianjun Zhang, Michael Janner, Yueying Li et al.ICLR 2023 · 3 citations
- In Pursuit of Pixel Supervision for Visual Pre-trainingLihe Yang, Shang-Wen Li, Yang Li, Xinjie Lei et al.CVPR 2026 · 13 citations
- General Policies, Representations, and Planning WidthBlai Bonet, Hector GeffnerAAAI 2021 · 28 citations
