Planning from Pixels in Atari with Learned Symbolic Representations
Andrea Dittadi, Frederik K. Drachmann, Thomas Bolander
摘要
Width-based planning methods have been shown to yield state-of-the-art performance in the Atari 2600 domain using pixel input. One successful approach, RolloutIW, represents states with the B-PROST boolean feature set. An augmented version of RolloutIW, pi-IW, shows that learned features can be competitive with handcrafted ones for width-based search. In this paper, we leverage variational autoencoders (VAEs) to learn features directly from pixels in a principled manner, and without supervision. The inference model of the trained VAEs extracts boolean features from pixels, and RolloutIW plans with these features. The resulting combination outperforms the original RolloutIW and human professional play on Atari 2600 and drastically reduces the size of the feature set.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- The Role of Pretrained Representations for the OOD Generalization of RL AgentsFrederik Träuble, Andrea Dittadi, Manuel Wuthrich, Felix Widmaier 等ICLR 2022 · 被引用 19 次
- Symbolic Distillation for Learned TCP Congestion ControlS. P. Sharan, Wenqing Zheng, Kuo-Feng Hsu, Jiarong Xing 等NeurIPS 2022 · 被引用 9 次
- Width-based Lookaheads with Learnt Base Policies and Heuristics Over the Atari-2600 BenchmarkStefan O'Toole, Nir Lipovetzky, Miquel Ramírez, Adrian R. PearceNeurIPS 2021
它引用的顶会 Paper3
- NVAE: A Deep Hierarchical Variational AutoencoderArash Vahdat, Jan KautzNeurIPS 2020 · 被引用 1,141 次
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski 等ICLR 2020 · 被引用 969 次
- Optimal Variance Control of the Score-Function Gradient Estimator for Importance-Weighted BoundsValentin Liévin, Andrea Dittadi, Anders Christensen, Ole WintherNeurIPS 2020 · 被引用 9 次
相关 Paper
- Discovering diverse athletic jumping strategiesZhiqi Yin, Zeshi Yang, Michiel van de Panne, KangKang YinSIGGRAPH 2021 · 被引用 46 次
- Scalable Decision-Making in Stochastic Environments through Learned Temporal AbstractionBaiting Luo, Ava Pettet, Aron Laszka, Abhishek Dubey 等ICLR 2025
- Efficient Planning in a Compact Latent Action SpaceZhengyao Jiang, Tianjun Zhang, Michael Janner, Yueying Li 等ICLR 2023 · 被引用 3 次
- In Pursuit of Pixel Supervision for Visual Pre-trainingLihe Yang, Shang-Wen Li, Yang Li, Xinjie Lei 等CVPR 2026 · 被引用 13 次
- General Policies, Representations, and Planning WidthBlai Bonet, Hector GeffnerAAAI 2021 · 被引用 28 次
