Counteractive RL: Rethinking Core Principles for Efficient and Scalable Deep Reinforcement Learning
Ezgi Korkmaz
Abstract
Following the pivotal success of learning strategies to win at tasks, solely by interacting with an environment without any supervision, agents have gained the ability to make sequential decisions in complex MDPs. Yet, reinforcement learning policies face exponentially growing state spaces in high dimensional MDPs resulting in a dichotomy between computational complexity and policy success. In our paper we focus on the agent's interaction with the environment in a high-dimensional MDP during the learning phase and we introduce a theoretically-founded novel paradigm based on experiences obtained through counteractive actions. Our analysis and method provide a theoretical basis for efficient, effective, scalable and accelerated learning, and further comes with zero additional computational complexity while leading to significant acceleration in training. We conduct extensive experiments in the Arcade Learning Environment with high-dimensional state representation MDPs. The experimental results further verify our theoretical analysis, and our method achieves significant performance increase with substantial sample-efficiency in high-dimensional environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- How to Lose Inherent Counterfactuality in Reinforcement LearningEzgi KorkmazICLR 2026
- Principled Analysis of Deep Reinforcement Learning Evaluation and Design ParadigmsEzgi KorkmazAAAI 2026
Builds on9
- Discovering Reinforcement Learning AlgorithmsJunhyuk Oh, Matteo Hessel, Wojciech M. Czarnecki, Zhongwen Xu et al.NeurIPS 2020 · 154 citations
- Learning and Planning in Complex Action SpacesThomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Mohammadamin Barekatain et al.ICML 2021 · 99 citations
- Can large language models explore in-context?Akshay Krishnamurthy, Keegan Harris, Dylan J. Foster, Cyril Zhang et al.NeurIPS 2024 · 95 citations
- Bootstrapped Meta-LearningSebastian Flennerhag, Yannick Schroecker, Tom Zahavy, Hado van Hasselt et al.ICLR 2022 · 62 citations
- Combining Q-Learning and Search with Amortized Value EstimatesJessica B. Hamrick, Victor Bapst, Alvaro Sanchez-Gonzalez, Tobias Pfaff et al.ICLR 2020 · 51 citations
Related papers
- Behavior From the Void: Unsupervised Active Pre-TrainingHao Liu, Pieter AbbeelNeurIPS 2021 · 258 citations
- Achieving Sample and Computational Efficient Reinforcement Learning by Action Space Reduction via GroupingYining Li, Peizhong Ju, Ness B. ShroffICLR 2024 · 1 citation
- Rich-Observation Reinforcement Learning with Continuous Latent DynamicsYuda Song, Lili Wu, Dylan J. Foster, Akshay KrishnamurthyICML 2024 · 2 citations
- Detecting Adversarial Directions in Deep Reinforcement Learning to Make Robust DecisionsEzgi Korkmaz, Jonah Brown-CohenICML 2023 · 16 citations
- Reward-Free RL is No Harder Than Reward-Aware RL in Linear Markov Decision ProcessesAndrew J. Wagenmaker, Yifang Chen, Max Simchowitz, Simon S. Du et al.ICML 2022 · 61 citations
