Lune

ICLR2023Top-tier venue

Provably Efficient Lifelong Reinforcement Learning with Linear Representation

Sanae Amani, Lin Yang, Ching-An Cheng

2023Year

Abstract

We theoretically study lifelong reinforcement learning (RL) with linear representation in a regret minimization setting. The goal of the agent is to learn a multi-task policy based on a linear representation while solving a sequence of tasks that may be adaptively chosen based on the agent's past behaviors. We frame the problem as a linearly parameterized contextual Markov decision process (MDP), where each task is specified by a context and the transition dynamics is context-independent, and we introduce a new completeness-style assumption on the representation which is sufficient to ensure the optimal multi-task policy is realizable under the linear representation. Under this assumption, we propose an algorithm, called UCB Lifelong Value Distillation (UCBlvd), that provably achieves sublinear regret for any sequence of tasks while using only sublinear planning calls. Specifically, for KK task episodes of horizon HH, our algorithm has a regret bound O~((d3+d′d)H4K)\tilde{\mathcal{O}}(\sqrt{(d^3+d^\prime d)H^4K}) using O(dHlog⁡(K))\mathcal{O}(dH\log(K)) number of planning calls, where dd and d′d^\prime are the feature dimensions of the dynamics and rewards, respectively. This theoretical guarantee implies that our algorithm can enable a lifelong learning agent to learn to internalize experiences into a multi-task policy and rapidly solve new tasks.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines