Model-Based RL; Partial Observability; Asymmetric RL; Representation Learning; World Model
Abstract :
[en] Much like humans benefit from guidance while learning, reinforcement learning algorithms may benefit from additional supervision beyond rewards. Leveraging additional information during training to learn better representations and behaviors has been the focus of asymmetric reinforcement learning. This learning paradigm has proven effective under partial observability when additional state information is available, but also under full observability when more refined state information is available. Focusing on model-based reinforcement learning, we study the effect of asymmetric learning on observation representations and on privileged information representations. First, we identify a limitation in the privileged information representations learned by an asymmetric model-based algorithm known as the Informed Dreamer. Then, we propose a novel asymmetric representation learning objective using latent guidance, resulting in a new algorithm called the Reinformed Dreamer. Experiments across several benchmarks show a more consistent improvement over Dreamer than previous asymmetric approaches.
Disciplines :
Computer science
Author, co-author :
Lambrechts, Gaspard ; McGill University > Department of Electrical and Computer Engineering ; Mila - Québec Artificial Intelligence Institute
Bolland, Adrien ; Université de Liège - ULiège > Montefiore Institute of Electrical Engineering and Computer Science
Ebi, Daniel; KIT - Karlsruhe Institute of Technology > Department of Computer Science
Ernst, Damien ; Université de Liège - ULiège > Montefiore Institute of Electrical Engineering and Computer Science
Language :
English
Title :
Reinformed Dreamer: An Asymmetric World Model Efficiently Trained through Latent Guidance