A Minimalist Approach to Offline Reinforcement Learning

Scott Fujimoto; Shixiang (Shane) Gu

2021 NIPS NeurIPS 2021

A Minimalist Approach to Offline Reinforcement Learning

Abstract

Offline reinforcement learning (RL) defines the task of learning from a fixed batch of data. Due to errors in value estimation from out-of-distribution actions, most offline RL algorithms take the approach of constraining or regularizing the policy with the actions contained in the dataset. Built on pre-existing RL algorithms, modifications to make an RL algorithm work offline comes at the cost of additional complexity. Offline RL algorithms introduce new hyperparameters and often leverage secondary components such as generative models, while adjusting the underlying RL algorithm. In this paper we aim to make a deep RL algorithm work while making minimal changes. We find that we can match the performance of state-of-the-art offline RL algorithms by simply adding a behavior cloning term to the policy update of an online RL algorithm and normalizing the data. The resulting algorithm is a simple to implement and tune baseline, while more than halving the overall run time by removing the additional computational overheads of previous methods.

🌉 Interdisciplinary Bridge — Deep Learning and Machine Learning and Reinforcement Learning

🐣 Hot Topic Early Bird — behavior cloning

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Robotics, Security & Privacy

Authors

Scott Fujimoto , Shixiang (Shane) Gu

Topics

Reinforcement Learning > Methods > Offline RL Machine Learning > Learning Types > Imitation Learning Deep Learning > Learning Types > Reinforcement Learning Machine Learning > Learning Types > Offline Reinforcement Learning

Keywords

deep reinforcement learning offline reinforcement learning behavior cloning value estimation policy regularization deep rl

Download PDF

Related papers

Mosaicking to Distill: Knowledge Distillation from Out-of-Domain Data 2021

On Model Calibration for Long-Tailed Object Detection and Instance Segmentation 2021

Test-Time Personalization with a Transformer for Human Pose Estimation 2021

NTopo: Mesh-free Topology Optimization using Implicit Neural Representations 2021

Scalable Intervention Target Estimation in Linear Models 2021