MIMEx: Intrinsic Rewards from Masked Input Modeling

Toru Lin; Allan Jabri

2023 NIPS NeurIPS 2023

MIMEx: Intrinsic Rewards from Masked Input Modeling

Abstract

Exploring in environments with high-dimensional observations is hard. One promising approach for exploration is to use intrinsic rewards, which often boils down to estimating "novelty" of states, transitions, or trajectories with deep networks. Prior works have shown that conditional prediction objectives such as masked autoencoding can be seen as stochastic estimation of pseudo-likelihood. We show how this perspective naturally leads to a unified view on existing intrinsic reward approaches: they are special cases of conditional prediction, where the estimation of novelty can be seen as pseudo-likelihood estimation with different mask distributions. From this view, we propose a general framework for deriving intrinsic rewards -- Masked Input Modeling for Exploration (MIMEx) -- where the mask distribution can be flexibly tuned to control the difficulty of the underlying conditional prediction task. We demonstrate that MIMEx can achieve superior results when compared against competitive baselines on a suite of challenging sparse-reward visuomotor tasks.

🌉 Interdisciplinary Bridge — Deep Learning and Machine Learning and Reinforcement Learning and Robotics

🧭 Keyword Pioneer — exploration in reinforcement learning

🐝 Cross-Pollinator — Artificial Intelligence, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Knowledge & Reasoning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Robotics

Authors

Toru Lin , Allan Jabri

Topics

Machine Learning > Learning Types > Self-Supervised Learning Machine Learning > Application Areas > Data Augmentation Reinforcement Learning > Methods > Deep RL Reinforcement Learning > Applications > Robotics Deep Learning > Learning Types > Self-Supervised Learning Machine Learning > Learning Types > Exploration Robotics > Applications > Robotics

Keywords

sparse reward intrinsic reward pseudo-likelihood estimation masked autoencoding conditional prediction exploration in reinforcement learning

Download PDF

Related papers

Risk-Averse Model Uncertainty for Distributionally Robust Safe Reinforcement Learning 2023

Generative Modeling through the Semi-dual Formulation of Unbalanced Optimal Transport 2023

Self-Supervised Motion Magnification by Backpropagating Through Optical Flow 2023

Diffused Task-Agnostic Milestone Planner 2023

Characterizing Graph Datasets for Node Classification: Homophily-Heterophily Dichotomy and Beyond 2023