Value Iteration Networks

Aviv Tamar; Yi Wu; Garrett Thomas; Sergey Levine; Pieter Abbeel

2016 NIPS NeurIPS 2016

Value Iteration Networks

Abstract

We introduce the value iteration network (VIN): a fully differentiable neural network with a `planning module' embedded within. VINs can learn to plan, and are suitable for predicting outcomes that involve planning-based reasoning, such as policies for reinforcement learning. Key to our approach is a novel differentiable approximation of the value-iteration algorithm, which can be represented as a convolutional neural network, and trained end-to-end using standard backpropagation. We evaluate VIN based policies on discrete and continuous path-planning domains, and on a natural-language based search task. We show that by learning an explicit planning computation, VIN policies generalize better to new, unseen domains.

🧭 Keyword Pioneer — differentiable planning

🐣 Hot Topic Early Bird — value iteration

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Robotics, Security & Privacy, Speech & Audio

Authors

Aviv Tamar , Yi Wu , Garrett Thomas , Sergey Levine , Pieter Abbeel

Topics

Reinforcement Learning > Methods > Deep RL Reinforcement Learning > Applications > Value Iteration

Keywords

value iteration convolutional neural network differentiable planning planning module

Download PDF

Related papers

Bayesian Intermittent Demand Forecasting for Large Inventories 2016

Dynamic Network Surgery for Efficient DNNs 2016

Beyond Exchangeability: The Chinese Voting Process 2016

Safe and Efficient Off-Policy Reinforcement Learning 2016

Tagger: Deep Unsupervised Perceptual Grouping 2016