A Survey on Reproducibility by Evaluating Deep Reinforcement Learning Algorithms on Real-World Robots

Nicolai A. Lynnerup; Laura Nolling; Rasmus Hasle; John Hallam

2019 CORL CoRL 2019

A Survey on Reproducibility by Evaluating Deep Reinforcement Learning Algorithms on Real-World Robots

Abstract

As reinforcement learning (RL) achieves more success in solving complex tasks, more care is needed to ensure that RL research is reproducible and that algorithms therein can be compared easily and fairly with minimal bias. RL results are, however, notoriously hard to reproduce due to the algorithms’ intrinsic variance, the environments’ stochasticity, and numerous (potentially unreported) hyper-parameters. In this work we investigate the many issues leading to irreproducible research and how to manage those. We further show how to utilise a rigorous and standardised evaluation approach for easing the process of documentation, evaluation and fair comparison of different algorithms, where we emphasise the importance of choosing the right measurement metrics and conducting proper statistics on the results, for unbiased reporting of the results.

🌉 Interdisciplinary Bridge — Computer Science and Natural Language Processing and Reinforcement Learning

🧭 Keyword Pioneer — algorithm evaluation

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Data Science & Analytics, Deep Learning, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Robotics

📈 Trend Setter — Applications

Authors

Nicolai A. Lynnerup , Laura Nolling , Rasmus Hasle , John Hallam

Topics

Natural Language Processing > Applications Reinforcement Learning > Methods > Deep RL Reinforcement Learning > Applications Reinforcement Learning > Applications > Robotics Data Science & Analytics > Applications Computer Science > Applications Artificial Intelligence > Core AI > Robotics

Keywords

deep reinforcement learning algorithm evaluation measurement metrics fair comparison

Download PDF

Related papers

On-Policy Robot Imitation Learning from a Converging Supervisor 2019

Learning by Cheating 2019

Object-centric Forward Modeling for Model Predictive Control 2019

Multi-Agent Manipulation via Locomotion using Hierarchical Sim2Real 2019

Combining Deep Learning and Verification for Precise Object Instance Detection 2019