Keyframing the Future: Keyframe Discovery for Visual Prediction and Planning

Karl Pertsch; Oleh Rybkin; Jingyun Yang; Shenghao Zhou; Konstantinos Derpanis; Kostas Daniilidis; Joseph Lim; Andrew Jaegle

2020 L4DC L4DC 2020

Keyframing the Future: Keyframe Discovery for Visual Prediction and Planning

Abstract

To flexibly and efficiently reason about dynamics of temporal sequences, abstract representations that compactly represent the important information in the sequence are needed. One way of constructing such representations is by focusing on the important events in a sequence. In this paper, we propose a model that learns both to discover such key events (or keyframes) as well as to represent the sequence in terms of them. We do so using a hierarchical Keyframe-Inpainter (KeyIn) model that first generates keyframes and their temporal placement and then inpaints the sequences between keyframes. We propose a fully differentiable formulation for efficiently learning the keyframe placement. We show that KeyIn finds informative keyframes in several datasets with diverse dynamics. When evaluated on a planning task, KeyIn outperforms other recent proposals for learning hierarchical representations.

🚀 Conference Pioneer — L4DC 2020

🌉 Interdisciplinary Bridge — Artificial Intelligence and Computer Vision and Machine Learning

📈 Trend Setter — Video Generation

🧭 Keyword Pioneer — sequence representation

🐝 Cross-Pollinator — Artificial Intelligence, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Machine Learning, Mathematics & Optimization, Natural Language Processing, Speech & Audio