Generating Videos with Scene Dynamics

Carl Vondrick; Hamed Pirsiavash; Antonio Torralba

2016 NIPS NeurIPS 2016

Generating Videos with Scene Dynamics

Abstract

We capitalize on large amounts of unlabeled video in order to learn a model of scene dynamics for both video recognition tasks (e.g. action classification) and video generation tasks (e.g. future prediction). We propose a generative adversarial network for video with a spatio-temporal convolutional architecture that untangles the scene's foreground from the background. Experiments suggest this model can generate tiny videos up to a second at full frame rate better than simple baselines, and we show its utility at predicting plausible futures of static images. Moreover, experiments and visualizations show the model internally learns useful features for recognizing actions with minimal supervision, suggesting scene dynamics are a promising signal for representation learning. We believe generative video models can impact many applications in video understanding and simulation.

🌉 Interdisciplinary Bridge — Computer Vision and Deep Learning

📈 Trend Setter — Generative Models

🧭 Keyword Pioneer — future prediction

🐣 Hot Topic Early Bird — video generation

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Robotics, Security & Privacy, Speech & Audio

Authors

Carl Vondrick , Hamed Pirsiavash , Antonio Torralba

Topics

Deep Learning > Models > Generative Models Computer Vision > Generation > Video Generation Deep Learning > Learning Types > Generative Models

Keywords

representation learning video generation video prediction future prediction generative adversarial network scene dynamics spatio-temporal model

Download PDF

Related papers

Bayesian Intermittent Demand Forecasting for Large Inventories 2016

Dynamic Network Surgery for Efficient DNNs 2016

Beyond Exchangeability: The Chinese Voting Process 2016

Safe and Efficient Off-Policy Reinforcement Learning 2016

Tagger: Deep Unsupervised Perceptual Grouping 2016