VQ3D: Learning a 3D-Aware Generative Model on ImageNet

Kyle Sargent; Jing Yu Koh; Han Zhang; Huiwen Chang; Charles Herrmann; Pratul Srinivasan; Jiajun Wu; Deqing Sun

2023 ICCV ICCV 2023

VQ3D: Learning a 3D-Aware Generative Model on ImageNet

Abstract

Recent work has shown the possibility of training generative models of 3D content from 2D image collections on small datasets corresponding to a single object class, such as human faces, animal faces, or cars. However, these models struggle on larger, more complex datasets. To model diverse and unconstrained image collections such as ImageNet, we present VQ3D, which introduces a NeRF-based decoder into a two-stage vector-quantized autoencoder. Our Stage 1 allows for the reconstruction of an input image and the ability to change the camera position around the image, and our Stage 2 allows for the generation of new 3D scenes. VQ3D is capable of generating and reconstructing 3D-aware images from the 1000-class ImageNet dataset of 1.2 million training images, and achieves a competitive ImageNet generation FID score of 16.8.

🌉 Interdisciplinary Bridge — Artificial Intelligence and Deep Learning

🧭 Keyword Pioneer — vector quantized autoencoder

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Robotics, Security & Privacy, Speech & Audio

Authors

Kyle Sargent , Jing Yu Koh , Han Zhang , Huiwen Chang , Charles Herrmann , Pratul Srinivasan , Jiajun Wu , Deqing Sun

Topics

Artificial Intelligence > Learning Paradigms > Transfer Learning Deep Learning > Architectures > Autoencoders Deep Learning > Models > Generative Models

Keywords

generative model view synthesis 3d-aware image generation vector quantized autoencoder

Download PDF

Related papers

PVT++: A Simple End-to-End Latency-Aware Visual Tracking Framework 2023

Periodically Exchange Teacher-Student for Source-Free Object Detection 2023

Stable and Causal Inference for Discriminative Self-supervised Deep Visual Representations 2023

Minimal Solutions to Uncalibrated Two-view Geometry with Known Epipoles 2023

3D Neural Embedding Likelihood: Probabilistic Inverse Graphics for Robust 6D Pose Estimation 2023