Single Image Depth Prediction With Wavelet Decomposition

Michael Ramamonjisoa; Michael Firman; Jamie Watson; Vincent Lepetit; Daniyar Turmukhambetov

2021 CVPR CVPR 2021

Single Image Depth Prediction With Wavelet Decomposition

Abstract

We present a novel method for predicting accurate depths from monocular images with high efficiency. This optimal efficiency is achieved by exploiting wavelet decomposition, which is integrated in a fully differentiable encoder-decoder architecture. We demonstrate that we can reconstruct high-fidelity depth maps by predicting sparse wavelet coefficients. In contrast with previous works, we show that wavelet coefficients can be learned without direct supervision on coefficients. Instead we supervise only the final depth image that is reconstructed through the inverse wavelet transform. We additionally show that wavelet coefficients can be learned in fully self-supervised scenarios, without access to ground-truth depth. Finally, we apply our method to different state-of-the-art monocular depth estimation models, in each case giving similar or better results compared to the original model, while requiring less than half the multiply-adds in the decoder network.

🌉 Interdisciplinary Bridge — Computer Vision and Deep Learning and Machine Learning

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Robotics, Security & Privacy, Speech & Audio

Authors

Michael Ramamonjisoa , Michael Firman , Jamie Watson , Vincent Lepetit , Daniyar Turmukhambetov

Topics

Machine Learning > Learning Types > Self-Supervised Learning Deep Learning > Architectures > Autoencoders Deep Learning > Architectures > Neural Networks Computer Vision > Analysis > 3D Vision Computer Vision > Analysis > Depth Estimation Computer Vision > Processing > Depth Estimation Deep Learning > Architectures > Encoder-Decoder

Keywords

self-supervised learning depth estimation monocular depth estimation image reconstruction monocular vision encoder-decoder architecture wavelet decomposition monocular depth prediction depth reconstruction

Download PDF

Related papers

Learning To Reconstruct High Speed and High Dynamic Range Videos From Events 2021

DeFLOCNet: Deep Image Editing via Flexible Low-Level Controls 2021

Vx2Text: End-to-End Learning of Video-Based Text Generation From Multimodal Inputs 2021

Coming Down to Earth: Satellite-to-Street View Synthesis for Geo-Localization 2021

Pose-Guided Human Animation From a Single Image in the Wild 2021