DPSNet: Multitask Learning Using Geometry Reasoning for Scene Depth and Semantics
Junning Zhang, Qunxing Su, Bo Tang, Cheng Wang, Yining Li
IEEE Transactions on Neural Networks and Learning Systems
Abstract
Multitask joint learning technology continues gaining more attention as a paradigm shift and has shown promising performance in many applications. Depth estimation and semantic understanding from monocular images emerge as a challenging problem in computer vision. While the other joint learning frameworks establish the relationship between the semantics and depth from stereo pairs, the lack of learning camera motion renders the frameworks that fail to model the geometric structure of the image scene. We make a further step in this article by proposing a multitask learning method, namely DPSNet, which can jointly perform depth and camera pose estimation and semantic scene segmentation. Our core idea for depth and camera pose prediction is that we present the rigid semantic consistency loss