Learning Composite Representations for Point Cloud Scene Understanding
Access status:
Open Access
Type
ThesisThesis type
Doctor of PhilosophyAuthor/s
Qiu, HaiboAbstract
In recent years, there has been significant interest in 3D point clouds from academia and industry. Unlike 2D images, point clouds consist of 3D points with Cartesian coordinates, offering an accurate, view-invariant representation of real-world scenes. Understanding 3D point cloud ...
See moreIn recent years, there has been significant interest in 3D point clouds from academia and industry. Unlike 2D images, point clouds consist of 3D points with Cartesian coordinates, offering an accurate, view-invariant representation of real-world scenes. Understanding 3D point cloud scenes is crucial for applications like autonomous driving, robotics, and AR/VR. However, comprehending these scenes is challenging due to their complex spatial structures and objects at various scales. Previous methods have often focused on learning representations that capture only local details or a specific scale, leading to suboptimal results. This thesis aims to advance learning composite representations for scene understanding by integrating multiple clues. It explores composite learning schemes from two angles: the data side and the model side. From the data perspective, it suggests projecting 3D point clouds into various 2D views and using multi-view feature fusion to learn composite representations. An end-to-end trainable geometric flow network is introduced to achieve this, enabling the learning and fusion of multi-view representations. The model side investigates the design of local basic operators and a global framework. The thesis introduces the collect-and-distribute block for local operators, which capture short- and long-range contexts simultaneously, effectively learning composite representations that incorporate sufficient contextual information. For the global framework, it addresses the challenge of multi-scale objects by employing high-resolution architectures that maintain high resolutions throughout the network and facilitate communication of multiple resolution features, efficiently learning multi-scale composite representations. Extensive experiments and analysis on popular benchmarks demonstrate the effectiveness of these approaches.
See less
See moreIn recent years, there has been significant interest in 3D point clouds from academia and industry. Unlike 2D images, point clouds consist of 3D points with Cartesian coordinates, offering an accurate, view-invariant representation of real-world scenes. Understanding 3D point cloud scenes is crucial for applications like autonomous driving, robotics, and AR/VR. However, comprehending these scenes is challenging due to their complex spatial structures and objects at various scales. Previous methods have often focused on learning representations that capture only local details or a specific scale, leading to suboptimal results. This thesis aims to advance learning composite representations for scene understanding by integrating multiple clues. It explores composite learning schemes from two angles: the data side and the model side. From the data perspective, it suggests projecting 3D point clouds into various 2D views and using multi-view feature fusion to learn composite representations. An end-to-end trainable geometric flow network is introduced to achieve this, enabling the learning and fusion of multi-view representations. The model side investigates the design of local basic operators and a global framework. The thesis introduces the collect-and-distribute block for local operators, which capture short- and long-range contexts simultaneously, effectively learning composite representations that incorporate sufficient contextual information. For the global framework, it addresses the challenge of multi-scale objects by employing high-resolution architectures that maintain high resolutions throughout the network and facilitate communication of multiple resolution features, efficiently learning multi-scale composite representations. Extensive experiments and analysis on popular benchmarks demonstrate the effectiveness of these approaches.
See less
Date
2024Licence
Copyright All Rights ReservedRights statement
The author retains copyright of this thesis. It may only be used for the purposes of research and study. It must not be used for any other purposes and may not be transmitted or shared with others without prior permission.Faculty/School
Faculty of Engineering, School of Civil EngineeringAwarding institution
The University of SydneyShare