Vision Transformer Advanced by Exploring Intrinsic Inductive Bias
Access status:
Open Access
Type
ThesisThesis type
Doctor of PhilosophyAuthor/s
Xu, YufeiAbstract
The vision models have experienced a paradigm shift from convolutional neural networks (CNNs) to transformers. Compared with convolutions, transformers can capture both short- and long-range dependencies, making them more adaptable for extensive datasets. However, this adaptability ...
See moreThe vision models have experienced a paradigm shift from convolutional neural networks (CNNs) to transformers. Compared with convolutions, transformers can capture both short- and long-range dependencies, making them more adaptable for extensive datasets. However, this adaptability comes at a cost: vision transformers are data-hungry and prone to overfitting with limited training data, restricting their applications in various vision tasks. This thesis aims to mitigate these shortcomings through advancements in architectural design and training methodologies, encompassing a comprehensive assessment involving various vision tasks. We investigate the data-hungry nature of transformers due to their lack of inductive bias. Our proposed remedy involves the incorporation of convolution blocks with multi-head self-attention (MHSA) mechanisms within each transformer block. This integration injects the inductive bias into the architecture, formulating the ViTAE model. Moreover, we present an innovative self-supervised learning approach, RegionCL, which bolsters the training process by emphasizing local information via region swapping. What’s more, a ViTPose-G model, based on ViTAE-G, is introduced and demonstrates exceptional performance in pose estimation tasks across various datasets.
See less
See moreThe vision models have experienced a paradigm shift from convolutional neural networks (CNNs) to transformers. Compared with convolutions, transformers can capture both short- and long-range dependencies, making them more adaptable for extensive datasets. However, this adaptability comes at a cost: vision transformers are data-hungry and prone to overfitting with limited training data, restricting their applications in various vision tasks. This thesis aims to mitigate these shortcomings through advancements in architectural design and training methodologies, encompassing a comprehensive assessment involving various vision tasks. We investigate the data-hungry nature of transformers due to their lack of inductive bias. Our proposed remedy involves the incorporation of convolution blocks with multi-head self-attention (MHSA) mechanisms within each transformer block. This integration injects the inductive bias into the architecture, formulating the ViTAE model. Moreover, we present an innovative self-supervised learning approach, RegionCL, which bolsters the training process by emphasizing local information via region swapping. What’s more, a ViTPose-G model, based on ViTAE-G, is introduced and demonstrates exceptional performance in pose estimation tasks across various datasets.
See less
Date
2023Licence
Copyright All Rights ReservedRights statement
The author retains copyright of this thesis. It may only be used for the purposes of research and study. It must not be used for any other purposes and may not be transmitted or shared with others without prior permission.Faculty/School
Faculty of Engineering, School of Civil EngineeringAwarding institution
The University of SydneyShare