Vision Transformer Advanced by Exploring Intrinsic Inductive Bias
| Field | Value | Language |
| dc.contributor.author | Xu, Yufei | |
| dc.date.accessioned | 2024-01-07T23:35:39Z | |
| dc.date.available | 2024-01-07T23:35:39Z | |
| dc.date.issued | 2023 | en |
| dc.identifier.uri | https://hdl.handle.net/2123/32048 | |
| dc.description | Includes publication | |
| dc.description.abstract | The vision models have experienced a paradigm shift from convolutional neural networks (CNNs) to transformers. Compared with convolutions, transformers can capture both short- and long-range dependencies, making them more adaptable for extensive datasets. However, this adaptability comes at a cost: vision transformers are data-hungry and prone to overfitting with limited training data, restricting their applications in various vision tasks. This thesis aims to mitigate these shortcomings through advancements in architectural design and training methodologies, encompassing a comprehensive assessment involving various vision tasks. We investigate the data-hungry nature of transformers due to their lack of inductive bias. Our proposed remedy involves the incorporation of convolution blocks with multi-head self-attention (MHSA) mechanisms within each transformer block. This integration injects the inductive bias into the architecture, formulating the ViTAE model. Moreover, we present an innovative self-supervised learning approach, RegionCL, which bolsters the training process by emphasizing local information via region swapping. What’s more, a ViTPose-G model, based on ViTAE-G, is introduced and demonstrates exceptional performance in pose estimation tasks across various datasets. | en |
| dc.language.iso | en | en |
| dc.rights | Copyright All Rights Reserved | en |
| dc.subject | Transformer | en |
| dc.subject | Inductive Bias | en |
| dc.subject | Self-supervised Learning | en |
| dc.subject | Pose Estimation | en |
| dc.subject | Multi-task Learning | en |
| dc.subject | Convolutional neural networks | en |
| dc.title | Vision Transformer Advanced by Exploring Intrinsic Inductive Bias | en |
| dc.type | Thesis | |
| dc.type.thesis | Doctor of Philosophy | en |
| dc.rights.other | The author retains copyright of this thesis. It may only be used for the purposes of research and study. It must not be used for any other purposes and may not be transmitted or shared with others without prior permission. | en |
| usyd.faculty | SeS faculties schools::Faculty of Engineering::School of Civil Engineering | en |
| usyd.degree | Doctor of Philosophy Ph.D. | en |
| usyd.awardinginst | The University of Sydney | en |
| usyd.advisor | Tao, Dacheng | en |
| usyd.include.pub | Yes | en |
Associated file/s
Associated collections