Show simple item record

FieldValueLanguage
dc.contributor.authorZhou, Dongzhan
dc.date.accessioned2023-03-30T04:37:09Z
dc.date.available2023-03-30T04:37:09Z
dc.date.issued2023en
dc.identifier.urihttps://hdl.handle.net/2123/31055
dc.descriptionIncludes publication
dc.description.abstractIn this thesis, we focus on three object perception tasks: neural architecture search in object recognition, object detection, and audio-visual localization. We develop novel approaches for these tasks and conduct extensive experiments to validate their effectiveness. First, we observe that the proxies in Neural Architecture Search (NAS) present different abilities to maintain rank consistency among candidates. We examine widely used reduction factors and investigate their influences on the object recognition task. Based on our observations, we discover some reliable reduced settings that enjoy high acceleration ratio and rank consistency simultaneously. These good settings can work well on existing NAS methods to further reduce search costs and achieve competitive accuracy. Second, we propose a novel pre-training paradigm for object detection, denoted as Montage pre-training, which requires only the target detection dataset and removes external data burdens. We carefully extract training samples and devise a novel input pattern to aggregate four samples in a montage style to raise pre-training efficiency. Considering the characteristics of object detectors, we propose an ERF-adaptive dense classification strategy to further benefit the subsequent detector training. Our Montage pre-training only consumes 1/4 computation resources compared with the standard pre-training counterparts but achieves on-par or even better performance. Finally, we propose a visual reasoning module to explicitly exploit the rich visual context semantics for the self-supervised audio-visual localization task. The learning objectives are elaborately designed to provide useful guidance for the extracted visual semantics and enhance the audio-visual interactions, which leads to stronger feature representations. Experiments on three benchmark datasets show that our approach can significantly boost the localization performance.en
dc.language.isoenen
dc.rightsCopyright All Rights Reserveden
dc.subjectcomputer visionen
dc.subjectneural networken
dc.subjectobject recognitionen
dc.subjectobject detectionen
dc.subjectmulti-modal learningen
dc.titleDesigning Deep Model and Training Paradigm for Object Perceptionen
dc.typeThesis
dc.type.thesisDoctor of Philosophyen
dc.rights.otherThe author retains copyright of this thesis. It may only be used for the purposes of research and study. It must not be used for any other purposes and may not be transmitted or shared with others without prior permission.en
usyd.facultySeS faculties schools::Faculty of Engineering::School of Electrical and Information Engineeringen
usyd.degreeDoctor of Philosophy Ph.D.en
usyd.awardinginstThe University of Sydneyen
usyd.advisorOuyang, Wanlien
usyd.include.pubYesen


Show simple item record

Associated file/s

Associated collections

Show simple item record

There are no previous versions of the item available.