Exploring Next-Generation Visual Object Detection and Tracking System
Access status:
Open Access
Type
ThesisThesis type
Doctor of PhilosophyAuthor/s
Feng, WeitaoAbstract
This thesis embarks on an exploration into the evolution of the next generation of visual detection and tracking systems, delineating four crucial subtopics: 1) model architecture: Investigating the impact of critical distractors and devising a model architecture explicitly tailored ...
See moreThis thesis embarks on an exploration into the evolution of the next generation of visual detection and tracking systems, delineating four crucial subtopics: 1) model architecture: Investigating the impact of critical distractors and devising a model architecture explicitly tailored to effectively manage them, 2) learning procedure: Delving into the intricacies of relation learning within joint tasks and endeavoring to elucidate fundamental principles governing the learning process, 3) scenario abilities: Examining the frame rate robustness of the joint task and formulating a novel pipeline to adeptly handle dynamic frame rates, and 4) data: Exploring unsupervised learning methodologies tailored for the joint task. These carefully chosen subtopics align with and address four critical challenges, namely the distractor problem, learning representations from videos, robust tracking, and the application of unsupervised learning in the context of videos. This research aims to contribute significant insights and advancements to the landscape of joint detection and tracking systems in the dynamic realm of deep learning. To address these challenges, several new methods have been raised in this thesis. For the distractor problem, the concept of ‘switcher’ has been proposed and a switcher-aware association has been introduced to utilize the information of distractors. For the relation learning problem, a similarity- and quality-guided attention mechanism has been introduced for more effective feature refinement. For the frame rate robustness problem, a frame rate agnostic framework has been proposed to conduct joint detection and tracking with unseen frame rate inputs. For the unsupervised learning problem, a video-based pseudo- labeling method has been developed for model pretraining. To evaluate the effectiveness of these approaches, both qualitative and quantitative experiments have been conducted. The experimental results have shown satisfactory progress.
See less
See moreThis thesis embarks on an exploration into the evolution of the next generation of visual detection and tracking systems, delineating four crucial subtopics: 1) model architecture: Investigating the impact of critical distractors and devising a model architecture explicitly tailored to effectively manage them, 2) learning procedure: Delving into the intricacies of relation learning within joint tasks and endeavoring to elucidate fundamental principles governing the learning process, 3) scenario abilities: Examining the frame rate robustness of the joint task and formulating a novel pipeline to adeptly handle dynamic frame rates, and 4) data: Exploring unsupervised learning methodologies tailored for the joint task. These carefully chosen subtopics align with and address four critical challenges, namely the distractor problem, learning representations from videos, robust tracking, and the application of unsupervised learning in the context of videos. This research aims to contribute significant insights and advancements to the landscape of joint detection and tracking systems in the dynamic realm of deep learning. To address these challenges, several new methods have been raised in this thesis. For the distractor problem, the concept of ‘switcher’ has been proposed and a switcher-aware association has been introduced to utilize the information of distractors. For the relation learning problem, a similarity- and quality-guided attention mechanism has been introduced for more effective feature refinement. For the frame rate robustness problem, a frame rate agnostic framework has been proposed to conduct joint detection and tracking with unseen frame rate inputs. For the unsupervised learning problem, a video-based pseudo- labeling method has been developed for model pretraining. To evaluate the effectiveness of these approaches, both qualitative and quantitative experiments have been conducted. The experimental results have shown satisfactory progress.
See less
Date
2024Licence
Copyright All Rights ReservedRights statement
The author retains copyright of this thesis. It may only be used for the purposes of research and study. It must not be used for any other purposes and may not be transmitted or shared with others without prior permission.Faculty/School
Faculty of Engineering, School of Electrical and Information EngineeringAwarding institution
The University of SydneyShare