Multimodal Integration for Natural Language Classification and Generation
| Field | Value | Language |
| dc.contributor.author | Zhang, Zhihao | |
| dc.date.accessioned | 2024-01-08T02:24:50Z | |
| dc.date.available | 2024-01-08T02:24:50Z | |
| dc.date.issued | 2023 | en |
| dc.identifier.uri | https://hdl.handle.net/2123/32061 | |
| dc.description | Includes publication | |
| dc.description.abstract | Multimodal integration is a framework for building models that can accept information from different types of modalities. Due to the recent success in the Transformer model and Pre-training Fine-tuning Techniques, Vision-and-Language Pre-training Models have been heavily investigated and they achieved State-of-the-Art in various of Vision-and-Language downstream tasks, such as Visual Question Answering, Image Text Matching and Image Captioning. However, most of the previous studies focus on improving the performance of the models and only provide accessible code for research purposes. There are several existing open-source libraries such as Natural Language Toolkit, OpenCV and HuggingFace, which combine and standardise the available models and tools for easy access, but applying these libraries still requires expertise in both Deep Learning and programming. Moreover, there has been no recent research aimed at establishing user-friendly multimodal question-answering platforms for non-deep-learning users. Therefore, the question of how State-Of-The-Art multimodal models can be easily applied by professionals in other domains remains open. Apart from the first challenge, there exists another challenge in the less-common domain. Since general multimodal domains such as street view, landscape, and indoor scenes have been extensively studied with current VL-PMs, while specific domains like medicine, geography, and esports have garnered less attention. Due to the difficulties in data collection, there aren't many publicly available multimodal datasets, and those that exist tend to be small. This scarcity poses challenges for model training. Consequently, the question of how to collect a comprehensive multimodal dataset in the esports domain and how to improve domain-specific multimodal models remains open. Therefore, the main focus of this thesis is integrating multimodal information for natural language classification and generation tasks by addressing the two challenges. | en |
| dc.language.iso | en | en |
| dc.rights | Copyright All Rights Reserved | en |
| dc.subject | Multimodal integration | en |
| dc.subject | Natural Language Processing | en |
| dc.title | Multimodal Integration for Natural Language Classification and Generation | en |
| dc.type | Thesis | |
| dc.type.thesis | Masters by Research | en |
| dc.rights.other | The author retains copyright of this thesis. It may only be used for the purposes of research and study. It must not be used for any other purposes and may not be transmitted or shared with others without prior permission. | en |
| usyd.faculty | SeS faculties schools::Faculty of Engineering::School of Civil Engineering | en |
| usyd.degree | Master of Philosophy M.Phil | en |
| usyd.awardinginst | The University of Sydney | en |
| usyd.advisor | Poon, Josiah | en |
| usyd.include.pub | Yes | en |
Associated file/s
Associated collections