Show simple item record

FieldValueLanguage
dc.contributor.authorDing, Yixiong
dc.date.accessioned2024-03-27T03:24:31Z
dc.date.available2024-03-27T03:24:31Z
dc.date.issued2024en
dc.identifier.urihttps://hdl.handle.net/2123/32413
dc.description.abstractPublic health educational resources developed by health institutions aim for high accessibility of information. The translation of these resources, which provide the public with a basic understanding of health risks and diseases, is commonly conducted by professional translators to cater to diverse linguistic and cultural backgrounds. In recent years, the global advancement of information technology has broadened the use of Machine Translation (MT) in online health education and promotion. MT tools such as Google Translate, DeepL, and ChatGPT have significantly improved performance, yet they face challenges posed by the language complexity, content complexity, and formality of professional medical resources. In this study, we leverage Natural Language Processing (NLP) and Machine Learning (ML) tools to harness the power of text classification, a vital task that assigns text to one or more predefined categories. We aim to develop machine learning classifiers within our newly proposed Multi-Dimensional Text Classification Pipeline (MD-TCP) framework. As a risk-prevention mechanism, MD-TCP assists medical professionals with limited knowledge of the patient's language and helps patients who wish to self-navigate. Our model predicts the likelihood of clinical mistakes or incomprehensible machine translation outputs based on the features of English source input to the machine translation systems. MD-TCP is a new, comprehensive pipeline for data mining and feature extraction that we developed to achieve this goal. The pipeline has demonstrated significant improvements in both of our datasets. Regarding Accuracy, AUC, Sensitivity, Precision, and Specificity, our method improved by 24% - 33% compared to baseline methods. This underscores the potential of machine learning, mainly when implemented through MD-TCP, in predicting translation errors, thereby ensuring more accurate and understandable translations for health resources across diverse populations.en
dc.language.isoenen
dc.rightsCopyright All Rights Reserveden
dc.subjectNatural Language Processingen
dc.subjectMachine Learningen
dc.subjectData Miningen
dc.subjectText Features Extractionen
dc.subjectText Classificationen
dc.subjectPublic Health Education and Promotionen
dc.titleAnticipating Hazards in Machine Translations of Public Health Resources via Advanced Text Classification Pipelinesen
dc.typeThesis
dc.type.thesisMasters by Researchen
dc.rights.otherThe author retains copyright of this thesis. It may only be used for the purposes of research and study. It must not be used for any other purposes and may not be transmitted or shared with others without prior permission.en
usyd.facultySeS faculties schools::Faculty of Engineering::School of Civil Engineeringen
usyd.degreeMaster of Philosophy M.Philen
usyd.awardinginstThe University of Sydneyen
usyd.advisorXU, Changen


Show simple item record

Associated file/s

Associated collections

Show simple item record

There are no previous versions of the item available.