Learning with Noisy Labels Incorporating Fairness and Privacy Concerns
Access status:
Open Access
Type
ThesisThesis type
Doctor of PhilosophyAuthor/s
Wu, SonghuaAbstract
In the era of big data, the volume of data is growing at a tremendous rate as data has been generated from multifarious sources, but without reasonable supervision. Consequently, the generated data are primarily imperfect, such as inaccurate, biased, and even invading privacy. ...
See moreIn the era of big data, the volume of data is growing at a tremendous rate as data has been generated from multifarious sources, but without reasonable supervision. Consequently, the generated data are primarily imperfect, such as inaccurate, biased, and even invading privacy. However, modern machine learning systems heavily rely on the quality of data. Noisy labels (inaccurate data) can be easily memorized by deep neural networks and thereby leads to overfitting and poor generalization problems; biased data guide models to give unfair predictions toward certain groups; privacy-invasion data are inherently harmful to the machine learning community. To deal with noisy labels, we propose two solutions to this problem from diverse perspectives. The first one from a noise reduction perspective transforms data points with noisy class labels to data pairs with noisy similarity labels. The second one from a curriculum learning perspective designs a curriculum selecting training data based on their dynamics over the course of training to learn the clean classifier and the transition matrix simultaneously. We provide general frameworks for learning fair classifiers with noisy labels. For statistical fairness notions, we rewrite the classification risk and the fairness metric in terms of noisy data and thereby build robust classifiers. For the causality-based fairness notion, we exploit the internal causal structure of data to model the label noise and counterfactual fairness simultaneously; we propose a denoised and unbiased estimator for the classification risk with respect to the accurately labeled data by employing the noisy data with indirect supervision and then learn the optimal model under the empirical risk minimization framework.
See less
See moreIn the era of big data, the volume of data is growing at a tremendous rate as data has been generated from multifarious sources, but without reasonable supervision. Consequently, the generated data are primarily imperfect, such as inaccurate, biased, and even invading privacy. However, modern machine learning systems heavily rely on the quality of data. Noisy labels (inaccurate data) can be easily memorized by deep neural networks and thereby leads to overfitting and poor generalization problems; biased data guide models to give unfair predictions toward certain groups; privacy-invasion data are inherently harmful to the machine learning community. To deal with noisy labels, we propose two solutions to this problem from diverse perspectives. The first one from a noise reduction perspective transforms data points with noisy class labels to data pairs with noisy similarity labels. The second one from a curriculum learning perspective designs a curriculum selecting training data based on their dynamics over the course of training to learn the clean classifier and the transition matrix simultaneously. We provide general frameworks for learning fair classifiers with noisy labels. For statistical fairness notions, we rewrite the classification risk and the fairness metric in terms of noisy data and thereby build robust classifiers. For the causality-based fairness notion, we exploit the internal causal structure of data to model the label noise and counterfactual fairness simultaneously; we propose a denoised and unbiased estimator for the classification risk with respect to the accurately labeled data by employing the noisy data with indirect supervision and then learn the optimal model under the empirical risk minimization framework.
See less
Date
2023Licence
Copyright All Rights ReservedRights statement
The author retains copyright of this thesis. It may only be used for the purposes of research and study. It must not be used for any other purposes and may not be transmitted or shared with others without prior permission.Faculty/School
Faculty of Engineering, School of Civil EngineeringAwarding institution
The University of SydneyShare