Causal Interpretation of Neural Networks
| Field | Value | Language |
| dc.contributor.author | Guo, Senhui | |
| dc.date.accessioned | 2023-09-25T03:19:46Z | |
| dc.date.available | 2023-09-25T03:19:46Z | |
| dc.date.issued | 2023 | en |
| dc.identifier.uri | https://hdl.handle.net/2123/31702 | |
| dc.description.abstract | In recent years, neural networks (NNs) have achieved great success in various challenging tasks. However, due to their black-box nature, it is still challenging to interpret how they process the data. As a result, NNs misled by spurious correlations often fail on out-of- distribution (o.o.d.) samples. Causal theory, on the other hand, has shown great potential in battling spurious correlations and interpreting complex systems. In this work, we use causal theory to tackle two tasks: NN modularity and causal representation learning. First, inspired by structural causal models (SCM), we propose a structural mechanism identification framework, namely, competitive disentanglement, to discover modular structures in NNs. The idea is to let neurons compete with each other in a specific setting such that neurons belonging to different modules will exhibit different behaviors. By observing these behaviors of neurons, we are able to disentangle them into different modules. Second, we explore the formulation of causal representation learning. The probability of causation (POC) has been regarded as a good measure of causal relevance. However, due to its counterfactual form, people have not been able to estimate the POC of representations consistently. We propose to model the representation learning process with a disentangled Variational Autoencoder (VAE) capable of producing counterfactual data that are rare in the training dataset. Thanks to the generative model in VAE, we are able to estimate the counterfactual POC directly and use them as part of the optimization goal to learn features with good causal properties. | en |
| dc.language.iso | en | en |
| dc.rights | Copyright All Rights Reserved | en |
| dc.subject | causal | en |
| dc.subject | causal inference | en |
| dc.subject | causal representation | en |
| dc.subject | interpretability | en |
| dc.title | Causal Interpretation of Neural Networks | en |
| dc.type | Thesis | |
| dc.type.thesis | Doctor of Philosophy | en |
| dc.rights.other | The author retains copyright of this thesis. It may only be used for the purposes of research and study. It must not be used for any other purposes and may not be transmitted or shared with others without prior permission. | en |
| usyd.faculty | SeS faculties schools::Faculty of Engineering::School of Electrical and Information Engineering | en |
| usyd.degree | Doctor of Philosophy Ph.D. | en |
| usyd.awardinginst | The University of Sydney | en |
| usyd.advisor | Ouyang, Wanli | en |
Associated file/s
Associated collections