Optimal Transport in Summarisation: Towards Unsupervised Multimodal Summarisation
Access status:
Open Access
Type
ThesisThesis type
Doctor of PhilosophyAuthor/s
Tang, Pik YeeAbstract
Summarisation aims to condense a given piece of information into a short and succinct summary that best covers its semantics with the least redundancy. With the explosion of multimedia data, multimodal summarisation with multimodal output emerges and extends the inquisitiveness of ...
See moreSummarisation aims to condense a given piece of information into a short and succinct summary that best covers its semantics with the least redundancy. With the explosion of multimedia data, multimodal summarisation with multimodal output emerges and extends the inquisitiveness of the task. Summarising a video-document pair into a visual-textual summary helps users obtain a more informative and visual understanding. Although various methods have achieved promising performance, they have limitations, including expensive training, lack of interpretability, and insufficient brevity. Therefore, this thesis addresses the gap and examines the application of optimal transport (OT) in unsupervised summarisation. The major contributions are as follows: 1) An interpretable OT-based method is proposed for text summarisation. It formulates summary sentence extraction as minimising the transportation cost of their semantic distributions; 2) An efficient and interpretable unsupervised reinforcement learning method is proposed for text summarisation. Multihead attentional pointer-based networks learn the representation and extract salient sentences and words. The learning strategy mimics human judgment by optimising summary quality regarding OT-based semantic coverage and fluency; 3) A new task, eXtreme Multimodal Summarisation with Multiple Output (XMSMO) is introduced. It summarises a video-document pair into an extremely short multimodal summary. An unsupervised Hierarchical Optimal Transport Network learns and uses OT solvers to maximise multimodal semantic coverage. A new large-scale dataset is constructed to facilitate future research; 4) A Topic-Guided Co-Attention Transformer method is proposed for XMSMO. It constructs a two-stage uni- and cross-modal modelling with topic guidance. An OT-guided unsupervised training strategy optimises the similarity between semantic distributions of topics. Comprehensive experiments demonstrate the effectiveness of the proposed methods.
See less
See moreSummarisation aims to condense a given piece of information into a short and succinct summary that best covers its semantics with the least redundancy. With the explosion of multimedia data, multimodal summarisation with multimodal output emerges and extends the inquisitiveness of the task. Summarising a video-document pair into a visual-textual summary helps users obtain a more informative and visual understanding. Although various methods have achieved promising performance, they have limitations, including expensive training, lack of interpretability, and insufficient brevity. Therefore, this thesis addresses the gap and examines the application of optimal transport (OT) in unsupervised summarisation. The major contributions are as follows: 1) An interpretable OT-based method is proposed for text summarisation. It formulates summary sentence extraction as minimising the transportation cost of their semantic distributions; 2) An efficient and interpretable unsupervised reinforcement learning method is proposed for text summarisation. Multihead attentional pointer-based networks learn the representation and extract salient sentences and words. The learning strategy mimics human judgment by optimising summary quality regarding OT-based semantic coverage and fluency; 3) A new task, eXtreme Multimodal Summarisation with Multiple Output (XMSMO) is introduced. It summarises a video-document pair into an extremely short multimodal summary. An unsupervised Hierarchical Optimal Transport Network learns and uses OT solvers to maximise multimodal semantic coverage. A new large-scale dataset is constructed to facilitate future research; 4) A Topic-Guided Co-Attention Transformer method is proposed for XMSMO. It constructs a two-stage uni- and cross-modal modelling with topic guidance. An OT-guided unsupervised training strategy optimises the similarity between semantic distributions of topics. Comprehensive experiments demonstrate the effectiveness of the proposed methods.
See less
Date
2023Licence
Copyright All Rights ReservedRights statement
The author retains copyright of this thesis. It may only be used for the purposes of research and study. It must not be used for any other purposes and may not be transmitted or shared with others without prior permission.Faculty/School
Faculty of Engineering, School of Civil EngineeringAwarding institution
The University of SydneyShare