Techniques and Data Structures to Extend Database Management Systems into Genomics Platforms
| Field | Value | Language |
| dc.contributor.author | Randeni Kadupitige, Sidath | |
| dc.date.accessioned | 2024-07-19T04:27:19Z | |
| dc.date.available | 2024-07-19T04:27:19Z | |
| dc.date.issued | 2024 | en |
| dc.identifier.uri | https://hdl.handle.net/2123/32821 | |
| dc.description | Includes publication | |
| dc.description.abstract | The recent coronavirus pandemic has shown that there is a great need for data velocity and collaboration between institutions and stakeholders regarding the sequencing and identification of variants in organisms. This data collaboration is hindered by the disparate nature of data schemes, metadata recording methods and pipelines between labs. This could be solved if there was an easy way to share and query data using methods and technologies that are common to most people involved in this field. This thesis aims to provide a guideline on how to adapt an off the shelf database system into a genomics platform. Leaning on the concept of data and processing co-location, we propose a list of requirements and a prototype implementation of said requirements in the scope of a next generation sequencing (NGS) pipeline. The data and the processes involved can be easily queried and invoked using a commonly known language such as SQL. Our implementation builds bio-data types and user-defined indexes to develop NGS related algorithmic logic inside a database system. We then leverage these algorithms to build a complete sequencing pipeline - from data loading to consensus sequence generation and variant identification. We also assess each stage of the pipeline to show how effective our methods are compared to existing command line tools. | en |
| dc.language.iso | en | en |
| dc.rights | Copyright All Rights Reserved | en |
| dc.subject | genomics | en |
| dc.subject | databases | en |
| dc.subject | bioinformatics | en |
| dc.subject | next generation sequencing | en |
| dc.subject | data centric computation | en |
| dc.subject | hardware optimized databases | en |
| dc.title | Techniques and Data Structures to Extend Database Management Systems into Genomics Platforms | en |
| dc.type | Thesis | |
| dc.type.thesis | Doctor of Philosophy | en |
| dc.rights.other | The author retains copyright of this thesis. It may only be used for the purposes of research and study. It must not be used for any other purposes and may not be transmitted or shared with others without prior permission. | en |
| usyd.faculty | SeS faculties schools::Faculty of Engineering::School of Civil Engineering | en |
| usyd.degree | Doctor of Philosophy Ph.D. | en |
| usyd.awardinginst | The University of Sydney | en |
| usyd.advisor | Roehm, Uwe | en |
| usyd.include.pub | Yes | en |
Associated file/s
Associated collections