DoctorateOpen Access

Development of a pipeline with next-generation sequencing data analysis on retinoblastoma disease

Is this your thesis?

This record came from a bulk archive import. If it’s yours, link it to your profile.

2020
0 views
0 downloads

Abstract (EN)

Next-generation sequencing (NGS), revolutionized genomic researches, is related to massively parallel deoxyribonucleic acid (DNA) sequencing technology. Although the cost of generating NGS data was decreased compared to initally emerging stages of this technology, its cost might still be somewhat a problem according to studied data. New strategies such as pool-seq and low-coverage data have been developed to overcome this cost problem. Despite decreasing cost, it is important to elucidate whether they are efficient in NGS studies. Within the scope of this thesis, a pipeline has been developed for pool-seq and low-coverage sequencing data obtained from tumors on retinoblastoma. Retinoblastoma is an eye malignancy in childhood that is initiated by RB1 mutation or MYCN amplification and can cause to the loss of vision of eye(s), and even sometimes life. In order to evaluate the effectiveness of the developed pipeline, obtained results on both the disease data with the required features and some other non-disease data exhibiting similar characteristics as much as possible were compared by working in conjuction with a standard counterpart. It has been observed that the developed pipeline is able to call larger number of variants and achieves to higher sensitivity and F-score values. Furthermore, results related to variants, variants called in disease-associated genes and variant types in retinoblastoma data are also presented. In order to evaluate the effectiveness of the developed pipeline more precisely, it is suggested to use cancer data with higher mutation rates and larger pools. Since the alignment step in NGS data analysis is highly time-consuming and also inherently compatible with the GPU, some versions of the alignment algorithms running on the CPU have been developed for GPU executions. BWA which is the alignment algorithm utilized within the developed pipeline and BarraCUDA which is GPU adaptation developed for running under the CUDA environment were examined on different data sets and their performance was evaluated in detail. Accordingly, it has been observed that BarraCUDA significantly reduces the computation time even on a single GPU and has a similar alignment rate to BWA as stated in all the data studied. In order to fully understand degree of the contribution of the GPU, BarraCUDA's performance on data with having a high number of reads and also the effect of using more than one GPU should be examined.

Author

Gülistan Özdemir Özdoğan

How to Cite

Gülistan Özdemir Özdoğan (Doctorate thesis). Development of a pipeline with next-generation sequencing data analysis on retinoblastoma disease, 2020, Ankara Yıldırım Beyazıt University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Ankara Yıldırım Beyazıt University