DoctorateOpen Access

Data integration for predicting gene expression

2019
0 views
0 downloads
Advisor: Prof. Dr. Hasan Oğul

Abstract (EN)

Protein synthesis is the basis of the sustainability of the living form. Small nucleotide sequences (micro-RNA) and other executive genes (Transcription Factor, TF) that regulate coding genes play an important role in the protein synthesis. The aim of this study was to investigate the effect of regulation information of micro-RNA and TFs on the performance of predicting the exact value of expressions of protein coding genes. In order to predict the exact value of gene expression, systematic approaches that includes regression-based models are introduced. First, linear, k-NN and Relational Vector Machine (RVM) regression models were applied to solve the common problem of missing data in gene expression measurements. The expression vectors used in the training phase of the regression model are generally composed of the expression values of the same gene that belongs to different experiments. After that, the effect of the inclusion of different gene expression values of the same experiment on these expression vectors was investigated. For this, the one-way data matrix, consisting of gene expression values, was transformed into a two-way data matrix using Two-way Collaborative Filtering method and the regression model was built with this new data matrix. It is observed that this new feature representation technique that is first used in this study for gene expression predicting increases the performance of predicting. In addition, the effect of integrating gene expression values of different cancer types on gene expression predicting is also investigated. Here, it is observed that the use of colon cancer data in model learning to predict the gene expression of prostate cancer increases prediction performance. There are many studies in the literature to determine the relationship between regulating molecules and genes using gene expression values. However, there are very limited studies based on predicting the exact value of gene expression by using these relations in the cell. Finally, miRNA-gene and TF-gene interaction information and gene expression values were integrated and the prediction performance outcomes obtained by using linear and RVM regression models were discussed. Euclidean, Affine Transformation and Bhattacharya distance measures were used in data integration approaches. Gene expression matrices from Gene Expression Omnibus; TF-gene regulation information from TRANSFAC; miRNA-gene regulation information from mirDB, mirTarbase and mirConnX were used. Spearman similarity coefficient, Pearson similarity coefficient and Root Mean Squared Error (RMSE) were used to evaluate the performance of predicting. It is observed that the performance of predicting gene expression is increased by integrating of miRNA-gene regulation information.

Author

Tuncay Bayrak

How to Cite

Tuncay Bayrak (Doctorate thesis). Data integration for predicting gene expression, 2019, Başkent University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Başkent University