Data integration for predicting gene expression
2019
0 views
0 downloads
Advisor: Prof. Dr. Hasan Oğul
Abstract (EN)
Protein synthesis is the basis of the sustainability of the living form. Small nucleotide sequences (micro-RNA) and other executive genes (Transcription Factor, TF) that regulate coding genes play an important role in the protein synthesis. The aim of this study was to investigate the effect of regulation information of micro-RNA and TFs on the performance of predicting the exact value of expressions of protein coding genes. In order to predict the exact value of gene expression, systematic approaches that includes regression-based models are introduced. First, linear, k-NN and Relational Vector Machine (RVM) regression models were applied to solve the common problem of missing data in gene expression measurements. The expression vectors used in the training phase of the regression model are generally composed of the expression values of the same gene that belongs to different experiments. After that, the effect of the inclusion of different gene expression values of the same experiment on these expression vectors was investigated. For this, the one-way data matrix, consisting of gene expression values, was transformed into a two-way data matrix using Two-way Collaborative Filtering method and the regression model was built with this new data matrix. It is observed that this new feature representation technique that is first used in this study for gene expression predicting increases the performance of predicting. In addition, the effect of integrating gene expression values of different cancer types on gene expression predicting is also investigated. Here, it is observed that the use of colon cancer data in model learning to predict the gene expression of prostate cancer increases prediction performance. There are many studies in the literature to determine the relationship between regulating molecules and genes using gene expression values. However, there are very limited studies based on predicting the exact value of gene expression by using these relations in the cell. Finally, miRNA-gene and TF-gene interaction information and gene expression values were integrated and the prediction performance outcomes obtained by using linear and RVM regression models were discussed. Euclidean, Affine Transformation and Bhattacharya distance measures were used in data integration approaches. Gene expression matrices from Gene Expression Omnibus; TF-gene regulation information from TRANSFAC; miRNA-gene regulation information from mirDB, mirTarbase and mirConnX were used. Spearman similarity coefficient, Pearson similarity coefficient and Root Mean Squared Error (RMSE) were used to evaluate the performance of predicting. It is observed that the performance of predicting gene expression is increased by integrating of miRNA-gene regulation information.
Author
Tuncay Bayrak
How to Cite
Tuncay Bayrak (Doctorate thesis). Data integration for predicting gene expression, 2019, Başkent University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Başkent University
- Classification of aircraft images(2025)
- A nietzschean reading of cormac Mccarthy's Blood Meridian Or the Evening Redness in the west and The Road(2021)
- An analysis of the alignment of English textbooks in Turkish primary schools with the 21st century skills(2025)
- The gastronomic heritage of tradesmen's restaurants: The case of Ankara(2025)
- The impact of vocational education on the skilled labor shortage: A study on the construction sector in Ankara province(2025)
- The effects of bankruptcy on litigation and follow-up processes(2019)
