Master'sOpen Access

Kanser türlerini ayırt edebilmek için biyoişaretçi tanımlaması

2019
0 views
0 downloads
Advisor: Dr. Öğr. Üyesi Zerrin Işık

Abstract (EN)

RNA-sequencing data provides measurements of mRNA (messenger RNA) levels of genes based on tissue or blood samples. The critical changes in transcriptome can be observed more accurately by using RNA-sequencing data that eventually helps to understand different behavior of the disease. In this study, different feature selection methods and machine learning algorithms were examined for accurate discrimination of cancer types by using RNA-sequencing data which was obtained from blood samples. In the analysis, six cancer types were compared with each other and healthy samples. Correlation coefficient and information gain analyses are applied as main feature selection methods. The selected genes are provided as the input of Support Vector Machine (SVM), Naïve Bayes (NB), and Random Forest (RF) machine learning algorithms, that were evaluated by applying 10-fold cross-validation. In the experimental results, machine learning algorithms achieved higher than 0.85 accuracies in the discrimination of hepatobiliary, lung, and pancreatic cancer types. When machine learning models are evaluated in terms of accuracy, RF and SVM were more successful than NB for many cases. A literature-based validation revealed that some of the genes used in classifiers might be promising biomarkers for discrimination of hepatobiliary and pancreatic cancers.

Author

Dr. Cem Buğra Alkan

How to Cite

Cem Buğra Alkan (Master Thesis). Kanser türlerini ayırt edebilmek için biyoişaretçi tanımlaması, 2019, Dokuz Eylül University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Dokuz Eylül University