Master'sOpen Access

Microarray data analysis with ensemble learning methods

2022
0 views
0 downloads
Advisor: Doç. Dr. Çiğdem Erol

Abstract (EN)

The Ensemble learning method is a machine learning technique that consists of combining many machine learning algorithms in order to have a better classification or prediction result than can be obtained from a single algorithm. Recently, ensemble learning has been shown as an intensive and widespread machine learning technique that is mostly used in many areas for different purposes such as pattern recognition, feature selection, and classification. Bagging, Boosting, Stacking, Random Forest are the most popular and widely used ensemble learning methods. One of the advantages of ensemble learning is the ability to deal with high- dimensional and complex data structures, that is to say ensemble learning helps reduce the variance of datasets and provides more accurate results. Microarray analysis is a bioinformatics method that aims to provide results that can help early detection and classification of diseases, to find genes or biomarkers related to diseases, and to drugs. Therefore, the data obtained by microarray analysis should be analyzed very carefully. In microarray data analysis, machine learning algorithms are used to predict disease states, to find disease-related genes and also to recognize patterns. However, one of the common major problems faced by these algorithms is the small sample size and the high dimensionality of the microarray datasets, often characterized by the presence of a high variance leading to the overfitting of models, and that high variance problem aforementioned has been highlighted in many studies. The aim of this thesis is to try to reduce the variance in microarray dataset by using ensemble learning methods. For this purpose, non-small cell lung cancer dataset (GSE19804) in the NCBI GEO database was used. Within the framework of this study, two types of machine learning models were applied; one is based on stacking ensemble learning and the other is based on simple algorithm. Seven of the 12 algorithms selected at the beginning of the analysis, namely Radial Support Vector Machine, Linear Discriminant Analysis, k-Nearest Neighbors, C5.0, CART, Feature Extraction Neural Networks, Generalized Linear Model algorithms were used in the analysis. As a result, stacking ensemble learning models showed lower variance compared to simple algorithm-based models.

Author

Dr. Tchare Adnaane Bawa

How to Cite

Tchare Adnaane Bawa (Master Thesis). Microarray data analysis with ensemble learning methods, 2022, İstanbul University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from İstanbul University