Microarray data analysis with ensemble learning methods
2022
0 views
0 downloads
Advisor: Doç. Dr. Çiğdem Erol
Abstract (EN)
The Ensemble learning method is a machine learning technique that consists of combining many machine learning algorithms in order to have a better classification or prediction result than can be obtained from a single algorithm. Recently, ensemble learning has been shown as an intensive and widespread machine learning technique that is mostly used in many areas for different purposes such as pattern recognition, feature selection, and classification. Bagging, Boosting, Stacking, Random Forest are the most popular and widely used ensemble learning methods. One of the advantages of ensemble learning is the ability to deal with high- dimensional and complex data structures, that is to say ensemble learning helps reduce the variance of datasets and provides more accurate results. Microarray analysis is a bioinformatics method that aims to provide results that can help early detection and classification of diseases, to find genes or biomarkers related to diseases, and to drugs. Therefore, the data obtained by microarray analysis should be analyzed very carefully. In microarray data analysis, machine learning algorithms are used to predict disease states, to find disease-related genes and also to recognize patterns. However, one of the common major problems faced by these algorithms is the small sample size and the high dimensionality of the microarray datasets, often characterized by the presence of a high variance leading to the overfitting of models, and that high variance problem aforementioned has been highlighted in many studies. The aim of this thesis is to try to reduce the variance in microarray dataset by using ensemble learning methods. For this purpose, non-small cell lung cancer dataset (GSE19804) in the NCBI GEO database was used. Within the framework of this study, two types of machine learning models were applied; one is based on stacking ensemble learning and the other is based on simple algorithm. Seven of the 12 algorithms selected at the beginning of the analysis, namely Radial Support Vector Machine, Linear Discriminant Analysis, k-Nearest Neighbors, C5.0, CART, Feature Extraction Neural Networks, Generalized Linear Model algorithms were used in the analysis. As a result, stacking ensemble learning models showed lower variance compared to simple algorithm-based models.
Author
Dr. Tchare Adnaane Bawa
Institution
How to Cite
Tchare Adnaane Bawa (Master Thesis). Microarray data analysis with ensemble learning methods, 2022, İstanbul University.
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from İstanbul University
- In the covid 19 pandemic of female employees at a university hospital attitudes and affecting factors in nutrition of 9 months-6 years old children(2022)
- The perception of the right-wing movements in Turkey as to the 27 May Coup: 1960-1980(2020)
- Economic and social life in the Ottoman Empire according to the 1890 year's news of La Turquie Newspaper(2022)
- Land regime in the Umayyads period(2022)
- Merkel hücreli karsinomda tanısal ve prognostik belirteçler(2022)
- Biotechnologically production of polyethylene terephthalate (PET) type plastic degrading enzyme petase in escherichia coli(2021)