Operational cancer classification using gene expression data
Is this your thesis?
This record came from a bulk archive import. If it’s yours, link it to your profile.
Abstract (EN)
Recent improvements in computer technologies, especially significant increase in processing power of central processing units, leads to usage of non ? linear models which represents physical and abstract problems better but require more memory and time, instead of simple, linear models.This study focuses on A. Statnikov?s article about multicategory cancer classification using of microarray gene expression data and optimization suggestions [1]. Before the training of support vector machines with the gene expression data which is gathered by microarray analysis, it is intented to accelerate the training and test speed process with both linear and non ? linear reduction methods. Reduction methods which are intented to be used are both implemented by using some algorithms and new interpretation of these algorithms. After that, these methods are tested according to their complexity, resource allocation and reduction performance. Therefore, by keeping the performance and success ratios of training and testing process above an acceptable treshold, it is intented to reduce the feature size in data sets as it will also increase the overall speed of the process.The results of the test show that, Independent Component Analysis (ICA), Kernel Principle Component Analysis (KPCA), Projection Pursuit Analysis (PPA) reduction algorithms used on data set failed to give any results due to excessive amount of features in data set by either locking down or terminating itself.With the usage of other algorithms which are Principle Component Analysis (PCA), Non ? Linear Principle Component Analysis (NLPCA), Self Organizing Maps (SOM), Linear Discriminant Analysis (LDA) and Correlation Analysis (CA), it is observed that the training and testing process times of the support vector machine is reduced variably. Taking this into consideration, most of the the features of the data set which is used in this study do not have any differentiative property and therefore have low - level of effect on the training and testing of the support vector machine. On the other hand, some features may become high ? level effective when combined together and form a sub group feature sets. So, by eliminating low ? level effective features and revealing high ? effective sub group features by feature selection and feature reduction, a significant improvement in both cost and time consume can be established.
Author
Namık Barış İdil
Institution
How to Cite
Namık Barış İdil (Master Thesis). Operational cancer classification using gene expression data, 2009, Başkent University, Bilgisayar Mühendisliği Bölümü.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Başkent University
- Classification of aircraft images(2025)
- A nietzschean reading of cormac Mccarthy's Blood Meridian Or the Evening Redness in the west and The Road(2021)
- An analysis of the alignment of English textbooks in Turkish primary schools with the 21st century skills(2025)
- The gastronomic heritage of tradesmen's restaurants: The case of Ankara(2025)
- The impact of vocational education on the skilled labor shortage: A study on the construction sector in Ankara province(2025)
- Effect of using eye mask and earplugs in preventing delirium in intensive care patients(2022)
