Early diagnosis of breast cancer with machine learning classification methods
Is this your thesis?
This record came from a bulk archive import. If it’s yours, link it to your profile.
Abstract (EN)
In the medical fields, the data mining approach has been widely used in recent years in order to make more accurate prediction of disease diagnosis by avoiding unnecessary methods and to help medical practitioners make decisions more quickly. Being able to make easy diagnosis, more accurate prediction and avoiding unnecessary biopsys would enable to have better outcomes of diseases, as well as reducing the cost of unnecessary methods and allows for increased clinical trials. The breast N is the highest mortality rate in women after lung N. The purpose of this study is to investigate early detection of breast N by machine learning methods, avoiding unnecessary biopsy. In previous studies, while variables including tumor information were used for diagnosis, in this study, variables that were mostly cultural and physical influences were used. In early diagnosis, it was investigated whether these variables were important or not, and it was also compared with studies done with different variables. Machine learning classification methods for breast N diagnosis have been used. Each method has its own size reduction in the variables according to the suitability of the method. The most effective variable according to each method varied. Two different applications were made on the same data set. The data set is divided into training and test set and applied in SPSS.18 Modular program. The second is the application in Weka with k-fold cross validation. Logistic regression and Bayesian network are the best result of the machine learning methods used. After that, they are followed by support vector machines, CRT and C5.0 algorithm in decision trees and neural network. Logistic Regression (78%) and Naive Bayes (78%) were the best results of machine learning methods used. These are followed by Support Vector Machines (76%), Decision Tree C5.0 algorithm (76%), and Artificial Neural Networks (74%) respectively. 3 fold cross validation performed in the case, the methods that give the best classification accuracy are Decision Tree C5.0 algorithm (76%), Logistic Regression (73%), Support Vector Machines (73%), Artificial Neural Networks (73%) and Naive Bayes (73%). Key Words: Machine Learning, Logistic Regression, Support Vector Machines, Decision Trees, Artificial Neural Networks
Author
Meliha Nur Durak
Institution
How to Cite
Meliha Nur Durak (Master Thesis). Early diagnosis of breast cancer with machine learning classification methods, 2017, Yıldız Technical University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Yıldız Technical University
- An investigation on the relationship between problem solving and critical thinking skill, and academic achievement of vocational and technical high school students(2017)
- Examining ?Historical housing structures" within the confines of protecting ecological balance(2012)
- Approximate solutions of integral equations(2012)
- The annotative dictionary of Kutadgu Bilig in terms of vocabulary(2013)
- Stepper motor speed control with labVIEW(2014)
- Determining supply chain risk factors in food industry(2014)