Master'sOpen Access

Evaluation of performance metrics and test techniques on various data sets in machine learning classification methods

Is this your thesis?

This record came from a bulk archive import. If it’s yours, link it to your profile.

2020
0 views
0 downloads

Abstract (EN)

One of the most important problems when creating a model in machine learning classification methods is the selection process of the best classifier. In the selection of the correct classifier, it is very important to select the data set, the training section reserved for creating the model and the test section used in the testing phase of the model. In the study, hold-out and 10-fold cross-validation methods were used for division. After the model is created, some metrics are used to evaluate the performance of the classifier. Within the scope of this thesis, nine classifiers have been applied to 32 data sets whose data distribution and decision class distribution are different from each other. The Python programming language, which is an open source language, and the Sklearn, Pandas, Numpy, Seaborn and Matplotlib libraries were used to create models. With these classifiers, models were created using hold-out and cross-validation methods. The complexity matrix was used to evaluate the performance of the classifiers and the accuracy, precision, recall, F1, MCC and AUC values were calculated for each model with the help of the confusion matrix. It was observed that MCC and AUC values gave more accurate results in the classifier selection in unbalanced data sets. When the obtained results are analyzed, although the hold-out method has yielded better results than the cross verification method in twenty datasets, the cross verification method was chosen while selecting the model. The reason for this may be a distribution of data that he never saw during the model training or testing in the hold-out method. When the models created with travel insurance and ph recognition data sets are analyzed, it is seen that in some cases, the high performance value is worse than the low performance value.

Author

Abdullah Alan

How to Cite

Abdullah Alan (Master Thesis). Evaluation of performance metrics and test techniques on various data sets in machine learning classification methods, 2020, Fırat University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Fırat University