Comparative analysis of data mining classification methods with data science survey data set
2021
0 views
0 downloads
Advisor: Dr. Öğr. Üyesi Arafat Şentürk
Abstract (EN)
Data Mining technology is a technology that is increasing its popularity day by day. One of the biggest reasons for its increasing popularity is the absence of a workspace limit. Data Mining technology, which belongs to the IT sector as a technical infrastructure, provides services to many sectors to provide convenience and advantage. Within the scope of the study, the preferred software language in Data Mining technology, the algorithm used, etc. The data set named "Data Science Questionnaire" is used, in which the criteria of the data scientists are accepted as input, and the output information about which sector they work in with inferences from this technical information preferred by data scientists. As a result of modeling the data set with Classification Algorithms C4.5 Algorithm, Random Forest Algorithm and K-Nearest Neighbor Algorithm, success rates evaluations are mentioned. While comparing the success rate of the models, the algorithms belonging to the Classification method used both the original and the processed data set. When model success rates are evaluated on the basis of data sets, the success rates of models created using the original data set are increased by 14-15% when modeled using the processed data set created after the data preprocessing stage. When the processed data set is modeled with selected classification algorithms (C4.5, Random Forest and KNN) and the default algorithmic features of these algorithms, the deviation rate is very low when the success rates are compared on the basis of algorithms. When the success rates of the algorithms are evaluated on the basis of the original data set used before the preprocessing and the processed data set used after the preprocessing, the deviation value becomes more evident. In addition, the success rate values were observed for the situations that would cause deviations in the model success rate, such as the "k" attribute value, which is specific to the KNN algorithm, taking different values or the Training-Test data set partitioning options. However, the effect of the mentioned situations on model success is not as clear as the effect of the preprocessing stage on model success. By inferring from these comparisons, the importance/effect levels of Data Mining stages were evaluated in order to create successful models, and Data Mining stages were interpreted by using the concepts of "cyclicality" and "subjectivity".
Author
Elvan Kübra Doğan
How to Cite
Elvan Kübra Doğan (Master Thesis). Comparative analysis of data mining classification methods with data science survey data set, 2021, Düzce University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Düzce University
- A review of Cem Akaş's novels(2021)
- New midpoint type inequalities for generalized fractional integrals(2021)
- Material culture in Mostarli Hasan Ziya'i Divan(2021)
- The life of Ebu'l-Hasen Ali b. Ahmed b. Muhammed en-Nîsâbûrî el-Vâhidî and his method in the tafsir named el-Vecîz fî Tefsîr-i Kitabi'l-Azîz(2021)
- Intertextuality in Alev Alatlı's novel's(2022)
- Visual interpretations on dark humor(2022)
