Hybrid method application to solve classification problems in imbalanced datasets
2022
0 views
0 downloads
Advisor: Doç. Dr. Duygu Yılmaz Eroğlu
Abstract (EN)
Today, the improvements of collecting data technologies and decisions depending on the data-based consequently increased the interest of data mining recently. This interest lead to studies in different data types. These days, besides of numeric and categorical data, visual recognition, voice recognition, text mining etc. has developed many real life and science study. In addition to the main real-life application areas such as biomedical informatics, pattern recognition, fraud detection, natural language processing, medical diagnosis, face recognition, text classification, fault diagnosis, anomaly detection, the number of studies in new technologies such as autonomous vehicles, Industry 4.0, unmanned aerial vehicles it increased. In some of these studies, it was encountered that the data sets were unbalanced, in other words, one class label was significantly dominant over the other class/classes. In this case, although the classifiers predict the majority class correctly but they cannot predict the minority class correctly. This makes serious problem on quality check, medical diagnossis etc. In this study, hybrid method proposed a solution the classification problem in imbalanced datasets. The aim is to prevent the overfitting problem caused by oversampling and valuable data loss caused by undersampling in imbalanced data, and to obtain successful classification results. Firstly, the studies on the classification of imbalanced data were examined. Then another method was proposed considering all the studies advantages and disadvantages. Hybrid method was applied to eight datasets, then these datasets were classified with different types of classifiers, and the results were compared with the results of the balanced data set with the SMOTE method, which is frequently used in imbalanced data classification problems. The obtained results confirmed the success of the proposed method. By using the input quality and process parameters in the real yarn data to predict yarn breaks, has presented a decision support system that can prevent yarns from entering the weaving with a high correct prediction rate.
Author
Mestan Şahin Pir
How to Cite
Mestan Şahin Pir (Master Thesis). Hybrid method application to solve classification problems in imbalanced datasets, 2022, Bursa Uludağ Üni̇versi̇ty.
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Bursa Uludağ Üni̇versi̇ty
- The effect of subthreshold bipolar disorder symptomatologyon neuropsychological profiles in children and adolescents withattention deficit and hyperactivity disorder(2022)
- Analysis of Ayman al Otoom's "Ya Sâhibay al-Sijn" in terms of structure and content in the context of prison literature(2022)
- The discrete divisions of Hanefi fakihs in the field of criminal law(2020)
- Bayt al-Hikmah and its importance during translation period(2020)
- New security problem in 21th century: Climate refugees(2020)
- Une etude sur les valeurs educatives des livres pour enfants de Daniel Pennac et leurs exploitations en fle(2020)