Filter Variable Selection Algorithm and Knowledge Discovery in Datasets
2019
0 views
0 downloads
Advisor: Ersin Kuset Bodur
Abstract (EN)
This dissertation addressed two aspects within the Data Mining field: filter variable selection, and knowledge discovery in datasets. A filter algorithm that serves to reduce the feature space in datasets, with special attention to healthcare data, was developed and tested. The algorithm binarizes the dataset, and then separately evaluates the risk ratio of each predictor with the response, and outputs ratios that represent the association between a predictor and the class attribute which translates to the importance rank of the corresponding predictor. The performance of the developed algorithm was compared against some existing feature selection algorithms on different datasets, using classification models. In the majority of the cases, the predictors selected by the new algorithm outperformed those selected by the existing algorithms. The proposed filter algorithm is therefore a reliable alternative for variable ranking in data mining classification with a dichotomous response. In the aspect of knowledge discovery in datasets, the relationship between employees’ psychological capital (PsyCap) and educational qualifications, and the relationship between employees’ PsyCap and organizational tenure was mined. The PsyCap and demographic data of 329 employees in the hospitality industry were collected. The odds ratio (OR) technique was deployed to measure the associations which revealed that, employees with higher educational qualifications are 2.6 times more likely to have positive psychological capital than those with lower educational qualifications. It was also discovered that employees who have stayed longer periods within the service of an organization are 3.6 times more likely to be seen as having positive psychological capital compared with those who have stayed shorter periods. The results of the two associations are statistically significant at p-value = 0.002 < 0.05 and p-value = 0.004 < 0.05, respectively. These findings will guide business owners on the calibre of employees to hire, retrench, or retain during general recruitment or retrenchment. Keywords: Data mining, Classification, Attribute selection, Odds ratio, Filter algorithm, Balanced classification accuracy
Author
Dr. Donald Douglas Atsa’am
How to Cite
Donald Douglas Atsa’am (Doctorate thesis). Filter Variable Selection Algorithm and Knowledge Discovery in Datasets, 2019, Eastern Mediterranean University, Department of Mathematics.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Eastern Mediterranean University
- An Investigation on Time and Cost Overrun in Construction Projects(2012)
- Radial Power-Law Position-dependent Mass, Cylindrical Coordinates, Spectral Signatures(2015)
- Predicting performance level of reinforced concrete structures subject to corrosion as a function of time(2012)
- Discussion of Conservation Approaches for the Selected Heritage Buildings in the Walled City of Famagusta(2019)
- High School Students' Learning Styles in North Cyprus(2011)
- Afyonkarahisar İl Merkezinde Yaşayan 18 Yaş ve Üzeri Kadınların Diyet Posasıyla İlgili Bilgi Düzeylerinin ve Posa Alım Miktarlarının Belirlenmesi(2018)
