Default prediction with graph theory in big data and interpretation of machine learning models
2021
0 views
0 downloads
Advisor: Prof. Dr. Suat Özdemir
Abstract (EN)
In recent years, the increase in the number of data sources, the decrease in the data collection, storage and processing costs, and the development of new methods for data analysis have led to the beginning of a new era called big data. Big data technologies have enabled the process of the data that could not be managed and processed before and explore valuable information hidden in the data. In this study, big data technology is used for the default prediction of companies. Default prediction has been studied for many years in the literature and its importance is increasing day by day. Two different default prediction models are proposed using Machine Learning (ML) and graph theory on a big data platform. In the study, credit, balance sheet and invoice datasets of more than 1 million real sector companies operated in Turkey between 2010 and 2018 are used. In the first model, two sub-models are created for credit and balance sheet datasets by using statistics and ML algorithms, and the probability scores obtained from these sub-models are combined to reach the best estimate in the final model. In the second model, graph theory is employed. It is based on the basic assumption that the internal dynamics of the companies, as well as the suppliers and customers, with whom they have commercial relations, are also important in the default of the companies. Therefore, a graph showing the commercial relationship is created using the invoice data of the companies. New variables that help explore default prediction are generated on the graph. These variables are further used in the second model. The results showed that both models achieved 0.81 and 0.82 Area Under Curve (AUC) scores, respectively. The higher prediction success of the second model showed that the new variables obtained from the graph contributed to the default prediction. Within the scope of the thesis, finally, a solution has been sought with Interpretable Machine Learning (IML) algorithms for the interpretability of the results, which is the most important criticism regarding the use of complex ML algorithms in default prediction. The interpretability results also indicated that IML gives consistent and reliable outcomes in explaining complex ML models.
Author
Dr. Mustafa Yıldırım
How to Cite
Mustafa Yıldırım (Doctorate thesis). Default prediction with graph theory in big data and interpretation of machine learning models, 2021, Gazi University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Gazi University
- The effect of computer-assisted and direct strategy teaching on reading comprehension(2021)
- Experimental development of the interfacial bond-slip model between textile reinforced mortar strips and masonry walls(2025)
- Deveplopment of semiconductor humidity sensors(2021)
- Death in plastic arts(2021)
- Consumption preferences of university students: Ankara Haci Bayram Veli University and Çankaya University examples(2021)
- The effect of health education and progressive muscle relaxation exercise on vasomotor symptoms and sleep problems in women with perimenopausal period(2021)
