DoktoraAçık Erişim

Default prediction with graph theory in big data and interpretation of machine learning models

2021
0 görüntülenme
0 i̇ndirme
Danışman: Prof. Dr. Suat Özdemir

Özet (EN)

In recent years, the increase in the number of data sources, the decrease in the data collection, storage and processing costs, and the development of new methods for data analysis have led to the beginning of a new era called big data. Big data technologies have enabled the process of the data that could not be managed and processed before and explore valuable information hidden in the data. In this study, big data technology is used for the default prediction of companies. Default prediction has been studied for many years in the literature and its importance is increasing day by day. Two different default prediction models are proposed using Machine Learning (ML) and graph theory on a big data platform. In the study, credit, balance sheet and invoice datasets of more than 1 million real sector companies operated in Turkey between 2010 and 2018 are used. In the first model, two sub-models are created for credit and balance sheet datasets by using statistics and ML algorithms, and the probability scores obtained from these sub-models are combined to reach the best estimate in the final model. In the second model, graph theory is employed. It is based on the basic assumption that the internal dynamics of the companies, as well as the suppliers and customers, with whom they have commercial relations, are also important in the default of the companies. Therefore, a graph showing the commercial relationship is created using the invoice data of the companies. New variables that help explore default prediction are generated on the graph. These variables are further used in the second model. The results showed that both models achieved 0.81 and 0.82 Area Under Curve (AUC) scores, respectively. The higher prediction success of the second model showed that the new variables obtained from the graph contributed to the default prediction. Within the scope of the thesis, finally, a solution has been sought with Interpretable Machine Learning (IML) algorithms for the interpretability of the results, which is the most important criticism regarding the use of complex ML algorithms in default prediction. The interpretability results also indicated that IML gives consistent and reliable outcomes in explaining complex ML models.

Yazar

Dr. Mustafa Yıldırım

Bu Yayına Nasıl Atıf Yapılır

Mustafa Yıldırım (Doctorate thesis). Default prediction with graph theory in big data and interpretation of machine learning models, 2021, Gazi University.

Anahtar Kelimeler

Lisans

Tüm Hakları Saklıdır

Bu eser belirtilen lisans koşulları altında paylaşılmaktadır.

Gazi University tezlerinden daha fazlası