Master'sOpen Access

Deep learning based large-scale document clustering with content-based similarity

2021
0 views
0 downloads
Advisor: Prof. Dr. Muhammet Ali Akcayol

Abstract (EN)

Nowadays, the size of data is increasing rapidly. It is not possible to process very large data in a short time even with today's technology. Hence, clustering, which requires organizing large numbers of large-scale documents into small numbers of interrelated and meaningful clusters, has become an important research topic. Deep learning methods, which have been successfully applied in many fields in recent years, can also be used successfully in unsupervised learning applications. In this study, a deep learning-based model has been developed for clustering based on content similarity in large-scale documents. CNN and LSTM networks have been used together in the developed deep learning model. A data set consisting of 386 English textbooks with a total size of 7.61 GB has been used to test the developed model. In experimental studies, 18 different clusters with an average accuracy of 66% have been obtained. The experimental results have shown that the clusters obtained with the developed model had higher success than the k-means and CURE clustering algorithms, which are widely used in the literature. The clusters created with the developed model have values of 0.65 NMI and 0.59 AMI. In addition, values of 0.81 and 0.95 have been obtained for the internal evaluation metrics Silhouette and Davies-Bouldin, respectively.

Author

Dr. Kevser Özdem

How to Cite

Kevser Özdem (Master Thesis). Deep learning based large-scale document clustering with content-based similarity, 2021, Gazi University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Gazi University