DoctorateOpen Access

Topic classification of official correspondences with natural language processing and machine learning

2024
0 views
0 downloads
Advisor: Prof. Dr. Resul Kara

Abstract (EN)

In line with digital advancements, official correspondence documents in public institutions are managed through Electronic Document Management Systems (EDMS). Appropriate determination of the Standard File Plan (SFP) codes of documents is important for correct archiving and archival destruction process. The SFP code information given to the document by the people who created the document may be written incorrectly for various reasons. To prevent these errors, it would be useful to develop applications that automatically detect the correct SFP code of documents. For this purpose, two different data sets were created in the study; initially, preprocessing was performed on these sets, followed by the application of various classification algorithms on the preprocessed data to detect, the documents' SFP codes. The results of the classification processes were compared and analyzed. In the analysis of the first dataset, the most successful classification results were obtained by using the correctly predicting the SFP code of 978 out of 1000 official correspondence documents with the Logistic Regression (LR) algorithm. In the analyses performed on the second dataset, the most successful classification results were obtained with the Non-Negative Matrix Factorization (NNMF) algorithm, which classified 1851 of 2100 documents into the correct subjects (SFP code) and achieved 88.14% success rate.

Author

Zeynep Bozdoğan

How to Cite

Zeynep Bozdoğan (Doctorate thesis). Topic classification of official correspondences with natural language processing and machine learning, 2024, Düzce University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Düzce University