Master'sOpen Access

Analysis of document datas of sakarya metropolitan municipality using data mining methods

2025
0 views
0 downloads
Advisor: Prof. Dr. Nilüfer Yurtay

Abstract (EN)

This thesis aims to contribute to the improvement of service quality by analyzing document management processes in public institutions. In this context, a comprehensive analysis was carried out with the clustering algorithm, one of the data mining methods, using the outgoing document data of Sakarya Metropolitan Municipality for the year 2016. Document management has a critical role in public institutions in line with the principles of transparency, traceability and efficiency. In this context, it is aimed to make data-based inferences for the improvement of processes through attributes such as the creation time of documents, distribution status, related document references and additional content. The methodology of the study is based on the CRISP-DM (Cross-Industry Standard Process for Data Mining) model, which is widely used in the data mining process. In this framework, firstly, the business problem was defined and then the data collection, data understanding, data preprocessing, modeling, evaluation and application stages were carried out systematically. Missing values in the data set were cleaned, categorical data were digitized and normalization processes were applied to make them suitable for analysis. In the modeling process, clusters were formed using the K-Means algorithm on the data set. Elbow method was used to determine the optimal number of clusters. Thus, it was revealed whether document density shows certain periodic patterns and how these patterns can be integrated into personnel leave planning. As the first step of the implementation phase, outgoing document data of Sakarya Metropolitan Municipality for the year 2016 were obtained and the data mining process was initiated on these data. The data set consists of a total of 16535 datas and 9 variables. These variables cover areas that are considered critical in document management systems such as record number, document duration, month information, directorate and department to which it belongs, distribution, attachment and interest information. The data set was first analyzed for missing data and 10 datas were identified as missing and these records were removed from the system. Then, categorical data were converted into numerical form and made suitable for the analysis process. In addition, data transformation processes were carried out to analyze the variables in a more meaningful way. At this stage, all data were digitized and a structure suitable for the k-means clustering algorithm was obtained. In the second stage, the modeling step was started in line with the CRISP-DM process and the k-means algorithm was applied on the data set. At this stage, Elbow method was first used to determine the optimal number of clusters and as a result of graphical analysis, k=4 was selected as the optimal number of clusters. Then, clusters were formed using RapidMiner software and the common characteristics of the documents in the same clusters were analyzed. The results of the analysis revealed that the intensity of document output increased significantly in some periods compared to other periods. It was also found that during these intensity periods, the rate of inclusion of interest or additional content to the document also increased. This increases not only the volume of document processing, but also document complexity and processing time. In the last stage of the implementation, inferences were made in terms of internal processes in line with the clusters obtained. It was observed that document processing times are prolonged, especially during periods of high work intensity, and if the number of personnel using annual leave increases during these processes, disruptions in service flow may occur. In this context, it is recommended that personnel leave planning be made in a data-driven manner, taking into account document mobility and content density. In this way, service continuity can be maintained and staff satisfaction can be increased. This application is an important example to show that data mining methods can be integrated into decision support mechanisms in large public institutions such as municipalities. In the analysis process, the effects of document attributes such as distribution, interest and attachment information on document processing time were also analyzed. Correlation analyses revealed that processing times were longer for documents that contained interest and distribution information. This indicates that such documents require more control and content evaluation. In addition, linear relationships between variables were evaluated using correlation matrices and some statistically significant relationships were identified. The findings contributed to the understanding of the structural factors affecting document processing times and provided important insights for improving document management processes. In this way, it is emphasized that data-driven approaches can guide not only technical but also strategic planning in public institutions. As a result, this thesis reveals that document management processes carried out in public institutions are not limited to archiving, recording or fulfilling legal obligations. With the impact of developments in information technologies, digitalized document management systems within public institutions offer a significant potential for both tracking and analyzing institutional data. In this context, systematic examination of document movements makes it possible to discover meaningful patterns for improving service processes by analyzing parameters such as processing times, document qualities and periodic intensities. The application carried out within the scope of this thesis has shown that the operational information contained in municipal documents can be analyzed with data mining techniques and that data-based outputs that will shed light on corporate decision-making processes can be produced as a result of these analyzes. The groupings obtained with the K-means clustering algorithm provide valuable information especially for human resources planning and workforce management. The document exit times and periodic document density data revealed the workload distribution of the organization throughout the year, which in turn showed which periods of annual leave would be more appropriate in terms of service continuity. For example, identifying months with high document mobility indicates that staff taking leave during these periods may have a negative impact on service performance. Such data-driven inferences provide solutions that are much more rational, objective and in line with the operational reality of the organization compared to the leave processes planned by conventional methods. Thus, by maintaining continuity and quality in public services, it can be aimed to increase citizen satisfaction as well as employee satisfaction. The findings and application results of this study not only contribute academically, but also provide practical recommendations for key management areas such as strategic planning, process management and resource allocation in local governments. While the model proposed in the thesis enables organizations to use their data infrastructure more effectively, it helps to develop managerial insights by integrating data mining methods into decision support mechanisms. Especially in organizations such as municipalities that are in direct contact with citizens, such analytical approaches play a critical role in ensuring the sustainability of services and using resources efficiently. In this respect, the thesis makes an applied contribution to both data science and public administration literature and sets an example for similar studies.

Author

Dr. Engin Uçar

How to Cite

Engin Uçar (Master Thesis). Analysis of document datas of sakarya metropolitan municipality using data mining methods, 2025, Sakarya University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Sakarya University