Yüksek LisansAçık Erişim

Clustering analysis system based on K-means algorithm and its application in the retail sector

2018
0 görüntülenme
0 i̇ndirme
Danışman: Prof. Dr. Ayla Şaylı

Özet (EN)

Developing and changing environmental conditions, globalization of the internet, competition with different research and development activities and marketing methods, and difficulties in customers' satisfaction are increasing the importance of information obtained from data day by day. The analysis of the information using some methods, the interpretation of the obtained results by the subject matter experts and making future forecasts from historical data can be stated as data mining. Data mining for companies and businesses is an important and strategic tool that facilitates decision making and allows decision makers to make quick decisions. Separating of data into similar groups, clustering of data, is one of the most basic methods in data mining. In this thesis, customer buying behaviors will be analyzed using the K-Means algorithm which is one of the non-hierarchical clustering methods. With the clustered data, it will be determined which brand, which product, when and how much is preferred by different customer profiles. The aim of this thesis is to create a system that will provide advantages such as both creating demand and meeting the right demand at the right time considering the customer preferences, and also realize data analysis using this system for the company. The system to be used for the analysis during the thesis was developed in the Java language and the obtained results were visualized by graphics and tables. Thus, a dynamic clustering analysis system was established for the K-Means algorithm. The data file to be used in the analysis belongs to Migros Ticaret A.S. on the global powers of retailing and consists of actual data. The data were stored in a table in the MS-SQL database, and all data preparation operations were performed on this table. Before the thesis, an article reviewing the brand loyalty of the customers named "Brand Loyalty Analysis System Using K-Means Algorithm" with the same data file was studied first. In our article, the number of clusters for data analysis was estimated, regardless of a method. In addition, the selection of the initial centers was conducted randomly. The results that are general brand loyalty, brand loyalty based on item and brand loyalty based on category were published in an international journal. However, the analysis system has been improved by using some methods for selecting the number of clusters and selecting the initial centers. In this thesis, the error for each value is calculated by choosing values from 2 to 20, and how many clusters of the data should be separated (determining the optimal value) has been determined using the Elbow method. For the determined value; Maximin, Katsavounidis, PCA-Part, Var-Part and K-Means++ methods have been used to find the initial centers. With the Elbow method which is chosen as the method of determining the optimal value, the effect of cluster selection of different initial centers has been investigated. Clustering results have been evaluated using the Silhouette and Calinski-Harabasz criterions and were presented to the company. The developed analysis system can also be used as a decision support system for other companies and businesses. Keywords: Data Mining, Clustering Analysis, K-Means Algorithm, Elbow Method, Selection of Initial Centers, Clustering Validation Criterions

Yazar

Merve Üstünel

Bu Yayına Nasıl Atıf Yapılır

Merve Üstünel (Master Thesis). Clustering analysis system based on K-means algorithm and its application in the retail sector, 2018, Yıldız Technical University.

Anahtar Kelimeler

Lisans

Tüm Hakları Saklıdır

Bu eser belirtilen lisans koşulları altında paylaşılmaktadır.

Yıldız Technical University tezlerinden daha fazlası