Forecasting modellings based on K- means algorithm
Is this your thesis?
This record came from a bulk archive import. If it’s yours, link it to your profile.
Abstract (EN)
The number of collected data in our epoch is increasing rapidly and will continue to increase continuously. Machine learning, a research field that is growing popularity day by day. Machine learning algorithms allow for the build models based on the data. It is very important to take advantages of the data and to extract valuable knowledges. Sub-headings of this field are generally called supervised learning algorithms, clustering algorithms and association rules. We will use clustering algorithms within the thesis. The study will determine how many clusters should be separated and then what we can do with the clustered data for reservation processes. The purpose of this thesis is to realize reservation optimization with the prediction of the conversion rate based on clustering algorithms. Hierarchical and K- Means algorithms have been studied. Since it is understood that the results obtained from the hierarchical clustering are not appropriate, therefore it has been decided to study in detail by K-means algorithms for the continuation of the study. A dynamic clustering is also achieved by using the K-means algorithm. After the data preparation steps such as outlier analysis, data scaling, filling in missing observations, it was decided how many clusters of data should be separated by Gap statistic value and Elbow method. The k to be determined at this stage is kept in a wide range in the first stage. The true k value between 2 to 15 is selected by the plotted graphs is determined by the Gap statistic and the Elbow method. Their errors are calculated. After the determined values, the clustering is done by K-means algorithms. If the true k values determined by these methods are different, the k values found with the clustering validation criteria by Davies - Bouldin, Dunn, Calinski - Harabasz, Wemmert Gancarski and Silhouette are more precisely determined. In this study, the optimal k value determination methods and clustering evaluation criteria are used together, and the answer is given to how many clusters of data should be separated. A conclusion of the study, on the tourism sector, we work on ten forecasting models by the use of the real data for tourism company. The results from the forecasting models are found, compared and then the model with a remarkable success is obtained. Keywords: Cluster Analysis, K-Means Algorithm, Elbow Method, Cluster Validation Techniques.
Author
Kemal Koşuta
Institution
Yıldız Technical University
Matematik Mühendisliği Bilim Dalı
How to Cite
Kemal Koşuta (Master Thesis). Forecasting modellings based on K- means algorithm, 2018, Yıldız Technical University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Yıldız Technical University
- A ruled-based software model in the preliminary design phase of housing: Example of Mardin(2008)
- An investigation on the relationship between problem solving and critical thinking skill, and academic achievement of vocational and technical high school students(2017)
- Numerical investigation of single bubble flow boiling both on a flat surface and on a cavity in horizontal pipe and pool boiling on a flat surface(2018)
- Examination of the categories in which buildings with LEED v3 and v4 certificates score the lowest points in Turkey(2018)
- Examining ?Historical housing structures" within the confines of protecting ecological balance(2012)
- Approximate solutions of integral equations(2012)