Master'sOpen Access

The usage of dissimilarity and similarity measures in statistics

2015
0 views
0 downloads
Advisor: Prof. Dr. Sadullah Sakallıoğlu

Abstract (EN)

Cluster analysis is defined as a multivariate statistical method for grouping observations which are unknown about their classifications, into clusters such that the observations in each cluster or group are similar and the observations across groups are as different as possible and getting some estimations about populations (Sharma, 1996). Dissimilarity measures are important in determining the efficiency of this method. In this study, the effects of similarity and dissimilarity measures on cluster analysis is investigated for binary data sets and as a validity index or criteria, the cophenetic correlation coefficient is used. For this purpose, the distance matrix depending on each of 47 measures, which of 37 are similarity, 10 are dissimilarity measures, is generated by using both real data and four artificial data sets which have different features and these datasets are analyzed. Besides, structures of binary similarity and dissimilarity measures are examined in terms of some theoretical conditions which are proposed to be satisfied. In other respects hierarchical clustering analysing methods are explained and their practical applications are examplifed. As a result of the findings, Hamann, Russell&Rao, Chord and Sokal&Sneath-1 measures are suggested as the measures providing best clustering structure.

Author

Hasan Yıldırım

How to Cite

Hasan Yıldırım (Master Thesis). The usage of dissimilarity and similarity measures in statistics, 2015, Çukurova University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Çukurova University