Master'sOpen Access

Büyük medikal veri setlerinin yaklaşık spektral öbeklenmesi için jeodezik tabanlı benzerlik ölçütleri

2015
0 views
0 downloads
Advisor: Yrd. Doç. Dr. İsa Yıldırım

Abstract (EN)

Clustering is the unsupervised classification to group patterns such as data, observations or feature vectors. Due to its extraction of clusters in a fast manner without supervised information, many different methods have been proposed in various applications. Among them, spectral clustering (SC) has been recently popular and successfully used in various areas such as image processing, computer vision and information retrieval, thanks to its ability to find irregularly shaped clusters and its independence from parametric cluster models. Spectral clustering is a manifold learning algorithm based on eigendecomposition of a graph Laplacian matrix constructed from pairwise similarities of the data points. Although this eigendecomposition produces higher clustering accuracies than the accuracies obtained by traditional methods, it has high computational cost and memory requirement which makes direct use of spectral clustering infeasible for clustering large datasets. To address this challenge, approximate spectral clustering (ASC) methods, which apply spectral clustering on a reduced set of data points (data representatives) selected by sampling or quantization, have been proposed. The ASC not only makes spectral clustering feasible for large datasets but also enables new information types to be used in similarity definition. In this thesis, data representatives are obtained by selective sampling and neural gas as sampling and quantization method respectively because of being the best methods in the literature for sampling and quantization. In addition, k-means++, a clustering algorithm which solves random initialization problem of k-means, is first used to obtain data representatives as a quantization method for ASC. In order to achieve high clustering accuracies with ASC, an important step is to determine the criterion to define the pairwise similarities of the selected data representatives. Traditionally, an Euclidean distance based Gaussian kernel with a global decay parameter (optimally set by experiments) is used. Alternatively local decay parameters are also used. However, this approach ignores new information types such as data topology, local density distribution and data manifold which are provided by ASC. To utilize all the available information for accurate similarity definition, geodesic based hybrid similarity criteria are proposed in this study. The geodesic distance between any two data representatives is their shortest path distance depending on a neighborhood graph. In addition to commonly used k-nearest neighbor graph, which requires a user-set parameter k, a weighted Delaunay triangulation (CONN) is employed to determine the optimal number of neighbors with respect to local data characteristics without any user-set parameter. CONN also indicates data topology and detailed local density distribution to be used in similarity definition. Based on these neighborhood graphs, the proposed geodesic distance based similarities are defined using Euclidean distance, local density distribution by CONN, and their fused approach. Therefore, these proposed similarity criteria represent better pairwise similarities because they use different combination of all available information for ASC. The clustering performance of the proposed criteria is evaluated by an extensive experimental study. Their advantages are first shown on three artificial datasets with different clustering challenges. Then, they are shown more successful than the commonly used Euclidean based similarity, using ten datasets (including medical data as well) from UCI Machine Learning Repository. Finally the proposed geodesic hybrid similarity criteria based ASC is applied on four sets of brain MR images obtained from MICCAI 2012 challenge on multi-modal brain tumor segmentation, achieving an accuracy of up to 80.06% without any supervised information. This accuracy is much successful than traditional clustering methods (compared to 66.55% obtained by k-means and compared to 76.52% obtained by ASC based on similarity using Euclidean distance) and very close to supervised accuracies existing in the literature. The extensive experiments show the outperformance of the proposed criteria and favor them as a successful approach in clustering both large and small/medium datasets.

Author

Dr. Berna Yalçın

How to Cite

Berna Yalçın (Master Thesis). Büyük medikal veri setlerinin yaklaşık spektral öbeklenmesi için jeodezik tabanlı benzerlik ölçütleri, 2015, Istanbul Technical University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Istanbul Technical University