Master'sOpen Access

Parallelization of K-means and DBSCAN algorithms and use on analysis of big data on Hadoop and performance and competence comparison

2015
0 views
0 downloads
Advisor: Doç. Dr. Gökhan Silahtaroğlu

Abstract (EN)

All actions in our life are now carried out through computers. This technology, which is now at the center of businesses performed in almost all sectors, day by day necessitates steps that facilitate and accelerate performance of business processes. Indeed, faults encountered in computer-based works result in significant losses for companies in the long run; while uncontrolled management of growing data constitute the essence of these problems. All operations performed with large data can cause many problems such as data storage, data analysis and data display. These problems may end up with many problems, especially data loss, and once again indicates the need for carrying out works in this field. In this thesis, methods for working with big data have been explored and applications for performing fast operations and obtaining more stable results with such data have been provided. In this context, methods for joint use of data mining algorithms for large data have been primarily considered with performance evaluations and parallelization of data minin algorithms DBSCAN and K-means have been analyzed. The Hadoop technology has been analyzed and performance comparison has been made with Pig, Hive and Impala, also projects have been examined where Hadoop technologies could be used. It has also been observed that data mining algorithms on Hadoop can be used with Mahout.

Author

Furkan Kayım

How to Cite

Furkan Kayım (Master Thesis). Parallelization of K-means and DBSCAN algorithms and use on analysis of big data on Hadoop and performance and competence comparison, 2015, İstanbul Beykent University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from İstanbul Beykent University