Master'sOpen Access

Distributed document classification and clustering on cloud

Is this your thesis?

This record came from a bulk archive import. If it’s yours, link it to your profile.

2015
0 views
0 downloads

Abstract (EN)

Big data is described as large and complex data sets which can not be stored and processed using traditional databases, file systems and algorithms.Therefore new technologies are being utilized to store and process these big data sets. Automatic content extraction, field and topic discovery, summarization or pattern recognition over very large document sets which are produced by various sources are the subjects of many current research. In this thesis, Big Data analysis, namely distributed classification and clustering algorithms are applied to a large data sets consisting Turkish scientific articles, To be able to run distributed Machine Learning algorithms Apache Mahout and Apache Spark are used. The servers needed for distributed classification and clustering algorithms are deployed on the Google Cloud, Amazon AWS and Microsoft Azure cloud computing infrastructures. Keywords: Big Data, Distributed Computing, Cloud Computing

Author

Selen Gürbüz

How to Cite

Selen Gürbüz (Master Thesis). Distributed document classification and clustering on cloud, 2015, Fırat University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Fırat University