DoktoraAçık Erişim

New map/reduce programming algorithm model in distributed hadoop clusters

2024
0 görüntülenme
0 i̇ndirme
Danışman: Prof. Dr. Resul Kara

Özet (EN)

Big data are often stored close to the locations where they are generated, owing to the cost of data transfer. These stored data are moved to a single location for processing or processed at that location. In the literature, it is possible to find different methods for processing data in distributed datacenters. In this study, we present a new method for data processing called GSelf-MapReduce. In the proposed method, shuffling is performed among heterogeneous datacenter (DC) that complete the data-processing process. To calculate the data processing cost of the reduced function of the DCs, a polynomial regression model was created using the data obtained in the test environment, and the coefficients obtained from this model were used in the decision process. The key/value pairs to be shuffled are distributed according to the cost of the DCs, considering their location. Because the data to be shuffled between DCs do not wait for all DCs to complete their jobs, the cost is reduced both in terms of the data to be moved and the data to be processed. The performance of the proposed method was compared with that of four different distributed data processing methods in the literature. As a result, this work generates 15% less shuffled data than the closest work.

Yazar

Dr. Emin Şeşen

Bu Yayına Nasıl Atıf Yapılır

Emin Şeşen (Doctorate thesis). New map/reduce programming algorithm model in distributed hadoop clusters, 2024, Düzce University.

Anahtar Kelimeler

Lisans

Tüm Hakları Saklıdır

Bu eser belirtilen lisans koşulları altında paylaşılmaktadır.

Düzce University tezlerinden daha fazlası