FLAGS framework and decentralized federated learning under device volatility
2023
0 views
0 downloads
Advisor: Prof. Dr. Öznur Özkasap ; Yrd. Doç. Dr. Barış Akgün
Abstract (EN)
Federated Learning (FL) has become a key choice for distributed machine learning. Initially focused on centralized aggregation, recent works in FL have emphasized greater decentralization supported by standardization of serverless interaction in the next-generation communication networks. However, the diversity of devices, data distributions, and communication settings, compounded by dynamic operating conditions, result in multiple challenges for Decentralized FL (DFL). There have been various approaches to DFL, from utilizing intermediate edge servers to fully device-to-device approaches. In decentralized settings, communication cost and learning performance are usually assessed together and certain trade-offs are made based on scenarios. However, there is a lack of existing work on comparing DFL approaches in an apples-to-apples manner in a multitude of scenarios and operating conditions. To bridge this gap between methods and their comparative analysis, we design and develop the Federated Learning Algorithms Simulation (FLAGS) Framework. One important challenge we noticed that most DFL methods struggle with is the extreme fluctuations in device availability, especially for purely decentralized approaches. This \textbf{device volatility} leads to poor learning performance. To address this issue, we investigate the effects of neighborhood selection, memory and multi-hop information passing on DFL performance. We introduce a fully decentralized FL approach that can operate under realistic and highly volatile device participation settings. The key contributions of this thesis are as follows: (i) development of a lightweight FL framework for benchmarking a large plethora of methods, (ii) analysis and comparison of multiple FL methods with this framework under multiple operating conditions, (iii) empirical analysis of various node selection strategies under heavy device volatility, and (iv) utilizing memory and relayed communication to enhance device-to-device FL by developing a novel algorithm that can operate under realistic operating conditions and heavy device volatility. Federated Learning supports a wide variety of node interactions and autonomous operations across the network edge. With the aim to encompass this multi-faceted heterogeneity, the FLAGS framework was proposed and developed as a lightweight FL implementation and testing platform. FLAGS framework allows for a wide range of device behaviors and cooperation mechanisms, enabling rapid testing of multiple FL algorithms. FLAGS's built-in features allow it to subject existing and novel FL algorithms to a wide range of data distributions, simulating the nodes with multiple neural networks as well as participation conditions ranging from homogeneous to highly volatile. Different network tiers and communication mechanisms enable various FL algorithms to be configured by employing various combinations of the aforementioned factors. In order to consolidate this very extensive FL landscape and offer an objective analysis of the major FL algorithms, comprehensive cross-evaluations for a wide range of operating conditions have also been conducted. Starting with the three foundational FL algorithms, including Hierarchical FL (HFL), Decentralized FL (DFL), and Gossip FL (GFL), this work evaluates six derived algorithms ranging from fully centralized to fully decentralized. The experiments indicate that fully decentralized FL algorithms achieve comparable accuracy under multiple operating conditions, including asynchronous aggregation and the presence of stragglers. Furthermore, DFL can also operate in noisy environments and with a comparably higher local update rate. However, the impact of extremely skewed data distributions on DFL is much more adverse than on centralized variants. The analysis of the cross-evaluation indicates that DFL performance is considerably impacted by node participation. This part of the thesis focuses on improving DFL performance under realistic and volatile device behavior. Node selection in various forms has been experimented with to improve both communication efficiency and convergence rate. We experimented with multiple node selection mechanisms and also proposed and evaluated a time-varying parameterized node selection method for DFL employing validation accuracy and its per-round change. The mentioned criteria are evaluated using both hard and stochastic/soft selection on sparse networks. The results indicate that the bias associated with node selection adversely impacts performance as training progresses, and a uniform random selection is preferable under extremely limited participation conditions. Continuing with volatile conditions, we investigate and propose mechanisms to improve DFL operating on sparse graphs in the presence of stragglers and non-participating nodes. We first propose two algorithms: Memory-Assisted DFL (MA_DFL) and Augmented-Graph Assisted DFL (AG_DFL). These algorithms employ memory and selective relaying to improve DFL performance. Both algorithms outperform the baseline DFL and gossip interaction for volatile node participation. Then, we propose a hybrid of these two algorithms, Memory and Augmented-Graph Assisted DFL (MAG_DFL), that employs memory and graph augmentation to improve the performance of DFL under highly volatile devices and extreme data conditions. The research conducted in this thesis evaluates the multi-faceted challenges to the DFL operation in volatile conditions and proposes mechanisms to improve its performance. Our work indicates that DFL holds the potential to assist learning operations distributed across the edge network. It may be used to augment the FL in the presence of costly upstream communication or limited connectivity. However, node density has a major impact on DFL, and sparse networks, along with volatile device behavior and non-IID distributions, tend to reduce its convergence rate. The enhanced neighborhood interaction and intelligent use of local information has the potential to improve DFL performance under such adverse conditions based on the presented results. The analysis, algorithms and results presented in this thesis pave the way for additional developments and more practical applications of DFL in the next-generation communication networks.
Author
Dr. Ahnaf Hannan Lodhı
Institution

Koç University
Bilgisayar Bilimi ve Mühendisliği Bilim Dalı
How to Cite
Ahnaf Hannan Lodhı (Doctorate thesis). FLAGS framework and decentralized federated learning under device volatility, 2023, Koç University.
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Koç University
- Obje tabanlı akıl danışma-tavsiye iletişimi tasarımına ilham kaynağı olarak Türk kahve falı(2017)
- Ekom-Eczacıbaşı'nın Rusya piyasasındaki pazarlama stratejileri(1995)
- Barok döneminde Balkanlar Osmanlı Avrupası'nda mimaride, dekorasyonda, himaye ve kültürel üretim modellerinde dönüşüm, 1718-1856(2006)
- De Rham-Witt kompleks(2011)
- Erteleme kısıtlı tek makine çizelgeleme(2014)
- Sarayda Osmanlı tütsüleme gelenekleri: Topkapı Sarayı buhurdanları(2015)