Master'sOpen Access

Classification of environmental sounds with convolutional neural networks

2021
0 views
0 downloads
Advisor: Dr. Öğr. Üyesi Özkan İnik

Abstract (EN)

The environmental activities consist of living or non-living things. The sound data that can represent the results of these activities and also provide information about the environment gains importance. Sound data is used to obtain basic information and the operation of activities such as noise pollution, traffic problems, security systems, smart tracking systems, health services, local services in cities. In this sense, Environmental Sounds Classification (ESC) gains critical importance. Due to the increasing amount of data and time constraints in analysis, new and strong artificial intelligence methods that enable automatic recognition of sound are needed. Due to the high performance that Deep Learning (DL) architectures have achieved in many different areas in recent years, it is aimed to perform ESC process with DL architectures. Within the scope of this thesis, the Convolutional Neural Networks (CNN) models, which are the basic architecture of Deep Learning, have been designed for the classification of three different CNN dataset. CNN models that have the highest accuracy have been obtained among more than one CNN models designed originally for each dataset. These data sets are ESC10, ESC50 and UrbanSound8K data sets, respectively. The sound recordings in these datasets have been converted to image formats that has 32x32x3 and 224x224x3 dimensions. Thus, dataset that a total of six different image format were obtained. The original CNN models developed to classify these data sets are named as ESC10_CNN32, ESC10_CNN224, ESC50_CNN32, ESC50_CNN224, URBANSOUND8K_CNN32 and URBANSOUND8K_CNN224, respectively. These models have been trained by performing 10-fold Cross Validation on the datasets. In the results obtained, the average accuracy rates of ESC10_CNN32, ESC10_CNN224, ESC50_CNN32, ESC50_CNN224, URBANSOUND8K_CNN32 and URBANSOUND8K_CNN224 models were found to be 80.75%, 82.25%, 54.55%, 72.15%, 88.60% and 84.63%, respectively. When the results obtained are compared with other studies in the literature on the same data sets, it was seen that the proposed models achieved better results in the ESC10 and UrbanSound8K data sets. In the ESC50 dataset, it was found to be better than other studies, except for one study.

Author

Dr. Yalçın Dinçer

How to Cite

Yalçın Dinçer (Master Thesis). Classification of environmental sounds with convolutional neural networks, 2021, Tokat Gaziosmanpaşa Üniversity.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Tokat Gaziosmanpaşa Üniversity