DoktoraAçık Erişim

Feature-based timbre comparison via machine learning and interface design

2025
0 görüntülenme
0 i̇ndirme
Danışman: Prof. Dr. Abdurrahman Tarikci

Özet (EN)

In contemporary times, machine learning technologies have become one of the most defining elements of the big data era. The rapid increase in the amount of data produced and stored in digital environments has rendered the processes of classifying, organizing, and interpreting such data unsustainable through manual methods. At this point, machine learning has assumed a revolutionary role in data organization and content-based information retrieval systems, owing to its capacity to automatically identify patterns within complex datasets and generate high-accuracy predictions based on these patterns. Particularly in the analysis of textual, visual, and auditory content, these technologies have not only facilitated access to information but also initiated a significant transformation in the automation of production processes. This thesis focuses on the reflections of this transformation within the domain of music production. With the advancement of digital audio technologies, music producers can now access millions of sampled sounds through online platforms. However, this vast diversity has made the selection of the right sound a considerable challenge in terms of both time and efficiency. Building upon this problem, the present study proposes an interface system based on timbral similarity analysis utilizing machine learning techniques. The aim of the research is to develop a system capable of automatically classifying and detecting similarities among sounds through content-based analysis, thereby accelerating sound selection and optimizing sound organization in the music production workflow. Within the scope of the study, a dataset consisting of 6,435 one-shot audio samples was constructed, encompassing seven instrument classes: kick, snare, rim, clap, cymbal, hat, and bass. From these sounds, 42 spectral and temporal features were extracted. The obtained features were trained using both classical machine learning models and a deep learning-based artificial neural network (ANN) architecture. The ANN model was re-trained using three different feature sets (all features, the top 20 most effective features, and the top 10 features), and their performances were compared based on accuracy, precision, recall, and F1-score metrics. The findings revealed that the deep learning architecture achieved higher accuracy, particularly for low-frequency instrument classes such as kick and bass. The representations obtained from the embedding layer of the ANN formed multidimensional vector spaces that encoded the timbral similarities of the sounds, allowing the similarity analysis to be directly integrated into the categorization process. A Streamlit-based interactive interface was developed as part of the thesis, enabling users to upload their own sound files and analyze timbral similarities among them. The interface can list similar sounds, visualize classification results, and operate locally in an offline environment. In conclusion, this study presents an innovative approach to integrating machine learning and deep learning technologies into data organization, sound analysis, and creative production processes within the field of music technology. Owing to its open-source structure, the developed system offers a reproducible, extensible, and practical solution for both academic researchers and music producers.

Yazar

Can Paşa

Bu Yayına Nasıl Atıf Yapılır

Can Paşa (Doctorate thesis). Feature-based timbre comparison via machine learning and interface design, 2025, Ankara Music and Fine Arts University.

Anahtar Kelimeler

Lisans

Tüm Hakları Saklıdır

Bu eser belirtilen lisans koşulları altında paylaşılmaktadır.

Ankara Music and Fine Arts University tezlerinden daha fazlası