DoctorateOpen Access

Automated audio captioning with acoustic and semantic feature representation

2023
0 views
0 downloads
Advisor: Doç. Dr. Mustafa Sert

Abstract (EN)

Today, audio data is increasing rapidly with the developing technology and the increasing amount of data. Therefore, there is a need for understanding and interpretation of the content of audio data by human-like systems. Generally, audio processing studies have focused on speech recognition, audio event/scene, and tagging to process audio data. Speech recognition aims to translate a spoken language into text. Audio event/scene and tagging studies make single or few-word explanations of an audio recording. Unlike the previous studies, automatic audio captioning aims to explain an environmental audio record with a natural language sentence. This thesis explores the importance of using semantic information to improve audio captioning performance after a detailed literature study on audio processing, image/video, and audio captioning. In this context, computational models have been developed using linguistic knowledge (subject-verbs), topic model, knowledge graphs, and acoustic events for audio captioning. As a methodology, the contributions of different features, word embedding methods, deep learning architectures and datasets, and the contribution of semantic information to audio captioning were examined. Within the scope of the studies, two publicly open audio captioning datasets were used. The success of the models proposed in the thesis was compared with the studies using the same datasets. The results show that the proposed methods improve AAC performance and give results comparable to the literature.

Author

Dr. Ayşegül Özkaya Eren

How to Cite

Ayşegül Özkaya Eren (Doctorate thesis). Automated audio captioning with acoustic and semantic feature representation, 2023, Başkent University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Başkent University