Stacked frequency-timeGRUs for continuous arousal recognition from musical audio
2021
0 views
0 downloads
Advisor: Prof. Dr. Yücel Yemez ; Prof. Dr. Engin Erzin
Abstract (EN)
We address the problem of continuous arousal detection for emotion recognition in musical audio pieces where emotions are represented in the two-dimensional arousal-valence space. We propose a novel method which is a combination of two recurrent neural networks using mel-spectrogram features: A bidirectional GRU network along the frequency dimension as a feature extractor, stacked with a GRU network along the temporal dimension, which is unidirectional for real-time adaptability. The method is evaluated on the MediaEval2015- Emotion in Music Dataset, achieving an RMSE of 0.215 which is better than the results reported by real-time adaptable state-of-the-art models.
Author
Aslıhan Çeliker
Institution
How to Cite
Aslıhan Çeliker (Master Thesis). Stacked frequency-timeGRUs for continuous arousal recognition from musical audio, 2021, Koç University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Koç University
- Obje tabanlı akıl danışma-tavsiye iletişimi tasarımına ilham kaynağı olarak Türk kahve falı(2017)
- State-building in multi-ethnic borderlands: Nationalizing Eastern Anatolia and Transylvania in interwar Turkey and Romania(2021)
- Cross-cultural and artistic dialogues in the seventeenth century constantinople/istanbul: The Iconography of Madonna della Misericordia and the Galata Icon(2024)
- Ekom-Eczacıbaşı'nın Rusya piyasasındaki pazarlama stratejileri(1995)
- Barok döneminde Balkanlar Osmanlı Avrupası'nda mimaride, dekorasyonda, himaye ve kültürel üretim modellerinde dönüşüm, 1718-1856(2006)
- De Rham-Witt kompleks(2011)
