Master'sOpen Access

Çoklu-ortam ses-görüntü işleme ile biometrik konuşmacı tanıma

2003
0 views
0 downloads
Advisor: Prof. Dr. Murat Tekalp ; Yrd. Doç. Dr. Engin Erzin ; Yrd. Doç. Dr. Yücel Yemez

Abstract (EN)

ABSTRACT In this thesis we present a multimodal text-dependent speaker identification system. The objective is to improve the recognition performance over conventional unimodal or bimodal schemes. The proposed system decomposes the information existing in a video stream into three modalities: voice, face texture and lip motion. Lip motion between successive frames is first computed in terms of eigenlip coefficients and then encoded as a feature vector. The feature vectors obtained along the whole stream are linearly interpolated to match the rate of the speech signal and then fused with mel frequency cepstral coefficients (MFCC) of the corresponding speech signal. The resulting joint feature vectors are used to train and test a Hidden Markov Model (HMM) based identification system. Face texture images are treated separately in eigenface domain and integrated to the system through decision-fusion. Experimental results are also included for demonstration of the system performance. IV

Author

Dr. Alper Kanak

How to Cite

Alper Kanak (Master Thesis). Çoklu-ortam ses-görüntü işleme ile biometrik konuşmacı tanıma, 2003, Koç University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Koç University