Yüksek LisansAçık Erişim

Çoklu-ortam ses-görüntü işleme ile biometrik konuşmacı tanıma

2003
0 görüntülenme
0 i̇ndirme
Danışman: Prof. Dr. Murat Tekalp ; Yrd. Doç. Dr. Engin Erzin ; Yrd. Doç. Dr. Yücel Yemez

Özet (EN)

ABSTRACT In this thesis we present a multimodal text-dependent speaker identification system. The objective is to improve the recognition performance over conventional unimodal or bimodal schemes. The proposed system decomposes the information existing in a video stream into three modalities: voice, face texture and lip motion. Lip motion between successive frames is first computed in terms of eigenlip coefficients and then encoded as a feature vector. The feature vectors obtained along the whole stream are linearly interpolated to match the rate of the speech signal and then fused with mel frequency cepstral coefficients (MFCC) of the corresponding speech signal. The resulting joint feature vectors are used to train and test a Hidden Markov Model (HMM) based identification system. Face texture images are treated separately in eigenface domain and integrated to the system through decision-fusion. Experimental results are also included for demonstration of the system performance. IV

Yazar

Dr. Alper Kanak

Bu Yayına Nasıl Atıf Yapılır

Alper Kanak (Master Thesis). Çoklu-ortam ses-görüntü işleme ile biometrik konuşmacı tanıma, 2003, Koç University.

Anahtar Kelimeler

Lisans

Tüm Hakları Saklıdır

Bu eser belirtilen lisans koşulları altında paylaşılmaktadır.

Koç University tezlerinden daha fazlası