Çok-kipli konuşmacı ve konuşma tanıma uygulamaları için dudak devinim öz niteliklerinde ayırıcı analiz
2005
0 views
0 downloads
Advisor: Prof. Dr. Murat Tekalp ; Yrd. Doç. Dr. Engin Erzin
Abstract (EN)
In this thesis a new multimodal speaker/speech recognition system that integrates audio, liptexture, lip geometry, and lip motion modalities is presented. There have been several studiesthat jointly use audio, lip intensity and/or lip geometry information for speaker identificationand speech recognition applications. This work proposes using explicit lip motioninformation, instead of or in addition to audio, lip intensity and/or geometry information, forspeaker identification and speech-reading within a unified feature selection and discriminationanalysis framework, and addresses two important issues: i) Is using explicit lip motioninformation useful? and ii) if so, what are the best lip motion features for these twoapplications? The best lip motion features for speaker identification are considered to be thosethat result in the highest discrimination of individual speakers in a population, whereas forspeech-reading, the best features are those providing the highest phoneme/word/phraserecognition rate. The audio modality is represented by the well-known mel-frequency cepstralcoefficients (MFCC) along with the first and second derivatives, whereas lip texture modalityis represented by the 2D-DCT coefficients of the luminance component within a boundingbox about the lip region. Several lip motion feature candidates are considered including densemotion features within a bounding box around the lip, lip contour motion features, lip shapefeatures, and combinations of them. Furthermore, a novel two-stage discriminant analysis isintroduced to select the best lip motion features for speaker identification and speech-readingapplications. The fusion of audio, lip texture and lip motion modalities is performed by the so-called Reliability Weighted Summation (RWS) decision rule. Experimental results show thatthe proposed discriminative analysis significantly improves the unimodal performance of thelip motion modality. Moreover, using explicit lip motion information in addition to audio andlip texture yields further performance gains in bimodal speaker/speech recognition systems.
Author
Dr. Hasan Ertan Çetingül
How to Cite
Hasan Ertan Çetingül (Master Thesis). Çok-kipli konuşmacı ve konuşma tanıma uygulamaları için dudak devinim öz niteliklerinde ayırıcı analiz, 2005, Koç University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Koç University
- Ekom-Eczacıbaşı'nın Rusya piyasasındaki pazarlama stratejileri(1995)
- Barok döneminde Balkanlar Osmanlı Avrupası'nda mimaride, dekorasyonda, himaye ve kültürel üretim modellerinde dönüşüm, 1718-1856(2006)
- Erteleme kısıtlı tek makine çizelgeleme(2014)
- Sarayda Osmanlı tütsüleme gelenekleri: Topkapı Sarayı buhurdanları(2015)
- Selçuk Rumları ve Gürcistan Krallığının Birbirlerine olan benzerlikleri: 13. Yüzyılda sanatsal değişim çerçevesi(2015)
- Obje tabanlı akıl danışma-tavsiye iletişimi tasarımına ilham kaynağı olarak Türk kahve falı(2017)
