Master'sOpen Access

Reconstruction of normal speech from dysphonic speech by using mixed excitation linear predictive coding method

2008
0 views
0 downloads
Advisor: Yrd. Doç. Dr. M. Elif Karslıgil

Abstract (EN)

Speech is one of the best effective ways of daily communication. While producing voice, vocal cords vibrate when air is forced through them and along the vocal tract. Voice shaped by palate, tongue and lips and gets form to speech. The intensity of the voice depends on the subglottic pressure produced by lungs. Voice can be adjusted by varying the shape and the tension in the vocal cords, and the pressure of the air behind them. Dysphony is the result of the neural, structural or pathological effects on the vocal cords or larynx and it causes undesirable changes in the pitch, amplitude and the quality of the speech.Medical and mechanical solutions like electrolarynx and voice prosthesis are proposed for the patients who have lost ability to speak. However, these techniques have infection risk or are poor quality. Although studies about enhancing the speech that is produced by these techniques have begun to be seen in the literature for 15 years, there is not a comprehensive study on alternative solution of medical and mechanical techniques. The proposed system in this thesis can be used as speech prosthesis for the dysphony patients.In the proposed system, MELP (Mixed Excitation Linear Predictive Coding) was used for synthesizing enhanced speech. Unvoiced phonemes were detected for dysphonic speech. Pitch and voicing were produced, formant modification was applied for the phonemes except unvoiced phonemes. Correlation between pitch and formant frequencies was used in order to produce pitch.Spectral distance was calculated and subjective listening tests were applied in order to discuss the synthetic speech quality. It is observed that similarity between synthetic speech and normal speech is %20 higher compared to the similarity between dysphonic speech and normal speech. In the future, an alternative method that can help patients who are lack of ability to communicate effectively can be developed by increasing the synthetic speech quality and constructing a real time embedded system.Keywords: Dysphony, Voice Modification Techniques, Mixed Excitation Linear Predictive Coding

Author

H. İrem Türkmen

How to Cite

H. İrem Türkmen (Master Thesis). Reconstruction of normal speech from dysphonic speech by using mixed excitation linear predictive coding method, 2008, Yıldız Technical University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Yıldız Technical University