Reconstruction of normal speech from dysphonic speech by using mixed excitation linear predictive coding method
2008
0 views
0 downloads
Advisor: Yrd. Doç. Dr. M. Elif Karslıgil
Abstract (EN)
Speech is one of the best effective ways of daily communication. While producing voice, vocal cords vibrate when air is forced through them and along the vocal tract. Voice shaped by palate, tongue and lips and gets form to speech. The intensity of the voice depends on the subglottic pressure produced by lungs. Voice can be adjusted by varying the shape and the tension in the vocal cords, and the pressure of the air behind them. Dysphony is the result of the neural, structural or pathological effects on the vocal cords or larynx and it causes undesirable changes in the pitch, amplitude and the quality of the speech.Medical and mechanical solutions like electrolarynx and voice prosthesis are proposed for the patients who have lost ability to speak. However, these techniques have infection risk or are poor quality. Although studies about enhancing the speech that is produced by these techniques have begun to be seen in the literature for 15 years, there is not a comprehensive study on alternative solution of medical and mechanical techniques. The proposed system in this thesis can be used as speech prosthesis for the dysphony patients.In the proposed system, MELP (Mixed Excitation Linear Predictive Coding) was used for synthesizing enhanced speech. Unvoiced phonemes were detected for dysphonic speech. Pitch and voicing were produced, formant modification was applied for the phonemes except unvoiced phonemes. Correlation between pitch and formant frequencies was used in order to produce pitch.Spectral distance was calculated and subjective listening tests were applied in order to discuss the synthetic speech quality. It is observed that similarity between synthetic speech and normal speech is %20 higher compared to the similarity between dysphonic speech and normal speech. In the future, an alternative method that can help patients who are lack of ability to communicate effectively can be developed by increasing the synthetic speech quality and constructing a real time embedded system.Keywords: Dysphony, Voice Modification Techniques, Mixed Excitation Linear Predictive Coding
Author
H. İrem Türkmen
How to Cite
H. İrem Türkmen (Master Thesis). Reconstruction of normal speech from dysphonic speech by using mixed excitation linear predictive coding method, 2008, Yıldız Technical University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Yıldız Technical University
- Examining ?Historical housing structures" within the confines of protecting ecological balance(2012)
- Approximate solutions of integral equations(2012)
- The annotative dictionary of Kutadgu Bilig in terms of vocabulary(2013)
- Stepper motor speed control with labVIEW(2014)
- Determining supply chain risk factors in food industry(2014)
- TiO2/Cu2O ince film fotovoltaik hücrelerin karakterizasyonu(2014)
