New feature approaches based on spatial lip points in visual-based lip reading applications
2021
0 views
0 downloads
Advisor: Ramazan Tekin
Abstract (EN)
As a social being, human beings often communicate with people by talking in order to meet their needs. The act of speaking takes place as a result of the joint use of both sight and hearing. While the sounds are produced during the speech, the forms of the lip can be clearly observed. Lip reading is the technique of understanding speech by analyzing the movement of the lips, face and tongue in cases where the voice is not heard or distorted. Visual speech information plays an important role in automatic lip reading, especially when the sound is distorted or inaccessible. Despite the success of audio-image-based lip reading, visual-only lip reading is a very difficult problem due to difficulties in distinguishing sounds with similar lip movements. In this study, some new attribute approaches are presented in order to increase the success rate in visual only lip reading applications. In this study, two separate data sets were used in speaker-independent and speaker-dependent prediction applications. These data sets; AVLetters2, in which 26 letters in the Latin alphabet are repeated seven (7) times by five (5) speakers, and AVDigits, in which the 10 digits 0-9 are repeated nine (9) times by six (6) speakers. First of all, the facial elements and lips are separated and the lip borders are marked with 20 points. Later, the attribute approaches based on these spatial points, named Center-Euclidean-Distance (CED), Symmetric-Euclidean-Distance (SED) and Neighbor-Points-Angles (NPA), are applied to classifiers. Finally, using the classification algorithms named K-Nearest Neighbor algorithm (KNN), Random Forest (RF), Support Vector Machine (SVM), lip reading analysis was performed from video images to determine 26 characters and 10 numbers. As a result of the analysis, the best success results were found to be 45.934% for the AVLetters2 data set with the RF-CED method and 67.407% for the AVDigits data set using the KNN-CED method. When compared to other visual-only studies on these data sets, it was seen that quite high and successful results were obtained.
Author
Dr. Hamdullah Tung
Institution
How to Cite
Hamdullah Tung (Master Thesis). New feature approaches based on spatial lip points in visual-based lip reading applications, 2021, Batman University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Batman University
- The transcription and evaluation of Mardin district's population book registered with the number of 3736(2016)
- The mediating role of emotional well-being in the effect of physical environment quality, food quality, and service quality on revisit ıntention in restaurant businesses(2026)
- Experımental study on the effects of graphene oxıde nanopartıcle added bıodıesel–dıesel fuel blends on engıne performance and emıssıons(2026)
- Muhi̇bbi̇'s life, literal personality and transparent text of Tuhfetu'l-Ahyâr (Research-text-lexi̇con)(2017)
- The position of women in the cinema of Nuri Bilge Ceylan in the context of inequality gender roles(2020)
- Examination of Mosques, Madrases and Qur'an courses in terms of Qur'an education (Mardin example)(2023)
