Yüksek LisansAçık Erişim

Lip reading with cascade object detection and artificial neural networks methods

2019
0 görüntülenme
0 i̇ndirme
Danışman: Dr. Öğr. Üyesi Cafer Bal

Özet (EN)

The conversion of image or sound into text or shape has become easier with the advancement of technology. Satisfactory results can be achieved even if the conversion of the said into text or figure is not at the desired level. However, it is not at a desirable level to fully elaborate what is said through the image search -lip reading process-. The reason for this is that the obstacles encountered in the lip reading process are more than the obstacles encountered in the processing of sound (in the analysis process). In this thesis, the speech content in video images was estimated by lip reading method. Lip reading is a complex process that requires the follow-up of lip movements by drawing the face, eyes, lip region, lips, and outer borders representing the lips in the image. In this study, firstly, basic image processing methods were examined and then the success of color spaces in determining the skin color was compared to find the face in the image. After finding the face in the image, the position of the face (right or left inclination) was checked.Since the inclination of the face means that the eyes and lips are also inclined, this required a recalculation of the outer boundary points of the lip according to the amount of inclination. The amount of inclination of the face was calculated according to the positions of the eyes. The eyes on the face were found by a hough transformation used to detect shapes such as ellipses, circles or lines within a defined range (region). It was used to recalculate the outer margins of the lip by locating the eyes relative to each other. Then, before and after the calculation of lip contours were compared. Then the lip area finding process was performed and results were compared using lip area finding methods (such as finding in the 1/3 of the face, use of skin color, and detection of chin and nostrils) were compared After the finding of the lip region, the lip and the outer borders of the lip were determined and the color of the skin was used in drawing the outer borders of the lip. Lip reading is controlled by 15 points. These are the 14 points representing the lip and whether the teeth are visible or not. From these 15 points, 3 different eigenvectors consisting of the values of 10, 11 and 19 points were formed. To monitor the lip movements, the differences between lip boundaries and thicknesses were evaluated with artificial neural networks and lip reading was performed, and the predicted accuracy rates were compared.

Yazar

Muhammed Halıcı

Bu Yayına Nasıl Atıf Yapılır

Muhammed Halıcı (Master Thesis). Lip reading with cascade object detection and artificial neural networks methods, 2019, Fırat University.

Anahtar Kelimeler

Lisans

Tüm Hakları Saklıdır

Bu eser belirtilen lisans koşulları altında paylaşılmaktadır.

Fırat University tezlerinden daha fazlası