Video-based isolated sign language recognition using convolutional neural networks
2024
0 views
0 downloads
Advisor: Doç. Dr. Ömer Kaan Baykan
Abstract (EN)
Sign language is a fundamental communication tool for millions of hearing-impaired individuals worldwide. However, understanding and using sign language is not a common skill among hearing individuals, which increases the risk of social isolation for the hearing-impaired. This thesis addresses the current limitations in word-based Sign Language Recognition (SLR) technologies, aiming to enhance the accuracy and generalizability of detection in this field. In this context, through three main studies, both the manual and non-manual elements of sign language are comprehensively analyzed, presenting deep learning-based systems. In the first study, the R3(2+1)D-SLR network, which combines the advantages of R3D and R(2+1)D convolutional blocks, was proposed. This network effectively extracts spatial and temporal features, providing high accuracy and robustness in sign language recognition. The sign language recognition system developed based on the R3(2+1)D-SLR architecture integrates data obtained from the signer's body, hands, and face, and classifies it using Support Vector Machines (SVM). The proposed system demonstrates significant improvements in accuracy and robustness against background variability by using visual pose data instead of RGB data. This system achieved test accuracies of 94,52% and 98,53% in signer-independent evaluations on the BosphorusSign22k-general and LSA64 datasets, respectively. The second study presents an innovative approach for the task of isolated SLR, focusing on integrating pose data with derived Motion History Images (MHI). The research combines spatial information obtained from body, hand, and face poses with three-channel MHI data that reflect the temporal dynamics of the sign. Notably, the developed finger-pose-based MHI feature captures the nuances of finger movements and gestures more successfully than current approaches in SLR. This feature enhances the system's accuracy and reliability by more accurately processing the rich details of sign language. Additionally, the use of linear interpolation to complete missing pose data has improved overall model performance. The combination of features obtained from the ResNet-18 model enhanced with Randomized Leaky Rectified Linear Units (RReLU) and classification through SVM has successfully addressed the interaction between manual and non-manual features. This integrated method has demonstrated competitive and superior results compared to existing methodologies, achieving accuracies of 96,94%, 94,87%, 98,68%, and 95,14% in experiments conducted on the BosphorusSign22k-general, BosphorusSign22k, LSA64, and GSL datasets, respectively. The third study introduces an innovative multi-channel approach focusing on the characteristics and configurations of fingers in sign language recognition. Based on visual finger pose data processed in separate channels, this approach is designed to provide detailed analysis of finger movements. The proposed Multi-Channel MobileNetV2 model utilizes multi-channel data on fingers to offer high accuracy and precision in the sign language recognition process. Additionally, the study incorporates non-manual features of sign language by processing body and face information derived from pose data. The proposed system has achieved notable accuracy rates of 97,15%, 95,13%, 98,93%, and 95,37% on the BosphorusSign22k-general, BosphorusSign22k, LSA64, and GSL datasets, respectively. These results highlight the generalizability and adaptability of the proposed method, proving its competitive superiority over existing studies in the sign language recognition literature. This thesis indicates that innovative approaches in sign language recognition technology have the potential to reduce communication barriers by more accurately capturing the richness and subtle details of sign language. All three studies have achieved high accuracy rates across different datasets, enhancing the effectiveness and reliability of SLR systems in practical applications.
Author
Dr. Ali Akdağ
How to Cite
Ali Akdağ (Doctorate thesis). Video-based isolated sign language recognition using convolutional neural networks, 2024, Konya Technical University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Konya Technical University
- Numerical and experimental in vestigation of optimization of Pelton turbine rotor design parameters in micro turbine size(2018)
- Comparison of some manufacturing costs according to various analysis parameters and other regulations of reinforced concrete structures with different floor systems(2018)
- The use of silica fume in self-compacting concretes affects the concrete compressive strength and adherence(2018)
- Load-bearing carrier system properties in the historical buildings repair and strengthening techniques for damages model analysis of Zenburi masjid(2018)
- Lateral rigidity improvement of deficient reinforced concrete structures with the use of user friendly systems(2018)
- Application of artificial intelligence methods to estimate monthly pan evaporation using meteorological data(2018)
