Master'sOpen Access

Sign language recognition using spatio-temporal features on Kinect rgb video sequences and depth maps

2013
0 views
0 downloads
Advisor: Yrd. Doç. Dr. Songül Albayrak

Abstract (EN)

Human Computer Interaction (HCI), has become widespread in academical researches and so in daily life with the growing up of computer systems in the last quarter of 20th century and first years of 21th century rapidly. Especially problem-focused researches and approaches in machine learning, robotic, image processing and computer vision improve HCI research area and make it grow up. New problems and requirements in life, also affect and direct the HCI based applications and research areas. HCI mainly includes approaches and applications in controlling of computer softwares, computer operating systems, robots and control of devices that can be controlled near-fear or far-away. Speech and image signals have been used in order to control computer systems, interact with these kind of systems and these approaches are preferred considering the system requirements and problems. Recent advances in computer vision enable vision based systems to be widespread in HCI systems. Although vision based systems have advantages on system simplicity and ease of use, can have disadvantages like illumination effects in environment and overlapping situations. So problem-focused approaches are proposed in researches in order to eliminate the effects of these kind of problems. Vision based systems have been used for interaction in controlling electronical devices and computer softwares. Also components that used for interaction should be recognized efficiently for interaction accuracy. So machine learning approaches carry weight in this kind of studies for system decisions that are based on recognition and training processes. Not only electrical signals like speech can be used for interactions with computer systems but also visual signs which are carried out using hands, face, head and human body can be used. Especially systems that are controlled by hand signs, have increased significantly in recent years. These systems also works includes researches on recognition of sign languages which are used by deaf community for communication. This thesis aims to improve the recognition success rates for sign language recognition by analyzing the related and existing approaches. Sign languages are visual languages consist of hand, face and body motions and which are used by deaf community for communication with each other and others. Sign languages are native languages of deaf people and constitute a important percent of their communication. So, not only recognition and interpretation of these languages via computer systems are important in technological advances, but also have quite importance in social perspective. Sign languages consist of static signs (postures) and non-static signs (gestures). Proposed sign language recognition system is performed on videos of non-static (dynamic signs). In feature extraction process, a two stage spatio-temporal structure is performed. Firstly, temporal features of signs are extracted using accumulated motion image approach which is based on the intensity differences of sequential video frames and these features are presented in a single image. In second step, these features, which are in time domain, are transformed into spatial features via Discrete Cosine Transform (DCT). DCT provides coefficients that contain higher energy and feature vectors, which will be used for recognition, are obtained by selecting these coefficients in different ratios via zig-zag scanning. K-Nearist Neighbor (K-NN) classifier is employed for performance evaluation. System performance is evaluated on a dataset which contains 20 words belong to American Sign Language (ASL) and totally 800 sign videos and evaluated on a second dataset, which is collected in scope of thesis works, contains 111 words belong to Turkish Sign Language (TSL) and totally 1002 sign videos. In system recognition process, different test samples are selected via K-fold cross validation for performance analysis on datasets. Thus, it is aimed to obtain more valid success rates by using every sample in train and test sets both. Depth information of RGB-D video sequences are used effectively for improving the recognition rates of Turkish Sign Language dataset which have sign video samples captured by a Kinect sensor. Non-static signs are successfully recognized by proposed system that extract spatio-temporal features using sequential motion differences and transformation methods and that use K-NN classifier. Proposed system also has recognition rates between %95-99 on ASL dataset and %80-98 on TSL dataset.

Author

Abbas Memiş

How to Cite

Abbas Memiş (Master Thesis). Sign language recognition using spatio-temporal features on Kinect rgb video sequences and depth maps, 2013, Yıldız Technical University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Yıldız Technical University