Age and gender identification by SMS text messages
Is this your thesis?
This record came from a bulk archive import. If it’s yours, link it to your profile.
Abstract (EN)
Age and gender identification from text documents became a popular subject for researchers within the text classification field. Over the last decades, the number of text-based social network applications such as Facebook messenger, Twitter, and short message services, has increased at a rapid space. That is why texting has become the most popular method of communication that has users' attention all around the globe. This research aims to predict the age for 8 different age ranges and to identify the gender of a text sender from their short text messages. The reason behind this research is that some people fake their age and gender in text-based messaging applications. Linguistic psychology shows how certain words and writing styles of different people can be used to identify their age and gender. In recent decades, researchers used different sets of features for age and gender identification of an author. However, feature set identifications will always be a barrier for researchers. In this study, 25 different experiments were applied for the Naïve Bayes, Support Vector Machine (SVM), and J48 algorithms based on changing the parameter settings to prepare a feature for the identification of age according to different age classes and gender. The text that an author used was preprocessed in different stages. To design a module for SVM, Naïve Bayes, and J48, Weka (data mining software) was used. The reason behind using these three algorithms is that SVM is the most accurate and powerful algorithm used in text classification, Naïve Bayes is the fastest at building a module, and J48 has the ability to choose the most biased features, can classify data without complex calculations, and it has ability to handle incomplete or noisy data. However, it still took a long time to create a module. The Short Message Service (SMS) text messages used for the training and testing stages in this study can be found on kaggle.com under the name "The National University of Singapore SMS Corpus". The highest accuracy for age prediction was in experiment number six, which yielded 70.9823% by SVM, Later, after the whole dataset been used as a training set and one of the age classes been used as a testing set at a time. The highest result recorded for the age between 16 to 20 that included 20649 instances by using the same parameters, it was 91.3361% which recorded by SVM algorithm, while the highest record for gender identification was in experiment number three, also gained by SVM, which it was 79.5869% according to application of different parameter settings.
Author
Ahmad Jamal Khdr Khdr
How to Cite
Ahmad Jamal Khdr Khdr (Master Thesis). Age and gender identification by SMS text messages, 2018, Fırat University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Fırat University
- Using social media as an integrated marketing communication tool(2018)
- Foundation of Dutch East İndia Company and her rising in İndonesia in the 17th century(2013)
- Color usage at Turkish Divan of Fuzûlî(2013)
- Transcript and evaluation of Number 317 Midyat Şer'iyye Record (Hicri 1328-1334 / G.C. 1910-1916)(2018)
- Examination of stress state between Doğanyol (Malatya) and Çelikhan (Adıyaman) on the east Anatolian fault zone(2020)
- Yavuzeli (Gaziantep) surrounding volcanic outcropping of rocks petrographic and geochemical features(2014)
