A study on sentiment analysis on social media using machine learning techiques
Is this your thesis?
This record came from a bulk archive import. If it’s yours, link it to your profile.
Abstract (EN)
In recent years, the use of machine learning techniques to analyze texts written by humans is attracting significant attention, according to the wide availability of these texts and their ease of access. As these texts are written by humans, the extraction of accurate knowledge requires intensive processing, known as Natural Language Processing (NLP). The main challenge that these techniques face is the enormous amount of information available in these texts and the complex relations among the features, i.e. words, and the knowledge required to be extracted. Accordingly, eliminating the words that has negative or no influence on the knowledge extraction can significantly improve the performance of NLP techniques, by reducing dimensionality and improving the efficiency of knowledge representation. In this study, we propose a new feature selection technique that uses vectors that represent the sentimental meaning of words and knowledge extracted about the influence of words on the performance of the classifiers. The proposed method uses an artificial neural network that is trained using reinforcement learning by monitoring the influence of removing each word in the training dataset. Word embedding is used to provide the vectors that represent these words, so that, even if a word is not included in the training, its rank can be predicted by the proposed method depending on the values of the vector generated for it and the knowledge about the most similar words that are considered during the training. Accordingly, no complex statistical computations are required for each word in the corpus, as well as any new words that can be added to the corpus in the future. The evaluation of the proposed method has shown its ability to predict the rank of each word that is not included in the training with 94.61% accuracy. Moreover, feature selection based on these ranks has been able to improve the performance of different types of classifiers, such as the Support Vector Machine (SVM), Naïve Bayes (NB) and Random Forest (RF), which use count vectors to represent the text, as well as the Convolutional Neural Network (CNN), Long- Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) classifiers, which rely on the word embedding vectors for the classification. Moreover, the GRU classifier has been able to achieve the highest performance, with 95.54% accuracy, compared to the other classifiers and state-of-the-art methods in the literature.
Author
Mohamed Guma Ibrahım Bodea
Institution
How to Cite
Mohamed Guma Ibrahım Bodea (Master Thesis). A study on sentiment analysis on social media using machine learning techiques, 2020, Kastamonu University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Kastamonu University
- Investigation of pre-school teacher candidates' mental models for the concept of biodiversity(2023)
- Examining the relationship between religious attitude, self-regulation, free will, and determination(2023)
- Mental models of day and night concepts of primary school fourth grade students(2022)
- Compliance with standards and the contribution of certification to brand value: A research in Kastamonu building materials sector(2023)
- Exegesis of Ahzâb 37. verse(2023)
- The effects of the Presidential Government System on the organizational structure of public administration(2023)
