A lexicon based method for subjectivity and sentiment analysis using an Arabic twitter corpus
Is this your thesis?
This record came from a bulk archive import. If it’s yours, link it to your profile.
Abstract (EN)
Sentiment analysis for social media is an interesting area of data mining for decision making in various domains. Therefore, continuous research is carried out in this area to cover the huge amount of data being pushed by users. Arabic is one of the ten important languages used in social media; therefore, interest in decision making anywhere needs knowledge about this. Twitter provides a platform for the exchange of opinions and ideas among users, leading decision making to building a knowledge base towards the development and planning of future outcomes. We present and illustrate how to obtain models with a high accuracy of classification by using the Lexicon-based approach. Our approach is implemented in three phases, beginning with preprocessing steps for Arabic words. The second phase discusses the extraction of more features relating to statistical and semantic orientations. We demonstrate how the extracted features (weight, score and negation) depend on two types of Arabic lexicon being clearly useful. Finally, the third phase applies a feature selection method with the Information Gain attribute evaluation and Ranker search method to find the features that have greater impact on the performance measures. We keep the features that have high rankings and remove those that have low rankings from the dataset. In the last two phases, we carry out our evaluations for all tasks using two machine-learning algorithms, namely K-Nearest Neighbor and Naïve Bayes. The accuracy for classification was found to have reached 93.56 with the Naïve Bayes classifier with a score feature, and this task determined which one of the two selected machine-learning models is more suitable for classifying the sentiment of Arabic tweets. Keywords: Arabic sentiment analysis, lexicon-based, feature extraction, feature selection, KNN, Naïve Bayes, Ranker, information gain attribute.
Author
Naseer Mohammed Jasım Al-buhruzı
Institution
How to Cite
Naseer Mohammed Jasım Al-buhruzı (Master Thesis). A lexicon based method for subjectivity and sentiment analysis using an Arabic twitter corpus, 2017, Çankaya University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Çankaya University
- Investigation of amazon and google for fault tolerance strategies in cloud computing services(2015)
- Exchange rate and inflation relationship: The case of Turkey(2023)
- Effects of the economic news on herd behavior(2023)
- Experimental analysis of effects of different network parameters on TCP / IP networks(2025)
- Reconstruction of patriarchy through matriarchy: A critique of gendered power structures in Naomi Alderman's The Power(2025)
- Characterization of under-hood airflow in construction equipment using experimental techniques(2025)