Master'sOpen Access

Improving sentiment analysis based deep learning by using feature selection

2021
0 views
0 downloads
Advisor: Doç. Dr. Fatih Özyurt

Abstract (EN)

In recent years, sentiment analysis has gained a great deal of research due to the dramatic growth in online social media use. For learning-based methods, many technically advanced strategies have been used to improve the performance of the forms. The sentiment analysis system uses natural language processing techniques and a sentimental vocabulary network, and one of the applications of deep learning in NLP is sentiment analysis. The most popular and successful type of RNN is the LSTM network. There is much research that uses the LSTM ability to analyze sentiment. However, large data volumes reduce the accuracy of LSTM network results in test data; in other words, over-fitting occurs. This problem occurs when there is a high correlation between independent variables. The model may not have high validity despite the high value of the correlation coefficient between the independent and dependent variables. In other words, although the model looks good, it does not have significant independent variables. Combining the LSTM network with feature selection methods can increase sentiment analysis accuracy to select effective features and solve them. In this study, we used the three benchmark datasets, namely (YELP, US Airline, and IMDB), then proposed a deep learning model (LSTM, Bi-LSTM, and GRU) to improve classification accuracy through using the feature selection method, namely, filter-based feature selection method Chi-square and Principle component analysis method PCA to select an optimal subset of features, and the performance of each measured and compared in terms of accuracy, precision, recall, and F1 score. After comparing the (LSTM, Bi-LSTM, and GRU) models with (Chi-2 and PCA) selected features, and also with the original feature set, the results show that feature selection methods significantly increase classification accuracy in all cases. . In the Yelp dataset, the maximum attained an accuracy of Bi-LSTM is 100% using chi-square. In the US Airline dataset, the maximum achieved accuracy of GRU-LSTM is 97.9% using chi-square. In the IMDB dataset, the maximum achieved accuracy of Bi-LSTM is 99.9% using chi-square.

Author

Dr. Mohammed Hussein Abdala

How to Cite

Mohammed Hussein Abdala (Master Thesis). Improving sentiment analysis based deep learning by using feature selection, 2021, Fırat University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Fırat University