Master'sOpen Access

Sentiment analysis of social network data using machine learning

2017
0 views
0 downloads
Advisor: Prof. Dr. Galip Aydın

Abstract (EN)

Opinion mining have shown to be a source of great value when it comes to collecting people's feedback about a product or an event as an example. The World Wide Web recently became very rich of people's opinions and thoughts so instead of filling some surveys for gathering feedback, due to the advance of sentiment analysis, we can now harvest thousands, millions or thousands of millions of opinions about what we are trying to develop, or thoughts about what we did recently. By the improvements that has occurred in the machine learning field we can now build a model that does some mathematical equations on large sets of data to give us the thoughts of millions of people about our work. With the many microblogging services online nowadays, these websites tend to have huge amounts of valuable information such as reviews, feedback, opinions and experience with specific products or events that can be benefited and gathered from by machines that work according to some machine learning model or algorithm. This is where sentiment analysis comes to existence, it is the field of Natural Language Processing (NLP) that uses machine learning models and techniques to collect subjective information from real world text (corpuses). In this research we aim to benefit from the opinions-rich social media network data by extracting the sentiments from them. We will use multiple sentiment analysis techniques (Paragraph Vector, Deep Learning for Java and Naïve Bayes-SVM) on multiple platforms (DL4J, Gensim and regular Machine Learning) all that on two different tasks; IMDB Movie Review Dataset and 1.6 million tweets, and change parameters (like epochs, text manipulation etc.) to see what affects the results of those models, thus ending up with a better model. We also compare the results of all models on different datasets in the Result chapter to find out which model scores higher accuracy in prediction plus the computation power and time needed for each of them. We try to use a distributed computing platform, Apache Spark, to achieve parallel and distributed computing power which really comes in handy when using repetitive calculation on big amounts of records (Big Data). Keywords: Big Data, Deep Learning, Sentiment Analysis, Spark, Paragraph Vector

Author

Dr. Alı Abas Alo Albabawat

How to Cite

Alı Abas Alo Albabawat (Master Thesis). Sentiment analysis of social network data using machine learning, 2017, Fırat University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Fırat University