Master'sOpen Access

Direction estimation for stock values by using classification algorithms on twitter data

2019
0 views
0 downloads
Advisor: Doç. Dr. Barış Koçer

Abstract (EN)

The stock market has always been the favorite investment instrument of investors with the advantages it provides. Some of them are easy buying and selling of the shares, size of proceed and the easy access to the data. In order to generate more revenue on this platform, investors were also constantly interested in the forward-looking direction forecasting of stocks. Many different techniques have been developed for this reason. In this study, we made the classification of twitter messages -which is one of the most widely used social media sharing platforms- in order to predict the increase or decrease of the demand for the stock, because of we considered that the main criterion of pricing of stock is "supply-demand" relationship. In our research, we analyzed the tweets for stocks of global companies such as Apple, Facebook, General Electric, General Motors, The Coca-Cola Company, McDonalds, Microsoft, Netflix, Pfizer Corporation, Tesla Motors, which are traded on the American Dow Jones (DJIA) stock exchange. We performed classification, using naive bayes, random forest, support vector machine, decision tree, k-nearest neighbor and artificial neural network classification algorithms and compared the performance results of these algorithms Bi-monthly data for the related stocks -covering April 2019 - May 2019- were obtained using the twitter web interface. Similarly, files containing the stock values used for determination of success level of direction estimation and testing purposes were also obtained from the www.eoddata.com web address. The labeling process was performed by 75 different participants which have various stock market knowledge, by reading the tweets one by one and marking them manually as positive, negative and neutral. In addition to senseless tweets, the tweets consisting of a link, consisting of an advertisement, vague tweets etc. were labeled as neutral. Because of containing a lot of rubbish tweets, the neutral class was ignored when calculating the success of direction estimation of the stock values. But they were not ignored the measuring the success of the classification algorithms. For the classification, firstly we cleaned the tweets by removing the inessential factors as punctuation marks, hyperlinks and web addresses, "tab" characters, hashtags, re-tweets, repeating similar tweets etc. We classified using machine learning techniques for classification on this clear dataset. Using the TF-IDF method, we digitized by calculating the frequency weights of all words in each data set for each tivit and converted it into a vector. We obtained the performance results by using 6 different classification algorithms on these digitized data. While we used the classification algorithms which we determined for classification process with machine learning method to estimate the class of tweets, we divided the data sets into 10 parts with the corss validation method in order to make it homogenous. We used each of these parts in order for verification purposes, and used the other parts in the training of the system. Finally, when all the parts were finished, we achieved the overall prediction success by averaging 10 tracks. At the end of the research, random forest algorithm gave the most successful result with 77.37% and the support vector machine gave the worst result with 61.41%. Also the best predictive success for the stocks we tried to predict the class belongs to GM with 83.3%, while the worst predicted success belongs to GE with 62.15%. In the evaluation of success results of the direction estimation of the stocks, the most successful prediction was made for KO with 96.5%, and the worst estimation was made for TSLA with 66.7%. In the calculation of these predictive successes, only positive and negative tagged tweets were taken into account, while neutral tweets were ignored. A positive tweet was considered as a asuccessful estimate unless the stock made a negative move the following day, and a tivit marked as negative was considered as successfully if the corresponding stock did not rise the following day. As a result of the findings, it was seen that twitter data can be used to make direction estimations of stocks and very successful results can be obtained.

Author

Dr. Mustafa Vehbi Türkalp

How to Cite

Mustafa Vehbi Türkalp (Master Thesis). Direction estimation for stock values by using classification algorithms on twitter data, 2019, Konya Technical University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Konya Technical University