Analysis of big data including terror terms with machine learning techniques
Is this your thesis?
This record came from a bulk archive import. If it’s yours, link it to your profile.
2019
0 views
0 downloads
Advisor: Dr. Öğr. Üyesi Mustafa Ulaş
Abstract (EN)
Thanks to technology and the internet, data is produced instantly. With the reproduction of this data, information extraction processes become important and the information to be extracted from the data can be very valuable. With the development of technology and the proliferation of communication networks, the size and value of information has come to the forefront today. In today's technology, many devices can produce instant data, record the data, and as these records increase, the data is considered insignificant and sent to the data store. By carrying out research and development efforts in this direction, many institutions and organizations realized that the rich resources underlie the complex and varied data have emerged as a structure called Big Data. It is seen that meaningful, useful and important data emerge thanks to large data processing methods. Big Data means not only large amounts of data, but data that cannot be processed by conventional methods. In this study, the Global Terrorism Database (GTD) data set, which includes the terrorist incidents that took place in the world where data obtained from international news agencies and sources, was discussed. Machine learning methods and analysis and classification operations were performed on this data set within the scope of big data. An application has been developed which predicts the organization of a terrorist incident. Within this framework, the Apache Spark-based open source big data processing tool, which uses the information about the type of weapon, type of weapon, country, region and target group in the attack, was developed using Python language. Six different machine learning algorithms have been applied by selecting the first 10 terrorist organizations performing the most attacks from the GTD dataset and comparisons and evaluations have been made between the performances of the algorithms. The highest accuracy rate in the applied algorithms is 98,2 % with K-Nearest Neighbor (KNN) algorithm. The logistic regression (LR) algorithm was specified according to the situation appropriate for the big data set. Keywords: Apache spark, Big data, Terrorist incidents, Machine Learning Algorithms
Author
Barış Karabay
How to Cite
Barış Karabay (Master Thesis). Analysis of big data including terror terms with machine learning techniques, 2019, Fırat University.
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Fırat University
- Using social media as an integrated marketing communication tool(2018)
- Foundation of Dutch East İndia Company and her rising in İndonesia in the 17th century(2013)
- Examination of stress state between Doğanyol (Malatya) and Çelikhan (Adıyaman) on the east Anatolian fault zone(2020)
- Color usage at Turkish Divan of Fuzûlî(2013)
- Yavuzeli (Gaziantep) surrounding volcanic outcropping of rocks petrographic and geochemical features(2014)
- Hizbu?t-Tahrir and the religions and political thoughts of Ercumend Özkan(2008)
