Application of text analysis processing using deep learning algorithms in big data
2019
0 views
0 downloads
Advisor: Prof. Dr. Ali Karcı
Abstract (EN)
With the high-speed developments in the IT world and the widespread use of the Internet, the diversity and amount of data produced on digital platforms has increased. The majority of this big data generated is textual content. However, it has become a difficult problem to process the increasing text data with traditional methods. For this reason, deep neural networks and neural network-based word embedding methods have been developed that perform highly successfully on big data technologies and especially big data. In this thesis, detailed analysis has been made on deep learning architectures used word embedding methods with big data technologies. When the studies were examined, it was seen that there were many natural language specific studies, especially English, but the number of Turkish studies was not sufficient. Therefore, Turkish was chosen as the target language of the study. However, three applications were developed in the thesis and two novel methods were proposed. In the first application, a big data application was made to determine the platform in which the studies would be conducted. In the second application, preprocessing studies were performed before text processing. In this context, the stopwords list for Turkish was generated for the first time by TF (Term Frequency) - IDF (Inverse Document Frequency) method. In the third application, a dataset (Dataset-1) consisting of very large Turkish unlabeled data has been generated. Word vectors were trained on this dataset using word embedding methods and the performances of different word embedding methods were compared. For the third application, a second Turkish dataset (Dataset-2) consisting of approximately 1,5 million data and 10 classes were generated. A method has been proposed on this data set where word vectors are used for the problem of text classification on different deep learning architectures with the transfer learning method as pre-trained word vectors. With this proposed method, current performance values on almost all models have been improved between 5-7%. As a second method, a new method called the dictionary method has been proposed. Since there is no spelling checker developed for Turkish, the misspelled words on Dataset-2 have been identified and LSTM (Long Short Term Memory), which is a deep learning model, has tried to identify the correct words instead. When the classification performance obtained as a result of the analysis was analyzed, it was seen that approximately 55.000 incorrect words were replaced with the correct words and the performance value was improved by 8.68%. With this thesis, two large Turkish datasets were generated in order to contribute to Turkish text processing. In addition, the largest Turkish word vectors ever trained on these datasets were generated and shared open to researchers.
Author
Dr. Murat Aydoğan
Institution
How to Cite
Murat Aydoğan (Doctorate thesis). Application of text analysis processing using deep learning algorithms in big data, 2019, İnönü University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from İnönü University
- Regional threats and opportunities to turkey's national economic security(2022)
- The effect of motivational interviews for primiparous pregnant women with low normal birth belief on medical and natural birth belief(2022)
- Knowledge, opinions and applications of pediatric nurses towards therapeutic games(2017)
- Martha Nussbaum's conception of justice: An inquiry on capability approach(2021)
- The effects of systemic pistacia eurycarpa yalt administration on alveolar bone loss and oxidative stress in rats with experimental periodontitis(2021)
- An examination of Dellâlzade İsmail Efendi's works found in Dârül-Elhan archive in terms of authority and navigation(2022)
