Master'sOpen Access

Detection of similarities with deep learning methods in Turkish texts

2019
0 views
0 downloads
Advisor: Prof. Dr. Ahmet Bedri Özer

Abstract (EN)

In today's technologies, the perception of natural languages by machines is an inevitable need for solving many problems. Thanks to the studies carried out on texts, software based solutions are used in many areas such as plagiarism detection, text-author matching, detection of text subject, and text summarization. It is expected that the texts will be interpreted and processed by computers in the exemplified works and many similar fields. In order to interpret the texts, it is expected that the structural and linguistic features of the language used should be comprehended by computers. There are various difficulties and problems to be solved in this process. One of these problems is the ability of computers to comprehend the texts semantically. If the computer is able to make sense from the text, it will provide a great solution to the aforementioned problems. In addition, it will be possible to offer software-based alternatives to problems requiring human perception in the solution by increasing user interaction to a very high level. Measuring the similarity of texts has been the subject of various studies in the literature as a problem addressed within this framework. In the past years, it has been possible to compare the texts only structurally, but in the recent years methods have been developed to detect various semantic similarities. A current method used to make semantic inferences from texts is deep learning based word representation method. This method makes it possible to determine the semantic affinity of words. In this study, as a result of examining Turkish texts both semantically and structurally, it is aimed to measure similarities with a common approach. The structural similarity was measured by Cosine Similarity and the semantic similarity was used by Word2Vec model. As a result of the thesis, a method which provides the common use of these two different approaches is proposed. In this context, experimental tests were conducted on both subject-oriented (Information Security) and general-language texts. The results obtained are shared in the findings and conclusions section which prove the success of the proposed method.

Author

İrfan Aygün

How to Cite

İrfan Aygün (Master Thesis). Detection of similarities with deep learning methods in Turkish texts, 2019, Fırat University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Fırat University