Detection of similarities with deep learning methods in Turkish texts
2019
0 views
0 downloads
Advisor: Prof. Dr. Ahmet Bedri Özer
Abstract (EN)
In today's technologies, the perception of natural languages by machines is an inevitable need for solving many problems. Thanks to the studies carried out on texts, software based solutions are used in many areas such as plagiarism detection, text-author matching, detection of text subject, and text summarization. It is expected that the texts will be interpreted and processed by computers in the exemplified works and many similar fields. In order to interpret the texts, it is expected that the structural and linguistic features of the language used should be comprehended by computers. There are various difficulties and problems to be solved in this process. One of these problems is the ability of computers to comprehend the texts semantically. If the computer is able to make sense from the text, it will provide a great solution to the aforementioned problems. In addition, it will be possible to offer software-based alternatives to problems requiring human perception in the solution by increasing user interaction to a very high level. Measuring the similarity of texts has been the subject of various studies in the literature as a problem addressed within this framework. In the past years, it has been possible to compare the texts only structurally, but in the recent years methods have been developed to detect various semantic similarities. A current method used to make semantic inferences from texts is deep learning based word representation method. This method makes it possible to determine the semantic affinity of words. In this study, as a result of examining Turkish texts both semantically and structurally, it is aimed to measure similarities with a common approach. The structural similarity was measured by Cosine Similarity and the semantic similarity was used by Word2Vec model. As a result of the thesis, a method which provides the common use of these two different approaches is proposed. In this context, experimental tests were conducted on both subject-oriented (Information Security) and general-language texts. The results obtained are shared in the findings and conclusions section which prove the success of the proposed method.
Author
İrfan Aygün
How to Cite
İrfan Aygün (Master Thesis). Detection of similarities with deep learning methods in Turkish texts, 2019, Fırat University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Fırat University
- Using social media as an integrated marketing communication tool(2018)
- Foundation of Dutch East İndia Company and her rising in İndonesia in the 17th century(2013)
- Examination of stress state between Doğanyol (Malatya) and Çelikhan (Adıyaman) on the east Anatolian fault zone(2020)
- Color usage at Turkish Divan of Fuzûlî(2013)
- Yavuzeli (Gaziantep) surrounding volcanic outcropping of rocks petrographic and geochemical features(2014)
- Hizbu?t-Tahrir and the religions and political thoughts of Ercumend Özkan(2008)
