A new method for computing the similarity of Turkish texts
2008
0 views
0 downloads
Advisor: Prof. Dr. A. Coşkun Sönmez
Abstract (EN)
To compute similarity rate of two texts, there are many methods. While some of these methods compute the similarity of texts by classical methods, some others can compute similarities more correctly and more closely to human brain by working more intelligently. The ones mentioned as the second are called Fuzzy String Similarity.Fuzzy string matching methods are generally improved through thinking English and texts in English. Thus, even they are succesfull for the texts in English, they usually may not be perform as the same in Turkish texts.Therefore, in this work a new similarity computing technic has been improved by modeling some cases oftenly encountered in matching Turkish texts and computing similarities. Especially, an intelligent method has been developed that perceives slips in spelling, arranges them, and computes similarity. Here, the concept mentioned as text may be both a word composed of a few letters and a long text composed of hundreds of words.In order to measure the success of the method improved, different computer users who have different specialities, at different levels, were asked to type different texts, and then these wrongly typed inputs were used. The similarity percent of the texts users typed false and their correct forms was computed by 3 different methods; The improved method, Edit Distance Similarity and Jaro-Winkler Similarity. So, their success was measured comparatively. Also, a software that finds repeated or similar records on the tables in any Oracle database by using 3 methods mentioned above was developed.It is hoped that this work may contribute to the common good of e-government workings with the aim of Turkish Natural Language Processing, finding similar records available in database systems, Turkish operating system, Turkish search engines, integration projects and providing integration between establishments.
Author
Bünyamin Dursun
Institution
How to Cite
Bünyamin Dursun (Master Thesis). A new method for computing the similarity of Turkish texts, 2008, Yıldız Technical University, Bilgisayar Mühendisliği Bölümü.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Yıldız Technical University
- Examining ?Historical housing structures" within the confines of protecting ecological balance(2012)
- Approximate solutions of integral equations(2012)
- The annotative dictionary of Kutadgu Bilig in terms of vocabulary(2013)
- Stepper motor speed control with labVIEW(2014)
- Determining supply chain risk factors in food industry(2014)
- TiO2/Cu2O ince film fotovoltaik hücrelerin karakterizasyonu(2014)
