Yüksek LisansAçık Erişim

A new method for computing the similarity of Turkish texts

2008
0 görüntülenme
0 i̇ndirme
Danışman: Prof. Dr. A. Coşkun Sönmez

Özet (EN)

To compute similarity rate of two texts, there are many methods. While some of these methods compute the similarity of texts by classical methods, some others can compute similarities more correctly and more closely to human brain by working more intelligently. The ones mentioned as the second are called Fuzzy String Similarity.Fuzzy string matching methods are generally improved through thinking English and texts in English. Thus, even they are succesfull for the texts in English, they usually may not be perform as the same in Turkish texts.Therefore, in this work a new similarity computing technic has been improved by modeling some cases oftenly encountered in matching Turkish texts and computing similarities. Especially, an intelligent method has been developed that perceives slips in spelling, arranges them, and computes similarity. Here, the concept mentioned as text may be both a word composed of a few letters and a long text composed of hundreds of words.In order to measure the success of the method improved, different computer users who have different specialities, at different levels, were asked to type different texts, and then these wrongly typed inputs were used. The similarity percent of the texts users typed false and their correct forms was computed by 3 different methods; The improved method, Edit Distance Similarity and Jaro-Winkler Similarity. So, their success was measured comparatively. Also, a software that finds repeated or similar records on the tables in any Oracle database by using 3 methods mentioned above was developed.It is hoped that this work may contribute to the common good of e-government workings with the aim of Turkish Natural Language Processing, finding similar records available in database systems, Turkish operating system, Turkish search engines, integration projects and providing integration between establishments.

Yazar

Bünyamin Dursun

Bu Yayına Nasıl Atıf Yapılır

Bünyamin Dursun (Master Thesis). A new method for computing the similarity of Turkish texts, 2008, Yıldız Technical University, Bilgisayar Mühendisliği Bölümü.

Anahtar Kelimeler

Lisans

Tüm Hakları Saklıdır

Bu eser belirtilen lisans koşulları altında paylaşılmaktadır.

Yıldız Technical University tezlerinden daha fazlası