Türkçe cümlelerde fonetik işaretlerin düzeltilmesi
2021
0 views
0 downloads
Advisor: Dr. Öğr. Üyesi İsmail Burak Parlak ; Prof. Tankut Acarman ; Dr. Öğr. Üyesi Cemal Okan Şakar
Abstract (EN)
This project focuses on correcting phonetic letters in Turkish. With the increase in the use of virtual keyboards on mobile and desktop devices, the Turkish word error rate has increased significantly due to the absence of Turkish letters or typos. Phonetic letter corrections are of great importance in order to remove this problem in the data that serves many text-based systems and to build a clean database. Because these letters are used extensively in languages such as Romanian, Arabic and Vietnamese, including Turkish, and the presence of these letters also reveals a great difference in terms of meaning. There are similar studies on this subject in Turkish, but one of the most important points of this study is the addition of the letters "â" and "î" to the set of diacritic letters. In this study, in order to overcome the mentioned problem, first of all, a raw dataset was prepared with approximately 5000 books on Turkish, both containing a wide variety of words and written in accordance with the spelling rules. This raw dataset was made ready for review and training the system to be created by going through several cleaning steps. First, N-Gram models were created with these datasets. Zipf and Mandelbrot distributions were observed on the generated models, and perplexity values were calculated to evaluate how well a language model was designed. After the high compatibility values obtained here, these data were used in the next learning and evaluation steps. In the training and evaluation of the model, which is the last step of the project, letter-based system training was carried out by using the "seq2seq" based learning model. A new artificial input set was created by randomly changing the Latin equivalents of the diacritic letters and phonetic letters in the words on the control sets separated from the main data set. The trained model was evaluated with this control set and success rates over 90% were obtained.
Author
Dr. Hüseyin Ekici
Institution
How to Cite
Hüseyin Ekici (Master Thesis). Türkçe cümlelerde fonetik işaretlerin düzeltilmesi, 2021, Galatasaray University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Galatasaray University
- International state responsibility arising from new space activities(2025)
- The liability of shareholders and organs for public debts in capital companies(2022)
- Karşı kültürel bir kimlik olarak taraftarlık: istanbul futbol tribünlerinde kimliksel yapılanış biçimleri çalışması(2014)
- Yeni roman: claude simon ve william faulkner(2014)
- Directors and officers liability insurance(2015)
- Langlands fonktörsellik ilkesi(2021)
