DoktoraAçık Erişim

Rule-based natural language processing methods for Turkish

Bu tez size mi ait?

Bu kayıt toplu arşivden geldi. Sizinse profilinize bağlayın.

2010
0 görüntülenme
0 i̇ndirme

Özet (EN)

In order to determine morphological properties of a language, a corpus which represents that language should be created. Many large scale corpora generated and have been used for Natural Language Processing (NLP) applications on many languages, such as English, German, Czech, etc, but any large scale Turkish corpora have not be generated yet.In this study, natural language processing methods for Turkish were developed by using rule-based approach, and also an infrastructure, Rule-Based Automatical Corpus Generation (RB-CorGen), to use the new developed methods was implemented. For testing RB-CorGen on Turkish, the roots, stems and suffixes were obtained from Turkish Linguistic Association (Türk Dil Kurumu, TDK) and Dokuz Eylul University, College of Literature Linguistic Department, the defined tags and grammatical rules were stored in XML formatted file, and documents, include nearly 95 million wordforms, were collected from five Turkish newspapers in electronic environment. The average success rates of Rule-Based Sentence Boundary Detection (RB-SBD) and Rule-Based POS Tagging (RB-POST) methods were determined as 99.66% and 92% respectively. It was seen that the success rate of RB-CorGen increases with the increasing number of rules.

Yazar

Özlem Aktaş

Bu Yayına Nasıl Atıf Yapılır

Özlem Aktaş (Doctorate thesis). Rule-based natural language processing methods for Turkish, 2010, Dokuz Eylül University.

Anahtar Kelimeler

Lisans

Tüm Hakları Saklıdır

Bu eser belirtilen lisans koşulları altında paylaşılmaktadır.

Dokuz Eylül University tezlerinden daha fazlası