DoctorateOpen Access

A sentence boundary detection system for Turkish textual documents

2022
0 views
0 downloads
Advisor: Prof. Dr. Selma Ayşe Özel ; Prof. Dr. Sera Yeşim Aksan

Abstract (EN)

This study aims to design a high-performance method for sentence boundary detection in accordance with its own language structure in order to increase the prevalence of Turkish in the digital world. When general-purpose sentence boundary detection applications were tested on corpora containing different usage patterns of Turkish, it was observed that these applications remained at low success rates. In the study within the scope of the thesis, a hybrid method is proposed by combining the FP-Growth method, which is used to generate association rules, and Recurrent Neural Network (RNN) methods. A set of features has been obtained, in which suffix and POS tag information are used together, as well as some formal rules for the tokens before and after the possible end-of-sentence character. Then, this data set was tested with RNN methods such as LSTM and BiLSTM in the Turkish National Corpus (TNC) sub-corpus. In general-purpose applications, it was observed that the f-score value, which was 77% in the tests performed with the same data set, reached 96% with the proposed method. Within the scope of the study, it was observed that the success reached similar values in different Turkish corpora. In addition, in the study, it was concluded that the POS tag information and suffix structures of the tokens for Turkish play an important role in the detection of the boundary of a sentence. Keywords: Sentence Boundary Detection, Recurrent Neural Network, FP-Growth, Turkish Natural Language Processing, Corpus Annotation, Grammatical Cues

Author

Dr. Yasin Bektaş

How to Cite

Yasin Bektaş (Doctorate thesis). A sentence boundary detection system for Turkish textual documents, 2022, Çukurova University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Çukurova University