Master'sOpen Access

Doğal dil işleme tekniklerinin incelenmesi ve seyahat-turizm sesli asistanı için Türkçe varlık ismi tanıma aracı geliştirilmesi

2020
0 views
0 downloads
Advisor: Doç. Dr. Ümit Deniz Uluşar

Abstract (EN)

The aim of this thesis is to conduct studies on development of a Turkish language mobile travel assistant in the travel industry using natural language processing techniques. While conducting this research, it is aimed to investigate the natural language processing done so far in the literature in languages other than Turkish and to measure the adaptability of these studies to Turkish. Turkish natural language processing studies provide resources for both our own language and other Turkish languages. Examining the success of academic studies in the field of Turkish natural language processing, determining how their accuracy rates can be increased and how the results found can be applied to practical applications will be done. Natural Language Processing or NLP is defined as the system of processing human language and understand the intention. It is a sub-field of artificial intelligence and focuses on topics such as speech recognition, morphological processing, disambiguation, dependency parsing and named entity recognition of the language. Although being researched under NLP, these topics are comprehensive and comprises of rich research all by themselves. The term of computers 'understanding' the language is not like the way human understand it. The computers should be provided with models using supervised, semi supervised and unsupervised approaches and data to use those given models to understand the language. Here, the success rates of understanding can be defined as how close a computer understands the intent comparing to a human speaking the language. Using the state-of-the-art models on NLP, the success rates of computers understanding the language increased with the increase of processing power of available computers and data. NLP can be used on areas such as sentiment analysis, machine translation, automatic speech recognition systems on various domains. In this article, I will present the academic studies focusing on the state-of-the-art NLP techniques and compare the results achieved by using those techniques on a Turkish Mobile Travel Assistant. When the literature is reviewed, we see that there are many studies in the field of Turkish NLP. In this context, methods such as recognizing the waves of the sound and transferring them to the text in the computer, text vocalization, morphological analysis-production, syntax analysis, semantic analysis are used and developed. Applications such as checking / correcting spelling errors, machine translation, information extraction, information retrieval, development of question-and-answer systems, summary extraction are collected under the definition of natural language processing. Voice assistant studies have been conducted before in Turkish. However, it was not widely used and remained in the research phase. In order to carry out studies in the field of NLP, machine learning methods are commonly used. In order to use machine learning techniques in this area, big data is needed in the subject. The data is not available due to the fact that the studies have remained at the research level or the data has been privatized by privately held companies. As the subject become more specific like travel sector, it becomes more difficult to access the data. Within the scope of this thesis, it is aimed to collect the voice intention in the field of travel and convert the voice into data using various NLP techniques and to provide this data as an input to machine learning methods. Although there are studies in the area of voice assistant, no studies have been carried out in the field of travel until today. Classification methods and machine learning models will be used during the development of an NLP Travel assistant. In this work, The Named Entity Recognizer (NER) tool utilizes transfer learning methods, namely Google BERT and ELECTRA. These models will be fine-tuned with 4 types of data. Those fine-tuned models will be evaluated with a unique test data that is common to all those data sets. The extended travel NER tool will be able to tag PERSON, LOCATION, ORGANIZATION, DATE, TIME entities. KEYWORDS: Machine Learning, Turkish Speech Recognition, Turkish Natural Language Processing, Turkish Natural Language Understanding, Turkish Morphological Analysis, Turkish Named Entity Recognition

Author

Dr. Deniz Gül Özcan

How to Cite

Deniz Gül Özcan (Master Thesis). Doğal dil işleme tekniklerinin incelenmesi ve seyahat-turizm sesli asistanı için Türkçe varlık ismi tanıma aracı geliştirilmesi, 2020, Akdeniz University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Akdeniz University