Development of a question-answer system with natural language processing techniques on Turkish texts
Is this your thesis?
This record came from a bulk archive import. If it’s yours, link it to your profile.
Abstract (EN)
Question-answer systems facilitate information access processes by providing fast and accurate answers to the questions expressed by users in natural language. Today, advances in natural language processing techniques increase the effectiveness of such systems and improve the user experience. However, in order for these systems to work effectively, the structural features of the language must be understood correctly. Traditional rule-based and knowledge retrieval-based systems are not able to analyse the contextual meaning of questions and texts deeply enough and therefore cannot produce satisfactory answers to complex questions. For this reason, Transformer-based models that can better capture the contextual and semantic integrity of the language have been developed. In this thesis, Transformer-based pre-trained language models such as BERTurk, ELECTRA and DistilBERTurk are used to develop a high-performance question-answer system for Turkish. It has been observed that there are many natural language-specific studies in the literature, including English, but the number of studies for Turkish is limited. Therefore, Turkish was preferred as the target language of the study. TQuAD (Turkish Question Answering Dataset) on Turkish & Islamic History of Science, THQuAD (Turkish Historic Question Answering Dataset) on Turkish & Islamic History of Science and Ottoman History, and BQuAD (Biology Question Answering Dataset) consisting of topics in the biology course were used in the experiments. For each dataset, experiments were performed with the same hyperparameters and the results were compared with exact match and F1 score performance metrics. It was observed that the models with case sensitivity obtained higher exact match and F1 scores. In addition, a more comprehensive dataset was created by combining THQuAD and BQuAD, and as a result of the experiments performed on this dataset with various hyperparameters, the best performance with 63.99 exact matches and 80.84 F1 scores was obtained in the BERTurk (Cased, 128k) model. This fine-tuned model was used in the design of the question-answer application. With this thesis, 1 question-answer model was created and shared openly for use in order to contribute to Turkish natural language processing studies.
Author
Mehmet Arzu
Institution
How to Cite
Mehmet Arzu (Master Thesis). Development of a question-answer system with natural language processing techniques on Turkish texts, 2024, Fırat University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Fırat University
- Using social media as an integrated marketing communication tool(2018)
- Foundation of Dutch East İndia Company and her rising in İndonesia in the 17th century(2013)
- Color usage at Turkish Divan of Fuzûlî(2013)
- Transcript and evaluation of Number 317 Midyat Şer'iyye Record (Hicri 1328-1334 / G.C. 1910-1916)(2018)
- Examination of stress state between Doğanyol (Malatya) and Çelikhan (Adıyaman) on the east Anatolian fault zone(2020)
- Yavuzeli (Gaziantep) surrounding volcanic outcropping of rocks petrographic and geochemical features(2014)