Master'sOpen Access

Büyük dil modellerinin Türkçe veri kümeleri üzerinde kapsamlı ve çok-istemli değerlendirmesi

2025
0 views
0 downloads
Advisor: Dr. Öğr. Üyesi Emre Uğur ; Prof. Ayşe Başar

Abstract (EN)

This study aims to comprehensively evaluate large language models (LLMs) on both understanding and trustworthiness tasks in Turkish language. In this study, we also conducted analyses on the prompt robustness of the models and the comparison of fine-tuned LLMs with their chat versions. Turkish is a resourced language, but it is an under-researched language. This led the Turkish language to be behind on recent advances in natural language processing (NLP). Having comprehensive evaluations and standardized benchmarks are important in the progress of Turkish LLMs, as they help to identify what is working. To contribute to the progress of Turkish LLMs, we evaluated 10 open-source LLMs on 11 tasks using 17 datasets. We used both originally Turkish datasets, and translated datasets. In some tasks, we used Turkish and multilingual pre-trained language models (PLMs), including TURNA, BERTurk, and mT5, as baselines. We found that the gemma2-9b-it is the clear best performer among the chat versions of the LLMs we evaluated, in both understanding and trustworthiness tasks. However, in fine-tuning experiments, there was no clear best performer and the best PLM had comparable results with the best LLM in most of the experiments. We found that there are significant differences in the performances of the models on paraphrased prompts, suggesting that robustness is an area that the LLMs need improvement on. We also observed that models like Trendyol-8B-chat-v2.0 and wiroai-turkish-llm-8b, which are adapted to Turkish by instruction tuning multilingual LLMs with Turkish instructions, usually outperform the models they are based on.

Author

Dr. Mustafa Burak Topal

How to Cite

Mustafa Burak Topal (Master Thesis). Büyük dil modellerinin Türkçe veri kümeleri üzerinde kapsamlı ve çok-istemli değerlendirmesi, 2025, Boğaziçi University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Boğaziçi University