DoctorateOpen Access

A new fusion reranking pipeline for Turkish datasets using fine-tuned RAG components

Is this your thesis?

This record came from a bulk archive import. If it’s yours, link it to your profile.

2025
0 views
0 downloads

Abstract (EN)

This study addresses the gap in the multilingual capabilities of Retrieval Augmented Generation (RAG) systems for the Turkish language, particularly in the medical domain. With the rise of Large Language Models (LLMs) and their widespread applications, the reliance on external knowledge through retrieval components has become crucial to mitigate hallucinations and improve response accuracy. However, most existing retrieval components, including embeddings and rerankers, are predominantly trained on English datasets, highlighting a significant limitation in multilingual and domain-specific capabilities. To address this, the study introduced Pubmed-RAG-TR, a Turkish-language medical dataset, and fine-tuned retrieval components on both Pubmed-RAG-TR and WikiRAG-TR, a Turkish RAG dataset. A novel RRF-based reranker pipeline was also developed to improve the context construction for LLMs. Experimental results demonstrated that fine-tuning retrieval components on domain-specific datasets significantly enhanced the retrieval and post-retrieval quality, improving the accuracy of LLM responses. The study concludes that incorporating domain-specific semantics into retrieval and reranking models can substantially boost the performance of RAG systems in multilingual contexts.

Author

Erdoğan Bıkmaz

How to Cite

Erdoğan Bıkmaz (Doctorate thesis). A new fusion reranking pipeline for Turkish datasets using fine-tuned RAG components, 2025, Çankaya University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Çankaya University