Master'sOpen Access

A hybrid approach supported by transformer-based data augmentation for turkish abstractive text summarization

2025
0 views
0 downloads
Advisor: Doç. Aysun Güran

Abstract (EN)

This dissertation investigates the optimization of transformer-based models for automated text summarization in Turkish, a language with limited linguistic resources and examines the influence of diverse data augmentation strategies on model efficacy. The study seeks to overcome the constraints of scarce training data, aiming to construct more reliable and precise summarization frameworks for Turkish. The research employs encoder-decoder models, namely MBART, MT5 and VBART which were trained on both unaltered datasets and enriched versions incorporating augmentation techniques such as back-translation, synonym substitution, and a combined approach integrating both methods. To evaluate the effects of these augmentation strategies, the study adopts a dual approach: quantitative assessment through ROUGE metrics, supported by robust statistical methods, and qualitative evaluation via content analysis to scrutinize the coherence and substantive quality of the generated summaries. The outcomes of this research are anticipated to contribute significantly to the advancement of Turkish text summarization systems and highlight the pivotal role of data augmentation in enhancing natural language processing for languages with limited resources.

Author

Umut Can

How to Cite

Umut Can (Master Thesis). A hybrid approach supported by transformer-based data augmentation for turkish abstractive text summarization, 2025, Doğuş University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Doğuş University