DoctorateOpen Access

Akademik bütünlükün sağlanması: insan tarafından yazılmış ve yapay zekâ tarafından üretilmiş metinleri ayırt etmek için transformer tabanlı ve toplu makine öğrenme yaklaşımları

2025
0 views
0 downloads
Advisor: Doç. Dr. Oğuz Ata

Abstract (EN)

The explosive development of large language models (LLMs) such as ChatGPT and Bard transformed text creation, offering new possibilities for academic communication and raising deep concerns regarding authenticity and authorship. As AI-created content has come to be almost indistinguishable from human writing, the need for open and reliable mechanisms to authenticate authorship has become crucial in sustaining academic trust. This work presents and illustrates two mutually complementary Natural Language Processing (NLP) frameworks the Transformer-Based Detection Framework and the Ensemble Learning-Based Detection Framework and contrasts them on three evenly balanced datasets: AI-GA, HWAI, and HAGT-1M. The former uses Sentence-BERT (SBERT) and RoBERTa embeddings together with Logistic Regression (LR) and Feed-Forward Neural Networks (FNNs) to detect deep semantic and syntactic trends. Across datasets, transformer configurations achieved strong performance: SBERT+LR: 93.66% (AI-GA), 91.37% (HWAI), 99.64% (HAGT-1M); RoBERTa+FNN: 91.13% (AI-GA), 90.29% (HWAI), 99.95% (HAGT-1M) with RoBERTa-FNN peaking at 99.95% on HAGT-1M. To ensure transparency and surmount the "black-box" limitation of transformers, the second method leverages ensemble learning by combining LR, SVM, Random Forest, Extremely Randomized Trees, XGBoost, AdaBoost, and SGD with soft voting over TF-IDF and linguistic feature sets (n-grams, lexical diversity, sentence length). The ensemble achieved 99.88% (AI-GA), 99.37% (HWAI), and 100% (HAGT-1M), outperforming prior state-of-the-art and demonstrating increased robustness on varied academic text. The current study focuses on reproducible preprocessing, data exploration, and careful threshold tuning regarding sensitivity and fairness, especially for non-native research scholars. Model linguistic authenticity as a three-dimensional property present in syntax, semantics, and stylistic coherence. Offer open, scalable, and transparent detection protocols suitable for universities, publishers, and research integrity offices. Future research will include the investigation of hybrid quantum machine learning (HQML), explainable-AI tools with more detail (e.g., SHAP/LIME/attention visualization), and multilingual extension in order to maintain cross-domain generalization.

Author

Dr. Layth Rafea Hazım Hazım

How to Cite

Layth Rafea Hazım Hazım (Doctorate thesis). Akademik bütünlükün sağlanması: insan tarafından yazılmış ve yapay zekâ tarafından üretilmiş metinleri ayırt etmek için transformer tabanlı ve toplu makine öğrenme yaklaşımları, 2025, Altınbaş University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Altınbaş University