Türkçe tıbbi veri setlerinde ICD ve ATC sınıflandırmaları için NLP modellerinin karşılaştırmalı analizi
Is this your thesis?
This record came from a bulk archive import. If it’s yours, link it to your profile.
Abstract (EN)
This study incorporates various methodologies to classify the Turkish medical texts with the aim of identifying ICD (International Classification of Diseases) & ATC (Anatomical Therapeutic Chemical) codes, leveraging various NLP approaches. The primary challenge involves comprehending the context and accurately categorizing medium to long documents into correct classes. Another challenge encompasses the acquisition of a labeled dataset with all categories, given the limited data resources in Turkish, which is a low resource language. Two distinct datasets are acquired for this study. The first dataset, which focuses on specific ICD-CM-10 C-Types, consists of medical summaries in Turkish language and was externally sourced. The second dataset, which is a new dataset including drug manuals covering all ATC types in Turkish is curated. After dataset creation, text processing has been implemented to obtain better performance results. SVC (Support Vector Classifier), FastText and various BERT models, such as BERTurk Uncased, BERTurk Large, ConvBERTurk, ElectraBERTurk and DistilBERTurk are used to classify the documents. Hyperparameter tuning is also applied to harness the potential of BERT and FastText, resulting in a notable 86% overall F1-Score on the Turkish medical summaries dataset and DistilBERTurk, the Turkish version of DistilBERT, resulted in 92.3% overall F1-Score on the same dataset. The most notable accomplishment, however, was for the Turkish Drug Manuals ATC code dataset, where the performance results, 96% F1-score using FastText, 95.9% F1-score using BERTurk Uncased and 94% using SVC models are obtained. These results highlight the effectiveness of the used approaches in Turkish text classification.
Author
Damla Büşra Özsönmez
Institution
How to Cite
Damla Büşra Özsönmez (Master Thesis). Türkçe tıbbi veri setlerinde ICD ve ATC sınıflandırmaları için NLP modellerinin karşılaştırmalı analizi, 2023, Galatasaray University.
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Galatasaray University
- Natural rights in humanism and transhumanism(2025)
- International state responsibility arising from new space activities(2025)
- Yeni roman: claude simon ve william faulkner(2014)
- Kentsel dünyanın 3D algısı için derin öğrenme tabanlı tespit ve segmentasyon(2025)
- Langlands fonktörsellik ilkesi(2021)
- Lawful use of data in machine learning-based artificial intelligence under the Turkish law(2021)
