Master'sOpen Access

Classification of medical documents according to diseases

2016
0 views
0 downloads
Advisor: Yrd. Doç. Dr. Alper Kürşat Uysal

Abstract (EN)

The number of documents produced on computers has increased exponentially every year, after the spreading use of the computers. Automatic text classification has become an important due to the exponential growth of texts on the Internet. Significant problems in text classification are the great number of features and misclassification are made accordingly. In this thesis, it is constructed of two different datasets containing English and Turkish abstract belonging to Turkish articles in the medical field. This dataset is similar structure to namely Ohsumed which is containing English medical text summary. In the literature, there is no dataset like Ohsumed datasets obtained from Turkish datasets to be used in academic studies. Various preprocessing, feature selection and successful classifiers in this field are used in automatic text classification stages. It has been investigated in the basis of languages how influences the performance of the classification according to whether stemming which differs in languages and one of the preprocessing steps applied or not. And also, the classification performance of different feature selection method has been investigated. Classifier performance which is another factor affecting the performance was analyzed by applying different classifiers. Finally, classification schemes that provide the best performance on the medical text summary in the same publication and different languages is determined. Keywords: Text Classification, Feature Selection Methods, Classification Algorithms, Preprocessing Steps

Author

Bekir Parlak

How to Cite

Bekir Parlak (Master Thesis). Classification of medical documents according to diseases, 2016, Anadolu University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Anadolu University