Master'sOpen Access

Document classification using improved word embeddings

2023
0 views
0 downloads
Advisor: Dr. Öğr. Üyesi Ayhan Akbaş

Abstract (EN)

In this study, one of our primary objectives is to develop a trained model capable of classification using deep learning techniques for prediction. Neural networks, especially in the realm of natural language processing, have demonstrated impressive results, notably in document classification. Researchers have focused extensively on classification prediction. Convolutional network models, recurrent networks, and other embedding mechanisms are employed where texts are extracted (embedded) from documents either at the sentence or word level. Historically, the Word2Vec model was utilized in natural language processing to extract words based on context. This was later augmented with Long Short-Term Memory (LSTM) networks. The use of N-gram properties, in context with the text and associations between words, has proven to enhance prediction accuracy in classification tasks. Previous studies have primarily based document classification on visual methods or formats, perhaps concentrating on titles and abstracts. However, this article posits that classification should be anchored in word inclusion. Utilizing a dataset comprised of 47,000 texts and topics, we employ word embeddings to determine document themes. These embedded words — vast textual content — are sorted into seven primary categories, serving as foundational classes in our dataset. This data then trains deep learning models designed for document classification (both for training and testing). Once trained, this model can autonomously classify documents based on embedded words and texts. Our approach begins by extracting words from the dataset's texts. Subsequently, two models are constructed using Word2Vec. The words undergo lemmatization, reverting to their original form. Superfluous elements, such as symbols and punctuation, are purged to ensure the text remains pure, concentrating solely on semantically significant words. These cleansed word series are then used to train the two models, aiming to establish correlations between words. Both models strive to construct associations based on word sequences within the text. The first model assigns vectors to words based on context and endeavors to predict context via these words. In contrast, the second model hinges on the interrelationships between words and predicts specific words based on classification, yielding a relational concept termed "Neighbor word". Finally, we employ a deep learning model rooted in Long Short-Term Memory (LSTM). This is buttressed by the relationships deduced from the two Word2Vec models. Evaluations between them are conducted to ascertain which offers superior performance in predictive classification

Author

Raad Saadı Mahmood Mahmood

How to Cite

Raad Saadı Mahmood Mahmood (Master Thesis). Document classification using improved word embeddings, 2023, Çankırı Karatekin Üniversitesi.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Çankırı Karatekin Üniversitesi