Multi-label and single-label text classification using standard machine learning algorithms and pre-trained BERT transformer
2023
0 views
0 downloads
Advisor: Doç. Dr. Abdulkadir Gürer
Abstract (EN)
Natural language processing (NLP) research has received a great deal of attention in recent times, because of the increasing availability of digital documents and the resulting need to access them in various ways. The explosion of digital text data demonstrates the need to develop diverse text processing and classification techniques. The most essential and vital challenge in NLP is text classification. It was proposed for this purpose to classify documents and texts into pre-determined categories based on their contents, and it has since become one of the most popular methods of implementing machine learning. The machine learning (ML) paradigm is one where a generic inductive approach learns to create a privately classified text using a set of classified texts and the features of the classes of interests. Furthermore, discovering the relevant information can help improve information retrieval efficiencies while reducing the overload of information. Traditional models typically require artificial methods for obtaining good sample attributes before classifying them using standard machine learning algorithms. Therefore, feature extraction restricts the method's effectiveness significantly. On the other hand, deep learning differs from typical models, which are getting more attention because they incorporate feature extraction into the model building approach by performing a series of nonlinear transformations that assist in transferring feature representations to outputs. Furthermore, deep learning algorithms avoid the need for experts to define rules and attributes, instead automatically providing high-level semantic representations for texts. Therefore, in these studies, we explore the capabilities of contextually embedding derived from pre-trained models like BERT, and make use of multi-label classification of text documents in a huge English news dataset, in addition to some traditional machine learning methods to be applied in a small English news dataset. Finally, another version of BERT, Arabic BERT, explores sentiment polarity toward extracted aspects in an Arabic hotel review dataset.
Author
Huda Alfıgı
Institution
How to Cite
Huda Alfıgı (Master Thesis). Multi-label and single-label text classification using standard machine learning algorithms and pre-trained BERT transformer, 2023, Çankaya University.
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Çankaya University
- Investigation of amazon and google for fault tolerance strategies in cloud computing services(2015)
- Exchange rate and inflation relationship: The case of Turkey(2023)
- Effects of the economic news on herd behavior(2023)
- Experimental analysis of effects of different network parameters on TCP / IP networks(2025)
- Reconstruction of patriarchy through matriarchy: A critique of gendered power structures in Naomi Alderman's The Power(2025)
- Characterization of under-hood airflow in construction equipment using experimental techniques(2025)