Master'sOpen Access

Deep learning-based preprocessing tools for Turkish natural language processing

Is this your thesis?

This record came from a bulk archive import. If it’s yours, link it to your profile.

2024
0 views
0 downloads

Abstract (EN)

As the demand for effective natural language processing (NLP) applications in Turkish continues to rise, the need of text preprocessing tools tailored to Turkish increases. The text preprocessing tools developed play a significant role in extracting information from raw data. As an initial step for any natural language application, these tools enhance the efficiency of complex tasks such as text summarization, questionanswering and machine translation. This thesis describes the development and evaluation of deep-learning based preprocessing tools specifically designed for Turkish. The developed deep-learning based tools tackle the issues of existing rule-based approaches for Turkish. In this thesis, preprocessing tools are developed with Long Short-Term Memory (LSTM) networks and transformers. The main advantage of these architectures is that they are able to adapt to the sequential flow of the natural language effectively. In this thesis, we investigate the performance of LSTMs and transformers architectures on Turkish preprocessing tasks including tokenization, sentence splitting, deasciification, vowelization, part-of-speech (POS) tagging, spelling correction and morphologic analysis-disambiguation as possible components of a preprocessing pipeline.

Author

Buse Ak

How to Cite

Buse Ak (Master Thesis). Deep learning-based preprocessing tools for Turkish natural language processing, 2024, Boğaziçi University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Boğaziçi University