Master'sOpen Access

Türkçedeki yüksek frekanslı yapım eklerinin semantiğinin keşfi için dağılımsal incelenmesi

2021
0 views
0 downloads
Advisor: Prof. Dr. Hüseyin Cem Bozşahin

Abstract (EN)

In agglutinating languages such as Turkish, the process of derivation is mostly performed by adding suffixes at the end of words. Most of the derivational suffixes carry a distinctive semantic content and representing them has an important role in computational tasks, such as question answering. In this thesis, we aim to explore the structure of some frequent Turkish derivational suffixes in distributional vector space by clustering word embedding vectors of them and analyzing their underlying semantic properties. Suffix vectors are obtained by subtracting the vector of the base form of the derived word from the derived word's word vector. We used a pre-trained word embedding model for obtaining word vectors and multiple unsupervised clustering algorithms with different parameters for clustering them. Our assumption is if a derivational suffix category manages to dominate one or more clusters, it is possible to obtain reliable representations of it in the distributional vector space. Our results show that many Turkish derivational suffix categories have this capability. We analyzed the underlying semantic structure of the generated clusters in terms of the thematic roles the suffixes are selecting, the UCCA labels and the UD relations the stem and the derived word can get.

Author

Dr. Gizem Nur Özdemir

How to Cite

Gizem Nur Özdemir (Master Thesis). Türkçedeki yüksek frekanslı yapım eklerinin semantiğinin keşfi için dağılımsal incelenmesi, 2021, Middle East Technical University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Middle East Technical University