DoctorateOpen Access

Corpus-driven semantic relations extraction for Turkish language

Is this your thesis?

This record came from a bulk archive import. If it’s yours, link it to your profile.

2014
0 views
0 downloads

Abstract (EN)

Identification of semantic relations is the core problem in many Natural Language Processing tasks. One of the important tasks is to build up ontology or to construct thesaurus/lexicon. The most popular and widely used lexical database, WordNet is developed by manually. So it is used as source and also comparable work for most of the studies. Although these types of lexicons are reliable and effective, their production can be troublesome and time-consuming in some cases. So acquisition of semantic relation automatically from large amount of electronic documents (corpora, dictionaries, newspapers, newswires, etc.) becomes more important. In this study, automatic and semi-automatic acquisition system for acquisition of hyponym/hypernym, meronym/holonym and synonym relations are handled from large corpus in Turkish Langage for nouns. For this purpose, some sort of methods is proposed to realize the model. The method for hyponym/hypernym relation relies on lexico-syntactic pattern and semantic similarity. Once the model has extracted the items using patterns, it applies similarity based elimination of the incorrect ones in order to increase precision. Second model is based on similarity based expansion in order to increase recall. Several scoring functions are within bootstrapping algorithm are applied. For meronym/holonym, lexico-syntactic patterns are utilized and adopted again to a Turkish huge corpus. Two different approaches are proposed to prepare patterns; one is based on pre-defined patterns that are taken from literature, second automatically produces patterns by means of bootstrapping method. Pre-defined patterns are categorized into two clusters; General and Dictionary-based patterns. Once these patterns help the system to extract matched cases, it proposes a list of part-whole pairs depending on their co-occur frequencies. For latter, bootstrapping model takes manually prepared unambiguous seeds to induce syntactic patterns and estimates their reliabilities. Then, system extracts pair instances then ranks them by instance reliability scoring. Additional, statistical selection is used on global data obtaining from all results of entire patterns, where global data refers to a whole-by-part matrix on which several association metrics such as information gain, T-score etc. are measured and compared Finally, how these patterns and statistical method improve the system accuracy especially within corpus-based approach and distributional feature of words is evaluated. For synonym relation, the main assumption is that synonym pairs show similar semantic and syntactic characteristics by the definition. They share same meronym/holonym and hypernym/hyponym relations. Contrary to synonymy, hypernymy and meronymy relations can be easily acquired by applying lexico-syntactic patterns to a corpus. Such acquisition might be utilized and ease detection of synonymy. Likewise, some particular syntactic relations are utilized such as object/subject of a verb etc. Machine learning algorithms were applied on all these acquired features. The first aim is to find out which syntactic and semantic features are the most informative and contributes most to the model. Performance of each feature is individually evaluated with cross validation. The model that combines all features shows promising results and successfully detects synonymy relation. Another model is proposed to extract synonym relation with using integration of some sort of sources such as WordNet, bilingual on-line dictionary and monolingual on-line dictionary. The main contributions of the study is considered as being first major attempt for Turkish hyponym/hypernym, meronym/holonym and synonym identification based on corpus-driven approach for Turkish Language. Second contribution is to use integrated approaches such as pattern-based method with statistical elimination and expansion, bootstrapping patterns, etc. for extracting relations. Third contribution is to use multiple resources such as WordNet, mono/bilingual on-line dictionaries, etc. and to integrate them.

Author

Tuğba Yıldız

How to Cite

Tuğba Yıldız (Doctorate thesis). Corpus-driven semantic relations extraction for Turkish language, 2014, Yıldız Technical University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Yıldız Technical University