DoktoraAçık Erişim

Sequence labeling with stacked conditional random fields

2015
0 görüntülenme
0 i̇ndirme
Danışman: Yrd. Doç. Dr. Mehmet Fatih Amasyalı

Özet (EN)

Sequence labeling is the production of an output sequence in return for an input sequence. Many issues (name entity recognition, machine translation, morphological analysis, resolving the sentence into its elements, etc.) of natural language processing based on the contents of the input and output sequence can be defined as sequence labeling. Sentence analysis and making out the meaning of a sentence are one of the main topics of natural language processing. If real meaning requiring saying the relevant sentence can draw, this sentence can convert into action by machines, translate from one language to other language or enable to get the emotive meaning of the sentence. Dependency Parsing determines the relationships and types of relationships between words within a sentence and is essential to the semantic analysis of a sentence. When attachment discrimination is defined as the problem of sequence labeling, two-output sequence (relationship type, related word) should be generated together. Analysis of a sentence depends on the sentence structure of the relevant language. Turkish is an agglutinative language and free-intrasentence arrangements of element. Therefore, it is a language difficult to analyze compared to other language families. Although some studies exist in the literature about Turkish, there have mainly been studies on English. Studies performed for Turkish were achieved a certain degree of accuracy with Malt Parser using a Support Vector Machines-based structure. When examining the studies performed for other languages, it is clear that new hypothesis should develop and test in order to increase this success. Our suggestion is that conditional random fields used often especially in solving the sequence labeling problems can be available in dependency parsing problem. However, the conditional random fields is a method of producing a single output. In order to overcome this challenge, dependency parsing being a problem with dual outputs (attachment type and connected word) is resolved by dividing into two parts. After, the results is provided as an output of the system by combining. Compared the studies carried out for Turkish with the results in the literature, it shows that a higher success rate was reached. Apart from Turkish, the method we recommended has also been tested for Swedish, Danish, Dutch and Portuguese languages. The success of studies in the literature has been exceeded to determine the kind of relationship. A poorer performance was exhibited to determine related word. This results from more variable of intra-sentence attachment structures of these languages other than Turkish. A more dynamic structure should develop to enhance the performance of the method developed as future work in other languages.

Yazar

Metin Bilgin

Bu Yayına Nasıl Atıf Yapılır

Metin Bilgin (Doctorate thesis). Sequence labeling with stacked conditional random fields, 2015, Yıldız Technical University.

Anahtar Kelimeler

Lisans

Tüm Hakları Saklıdır

Bu eser belirtilen lisans koşulları altında paylaşılmaktadır.

Yıldız Technical University tezlerinden daha fazlası