Multilingual identification of verbal multiword expressions using bidirectional long short-term memory based architectures
Is this your thesis?
This record came from a bulk archive import. If it’s yours, link it to your profile.
Abstract (EN)
Verbal multiword expression (VMWE) identification is a challenging task for many natural language processing studies. In this study, sequence tagging approach accompanied with stochastic models and variants of IOB tagging scheme is used for VMWE identification. In the scope of this thesis, a VMWE annotated Turkish corpus is constructed as the first part of the PARSEME shared task 1.1 which is constructing VMWE annotated corpora in many languages. Additionally, a multilingual system called Deep-BGT is developed as the second part of the shared task which is developing language-independent VMWE identification systems using the corpora constructed in the first part. The Turkish corpus is one of the biggest corpora in the shared task. The training and test corpora that were published in the PARSEME shared task 1.0 are updated as the PARSEME shared task 1.1 training and development corpora according to the new guidelines. A new test corpus is constructed from scratch. Deep-BGT uses the bidirectional Long Short-Term Memory model with a Conditional Random Field layer on top (BiLSTM-CRF). To the best of our knowledge, this study is the first one that employs the BiLSTM-CRF model for VMWE identification. Deep-BGT was ranked the second in the open track in terms of the general ranking metric. Moreover, a novel tagging scheme called bigappy-unicrossy is introduced to rise to the challenge of overlapping VMWEs. Finally, the VMWE identification system is advanced by evaluating a subset of hyperparameters which consists of tagging scheme, number of units, number of BiLSTM layers, and classifier. A comprehensive analysis of BiLSTM based architectures for multilingual identification of VMWEs is presented accordingly.
Author
Gözde Berk
How to Cite
Gözde Berk (Master Thesis). Multilingual identification of verbal multiword expressions using bidirectional long short-term memory based architectures, 2019, Boğaziçi University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Boğaziçi University
- İş zekası uygulamalarında üretken yapay zekanın benimsenmesini etkileyen faktörlerin araştırılması(2025)
- Nükleer güç, emek ve çevre: Akkuyu NGS(2023)
- Darağacının ardında: Türkiye'de idam cezası, hukuk ve yasama performansı (1926-1990)(2025)
- Doğaya atfedilen değerler, doğayla bağ, çevre dostu davranış ve esenlik: İstanbul'daki kent parkları ziyaretçileri üzerine bir vaka çalışması(2025)
- Türkiye'de bölgesel kalkınma ajanslarının çevre yönetişimindeki rolü üzerine bir değerlendirme: Trakya Bölgesi üzerine bir vaka çalışması(2025)
- Türkiye'de süt üretiminin politik ekolojisi: Değişen pratikler, kırsal geçim kaynakları ve süt hayvanları(2025)