Identification of verbal multiword expressions using deep learning architectures and representation learning methods
Is this your thesis?
This record came from a bulk archive import. If it’s yours, link it to your profile.
Abstract (EN)
Understanding multiword expressions (MWEs) plays an instrumental role in Natural Language Processing applications such as parsing and machine translation. MWE identification is a task that automatically detects and classifies MWEs in running text. As with the basic characteristics of MWEs, significant challenges exist in MWE identification. Considering the recent attempts of the PARSEME network on verbal multiword expressions (VMWEs), we focus on the identification of VMWEs. We update the PARSEME Turkish train and test corpora 1.0 (2017) as the PARSEME Turkish train and development corpora 1.1 (2018). We construct the PARSEME Turkish test corpus 1.1. In addition, we develop a multilingual VMWE identification system based on bidirectional long short term memory with conditional random fields networks accompanied with the gappy 1-level tagging scheme. To extend our study, we examine the impact of data representation format on the VMWE identification task. We introduce the bigappy-unicrossy tagging scheme to recognize overlaps in sequence labelling tasks. Our results show that data representation format is important to identify discontinuous VMWEs. Moreover, we enhance our neural VMWE identification model with automatically learned embeddings by neural networks to respond to the variability challenge. We compare character-level convolutional neural networks and character-level bidirectional long short-term (BiLSTM) networks. We analyze two different schemes to represent morphological information using BiLSTM networks. Our results demonstrate that character embeddings and morphological embeddings improve performance in general. The choice of representation learning method depends on language.
Author
Berna Erden
How to Cite
Berna Erden (Master Thesis). Identification of verbal multiword expressions using deep learning architectures and representation learning methods, 2019, Boğaziçi University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Boğaziçi University
- İş zekası uygulamalarında üretken yapay zekanın benimsenmesini etkileyen faktörlerin araştırılması(2025)
- Nükleer güç, emek ve çevre: Akkuyu NGS(2023)
- Darağacının ardında: Türkiye'de idam cezası, hukuk ve yasama performansı (1926-1990)(2025)
- Doğaya atfedilen değerler, doğayla bağ, çevre dostu davranış ve esenlik: İstanbul'daki kent parkları ziyaretçileri üzerine bir vaka çalışması(2025)
- Türkiye'de bölgesel kalkınma ajanslarının çevre yönetişimindeki rolü üzerine bir değerlendirme: Trakya Bölgesi üzerine bir vaka çalışması(2025)
- Türkiye'de süt üretiminin politik ekolojisi: Değişen pratikler, kırsal geçim kaynakları ve süt hayvanları(2025)