Master'sOpen Access

Using transformer networks for detection andnormalization of named entities in biomedical texts

Is this your thesis?

This record came from a bulk archive import. If it’s yours, link it to your profile.

2021
0 views
0 downloads

Abstract (EN)

The increasing difficulty of retrieving relevant information from rapidly growingliterature has raised the interest for natural language processing (NLP) systems in thebiomedical domain. In many of these systems, detection of named entities such asdiseases, genes, and molecules (named entity recognition) and matching them to thecorresponding entries in ontologies (normalization) are important intermediate steps.As these two tasks are related and datasets in this domain are relatively small, multi-task learning has been frequently used in the literature for this problem. Meanwhile,in recent years, the success of transformer-based pre-trained language models suchas BERT in various NLP tasks has led them to be also applied in the biomedicaldomain. The different characteristics of biomedical text such as abbreviations andspecific terminology motivated the development of new language models, which weretrained specifically for this domain using a biomedical corpus. In this study, we proposea multi-task learning approach for named entity recognition and normalization byutilizing transformer-based pre-trained language models. To enable the optimal sharingof information, both tasks are formulated with text span embeddings obtained witha common encoder network. Promising results are obtained and compared with theresults of state-of-the-art systems from the literature for commonly used named entityrecognition datasets.

Author

İlkay Ramazan Pala

How to Cite

İlkay Ramazan Pala (Master Thesis). Using transformer networks for detection andnormalization of named entities in biomedical texts, 2021, Boğaziçi University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Boğaziçi University