Master'sOpen Access

Biomedical entity normalization using clustering and text similarity

Is this your thesis?

This record came from a bulk archive import. If it’s yours, link it to your profile.

2024
0 views
0 downloads

Abstract (EN)

Extensive biomedical texts accumulate daily in the medical literature. The accurate identification of biological entities is of crucial importance for biomedical research, as well as for medical diagnosis and treatment, and promises significant advances in healthcare. Named Entity Recognition (NER), the recognition of entities in a text, and Named Entity Normalization (NEN), the linking of entities with their corresponding identifiers, are two related tasks that are still under investigation in natural language processing (NLP). These tasks are important to ensure the integrity of data in biological and medical databases. Normalizing biomedical entities in medical texts with the corresponding identifiers in biomedical ontologies or dictionaries is a major challenge, which is compounded by factors such as localization, unexpected abbreviations and synonyms. This challenge becomes even greater when similar words correspond to different entities and, conversely, lexically different entities have the same identity. In this thesis, we propose a NEN system that matches biological entities with their corresponding identifiers in an ontology or dictionary. Our method uses a clustering approach in combination with text similarity, using BERT-based contextual word vector representations and string similarity to normalize entity mentions. Promising results have been obtained in benchmark datasets for disease and symptom normalization compared to more complicated supervised approaches. The results show that despite its simplicity, our proposed approach is effective for named entity normalization and can be efficiently adapted to different languages and domains.

Author

Berke Kavak

Institution

How to Cite

Berke Kavak (Master Thesis). Biomedical entity normalization using clustering and text similarity, 2024, Boğaziçi University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Boğaziçi University