Biomedical entity normalization using clustering and text similarity
Is this your thesis?
This record came from a bulk archive import. If it’s yours, link it to your profile.
Abstract (EN)
Extensive biomedical texts accumulate daily in the medical literature. The accurate identification of biological entities is of crucial importance for biomedical research, as well as for medical diagnosis and treatment, and promises significant advances in healthcare. Named Entity Recognition (NER), the recognition of entities in a text, and Named Entity Normalization (NEN), the linking of entities with their corresponding identifiers, are two related tasks that are still under investigation in natural language processing (NLP). These tasks are important to ensure the integrity of data in biological and medical databases. Normalizing biomedical entities in medical texts with the corresponding identifiers in biomedical ontologies or dictionaries is a major challenge, which is compounded by factors such as localization, unexpected abbreviations and synonyms. This challenge becomes even greater when similar words correspond to different entities and, conversely, lexically different entities have the same identity. In this thesis, we propose a NEN system that matches biological entities with their corresponding identifiers in an ontology or dictionary. Our method uses a clustering approach in combination with text similarity, using BERT-based contextual word vector representations and string similarity to normalize entity mentions. Promising results have been obtained in benchmark datasets for disease and symptom normalization compared to more complicated supervised approaches. The results show that despite its simplicity, our proposed approach is effective for named entity normalization and can be efficiently adapted to different languages and domains.
Author
Berke Kavak
Institution
Boğaziçi University
Bilgisayar Bilimi ve Mühendisliği Bilim Dalı
How to Cite
Berke Kavak (Master Thesis). Biomedical entity normalization using clustering and text similarity, 2024, Boğaziçi University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Boğaziçi University
- Investigating the factors affecting the acceptance of generative artificial intelligence in business intelligence applications(2025)
- Nükleer güç, emek ve çevre: Akkuyu NGS(2023)
- Behind the gallows: Capital punishment, law, and legislative performance in Turkey (1926-1990)(2025)
- Exploring the values for nature, nature connectedness, pro-environmental behaviour, and well-being: A case study on urban park visitors in Istanbul(2025)
- An assessment on the role of regional development agencies in environmental governance in Türkiye: A case study on Thrace Region(2025)
- Political ecology of milk production in Türkiye: Changing practices, rural livelihoods, and dairy animals(2025)