DoctorateOpen Access

A framework for open-domain question answering system

2024
0 views
0 downloads
Advisor: Doç. Dr. Baha Şen

Abstract (EN)

In general, Question Answering (QA) can be defined as an automatic process that is capable of understanding questions posed in a natural language such as English and responding exactly with requested information. An "ideal" QA system has a highly complex architecture. Because this system has to determine the desired information in the question, find the requested information from suitable sources, extract information, and then create an answer. Users prefer a QA system to find precise answers to their questions rather than inspect all related documents relevant to search queries. The studies on the systems automatically answering natural language questions started in the 1960s. It has become the leading research area within the information retrieval community, with the QA track started in 1999 under the Text Retrieval Conference (TREC). Contrary to open domain QA systems, fewer researchers are working on medical, domain-specific question answering. Due to the continuous increase in information produced in the biomedical field, there is an increasing need for biomedical QA, especially for the public, medical students, healthcare professionals, and biomedical researchers. In a sense, biomedical QA is one of the most critical applications of the real world. In this study, a question-answering system was developed for the biomedical field. Four different ranking algorithms (Vector Space Model, Okapi BM25, Query Likelihood with Dirichlet Smoothing, and the Jelinek–Mercer Smoothing Model) were tested for the system's document retrieval component. The best performance was achieved with the Query Likelihood with Dirichlet Smoothing ranking algorithm using a query expanded with MESH terms. In the answer extraction component, in addition to text similarity, Named Entity Recognition (NER), UMLS Concept Unique Identifiers (CUIs), UMLS Semantic Types, and UMLS Semantic Group features were employed to find sentences that might be the answer. The F1 score based solely on text similarity was increased from 0.27 to 0.39, achieving an approximate 44% performance improvement. Based on transformer architecture, the BERT language model was trained for the biomedical field and fine-tuned for the biomedical question-answering system using the SQuAD and BioASQ 9b train datasets. For factoid questions in the BioASQ 9b test datasets, a 0.72 MRR score was achieved.

Author

Harun Bolat

How to Cite

Harun Bolat (Doctorate thesis). A framework for open-domain question answering system, 2024, Ankara Yıldırım Beyazıt University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Ankara Yıldırım Beyazıt University