Yüksek LisansAçık Erişim

Proteinlerdeki allosterik bölgelerin gelişmiş tahmini için protein dil modellerini (pLM'ler) ML/DL yaklaşımlarına dahil etme potansiyelinin araştırılması

2023
0 görüntülenme
0 i̇ndirme
Danışman: Prof. Dr. Attila Gürsoy ; Prof. Dr. Zehra Özlem Keskin Özkaya

Özet (TR)

Allostery, the process by which binding at one site perturbs a distant site, is being rendered as a key focus in the field of drug development with its substantial impact on protein function. Allosteric drugs activate or inhibit proteins and offer advantages over non-allosteric drugs. However, the identification of allosteric sites is a challenging task due to unavailability of huge dataset, their distance, and lack of conservation across protein structures. A variety of computational techniques have been developed in the past to predict allosteric sites, such as Normal Mode Analysis (NMA), Molecular Dynamics (MD), and Machine Learning (ML), utilizing both static pocket characteristics and the dynamics of proteins; the performance of these methods needs further improvement. This research investigates the potential of incorporating Protein Language Models (pLMs) into ML and/or DL approaches to improve prediction of allosteric residues, based on the fact that the pLMs (e.g., ProtBERT from the family of ProtTrans pLMs: based on BERT architecture) effectively capture the spatial relationship among residues, eventually contributing to identification of allosteric sites/pockets. ProtBERT-BFD (ProtTrans) was fine-tuned on the Allosteric Dataset (ASD) of protein sequences, which predicts the allosteric residues with an F1 score of 61.54% on the test dataset. Several ML and DL approaches were utilized including XGBoost, SVM, AutoML, and GNNs. With the inclusion of fine-tuned pLM features, all of the aforementioned approaches improve the prediction performance of allosteric sites over previous studies by a considerable margin. XGBoost, being the highest performing model in this study, improves the results by combining the features extracted from finetuned ProtBERT with pocket features extracted by FPocket, resulting in an F1 score of 75.76% for allosteric pockets/sites. Case studies have been performed on proteins with known allosteric sites in addition to the case study to predict novel allosteric sites on new proteins.

Yazar

Dr. Moaaz Ur Rehman Azhar Khokhar

Bu Yayına Nasıl Atıf Yapılır

Moaaz Ur Rehman Azhar Khokhar (Yüksek Lisans Tezi). Proteinlerdeki allosterik bölgelerin gelişmiş tahmini için protein dil modellerini (pLM'ler) ML/DL yaklaşımlarına dahil etme potansiyelinin araştırılması, 2023, Koç University.

Anahtar Kelimeler

Lisans

Tüm Hakları Saklıdır

Bu eser belirtilen lisans koşulları altında paylaşılmaktadır.

Koç University tezlerinden daha fazlası