Theses supervised by Prof. Dr. Attila Gürsoy ; Prof. Dr. Zehra Özlem Keskin Özkaya

12 theses · Koç University

Master'sOpen AccessEN

LISA: A fast filtering algorithm for structural alignment

Protein - Protein Interactions (PPI) cause vital processes such as growing, division or maintenance of a cell. It is an important topic in order to understand the functioning of any cell with itself and others. Even though the most reliable techniques are experimental for investigating the properties of different structures, numerical methods are much faster with negligible error. Nevertheless, predictions of PPI needs improvements in order to derive more accurate results faster. In this thesis, we are implementing a two-step hashing algorithm in order to increase the speed of prediction of PPI structures. In the first part, a large number of interfaces are classified by their different properties such as bond angles, dihedral angles and distances between Carbon Alpha (CA) atoms. For each non-consecutive residue that has the distance of 4 $\AA$ to 13 $\AA$, between CA atoms,we are calculating necessary angles and distances with selected CA atoms and their consecutive neighbors. Then, we are classifying fragments of interfaces with similar properties in the first phase of the algorithm by using a hash table. In the second phase, we are comparing a given protein using the same properties that calculated for templates, and score them by their similarity. Proposed algorithm is developed for filtering the dissimilar interfaces and limiting the possible number of interfaces that can be used in Template-Based PPI prediction protocols. In addition, it is useful for reducing the computation time of any structural alignment algorithm to find input templates for a docking algorithm by returning a filtered subset from a given template dataset.

FiltrationCell divisionCell growth processes+1
Emre Küçük
Koç University · Institute of Graduate Studies in Science
2022
00
Master'sOpen AccessEN

Investigating the potential of incorporating protein language models (pLMs) into ML/DL approaches for enhanced prediction of allosteric sites in proteins

Allosteri, proteinin bir bölgesindeki değişikliğin, mesela başka bir moleküle bağlanmanın, proteinin uzak bir bölgesini etkilediği süreç olarak tanımlanabilir. Allosteri protein fonksiyonu üzerindeki önemli etkisi sebebiyle ilaç geliştirme alanında önemli bir odak noktasıdır. Allosterik ilaçlar proteinleri aktive veya inhibe edebilir, allosterik olmayan ilaçlara göre avantajlar sunar. Bununla birlikte, allosterik bölgelerin tanımlanması zorlu bir iştir. Geçmişte allosterik bölgeleri tahmin etmek için Normal Mod Analizi (NMA), Moleküler Dinamik (MD) ve Makine Öğrenimi (MÖ) gibi hem statik cep özelliklerini hem de proteinlerin dinamiklerini kullanan çeşitli hesaplama teknikleri geliştirilmi olmakla birlikte bu yöntemlerin performansının daha da geliştirilmesi gerekmektedir. Bu araştırmada, pDM'lerin (örneğin, ProtTrans pDM ailesinden BERT mimarisine dayalı ProtBERT'in) allosterik kalıntıların tahminini iyilrştirmek için Protein Dil Modellerini (pDM'ler), MÖ ve/veya DÖ yaklaşımlarıyla birlikte kullanılma potansiyelini araştırılıyor. Tezde, amino asitler arasındaki mekansal ilişkiyi etkili bir şekilde öğrenerek, sonuçta allosterik alanların/ceplerin tanımlanmasını hedeflenmektedir. ProtBERT-BFD (ProtTrans), test veri kümesinde %61,54'lük bir F1 puanıyla allosterik kalıntıları tahmin eden protein dizilerinin Allosterik Veri Kümesine (AVK) göre ince ayar yapılmıştır. XGBoost, SVM, AutoML ve GNN'ler dahil olmak üzere çeşitli MÖ ve DÖ yaklaşımlarından yararlanılmıştır, İnce ayarlı pDM özelliklerinin dahil edilmesiyle, yukarıda belirtilen yaklaşımların tümü, allosterik bölgelerin tahmin performansını önceki çalışmalara göre önemli bir farkla artırdığı bulunmuştur. Bu çalışmada en yüksek performansa sahip model olan XGBoost, ince ayarlı ProtBERT'ten çıkarılan özellikleri FPocket tarafından çıkarılan cep özellikleriyle birleştirerek sonuçları iyileştiriyor ve allosterik cepler/bölgeler için %75,76'lık bir F1 puanı erişmektedir. Bilinen allosterik bölgelere sahip proteinler üzerinde örnek çalışmaların yanı sıra, farklı proteinler üzerindeki yeni allosterik bölgeleri tahmin etmek için de çalışmalar yapılmıştır.

Moaaz Ur Rehman Azhar Khokhar
Koç University · Institute of Graduate Studies in Science
2023
00
DoctorateOpen AccessEN

Multiview contrastive autoencoder-transformer approach for protein-protein interface representation: Unveiling biological and functional insights

Protein-protein interactions (PPIs) play pivotal roles in various biological processes, orchestrating cellular functions essential for life. The interfaces where these interactions occur serve as focal points for understanding the mechanisms underlying disease pathways. Accurate representation of these interfaces is crucial for deciphering their biological significance and designing therapeutic interventions. This thesis introduces a novel approach for representing protein-protein interfaces using a graph-based multiview contrastive autoencoder combined with a transformer, which learns representations from a large dataset. Comprehensive evaluations demonstrate the method's effectiveness in capturing the structural and functional characteristics of protein-protein interfaces. The learned representations are applied to tasks such as biological relevance prediction, biological vs. crystal classification, and Gene Ontology term prediction, showcasing their versatility and utility in understanding PPIs. By integrating explainable AI techniques, key features contributing to model predictions are identified, enhancing the interpretability of the results. A detailed case study illustrates the practical application of these methods, highlighting their potential to provide actionable insights for biological research and drug discovery. Overall, this thesis advances the understanding of protein-protein interactions by providing interpretable representations that capture the complex structural and functional characteristics of interfaces, thereby facilitating biomedical studies and therapeutic developments.

Damla Övek
Koç University · Institute of Graduate Studies in Science
2024
00
DoctorateOpen AccessEN

Crosstalk between cardiovascular and cognitive diseases: deciphering molecular mechanisms of vascular cognitive impairment

Vascular cognitive impairment (VCI) is a growing public health concern with significant implications for human health. VCI is an understudied complex disease; therefore, this work aims to understand this disease by studying complex molecular interactions between cardiovascular (CVD) and cognitive diseases (CD). This thesis analyzes this crosstalk by building and examining protein-protein interaction (PPI) networks related to CVD and CD. By analyzing alternative protein conformations and mutations in interfaces of interactions, this thesis suggested three mutations in the kinase DYRK1A (V165I, S337P, and D401G) may be essential for VCI via its interaction with APP. Our results indicated that chemokine-related, and stress response-related pathways are likely related to Blood-Brain-Barrier dysregulation in VCI. We found mutant and wild-type VCP, XRCC4, and LIG4 conformations interacting with BRCA1. The analysis of the effect of transcription factors on CVD-CD crosstalk showed that JUN, CREB1, NFKB1, ESR1, and NR3C1 are crucial for VCI regulation, particularly the interaction between the mutant conformation of NFKB1 (structure: 2O61 chain B) and wild-type conformation of NR3C1 structure (3H52 chain B). Lastly, we clustered and predicted disease labels with machine-learning models. We found that GUILD scores and ESMs as features could be crucial when training a model to predict disease labels. Clustering resulted in four clusters enriched in pathways that support our previous results and suggest that RhoGTPase Signaling could play an essential role in the CVD-CD crosstalk and VCI progression. This thesis uncovered biomarkers, protein-protein interactions (protein structures/conformations that interact), and potential therapeutic targets to benefit the scientific and medical communities. This work also proposed pathways linking CVD and CD to suggest new approaches for VCI intervention.

Melisa Ece Zeylan
Koç University · Institute of Graduate Studies in Science
2025
00
Master'sOpen AccessEN

Protein bağlanma bölgesi ve protein yüzeyi yapısal hizalanması aracı prototipi iMatch'in performans ve doğruluk analizi

Protein structure alignment is the task of finding an optimal transformation between two protein structures that minimizes the distances between two molecules. Since 3D protein structure is one of the most important characteristics of proteins and the number of deposited structures in the Protein Data Bank (PDB) is exponentially increasing, the structural alignment of proteins has been studied extensively. The structural alignment of proteins can provide evolutionary information, allow prediction of function, identify homologs and provide a classification mechanism. In this thesis, a new structural alignment method called iMatch has been proposed. iMatch is based on the pairwise structural alignment component of MultiProt and is designed as the structural alignment component for alignment of binding sites onto protein surfaces. iMatch uses the object recognition method geometric hashing to identify the similar local regions on the proteins. These local regions are then clustered according to their transformation similarities and extended further to obtain a final global alignment. The heuristics used throughout the alignment process and the working mechanism of iMatch has been explained in detail. iMatch is a highly customizable alignment method due to its parametric nature. This flexibility allows the configuration of iMatch according to the task in hand for better accuracy. Furthermore, this parametric nature grants iMatch the ability to be adjusted for the binding site – surface alignment. The default optimal parameters have been identified by a two-step optimization process and each parameter's effect to the alignment speed and accuracy has been inspected thoroughly for future studies on binding site – surface alignment. The overall performance of iMatch has been evaluated on three different well-known datasets against state-of-the-art structural alignment methods existing today.

Deniz Demircioğlu
Koç University · Institute of Graduate Studies in Science
2015
00
Master'sOpen AccessEN

Prısm 2.0: Prism'in iyileştirilmesi ve protein - protein etkileşimlerinin yapısal modellenmesini yapan bir internet sunucusunun tasarlanmasi

Biological processes in the cell mainly are carried out by protein – protein interaction networks among numerous proteins. One of the goals of the computational biology is to use mainly computational methods to predict interaction among proteins and understand how a cell functions. Today, there are many experimental and computational approaches used extensively to generate protein interaction data. One of the pioneer computational tools is PRISM. PRISM is the first template based protein – protein interaction prediction tool developed. PRISM has proven its success at many publications and used widely for prediction and structural modeling of protein-protein interactions as a stand-alone tool. However, its application to pathways, a network of protein-protein interactions, has both performance and usability limitations. In this thesis, a new version of PRISM (PRISM2.0) has been designed and implemented to make it efficient and easy-to-use for network of proteins predictions. PRISM2.0 is built from loosely coupled modules, therefore it is easy to use and develop further. A web server version of PRISM2.0 has been developed to create a protein–protein interaction repository. As the repository grows, it is expected to be a rich resource for protein-protein interactions. Users can enjoy using the web server without any installation or any further information about computer world rather than browsing the web page. We believe these tools will be beneficial for the community who is interested structural modeling of biological pathways.

Alper Başpınar
Koç University · Institute of Graduate Studies in Science
2015
00
Master'sOpen AccessEN

İlaç hedef dışı proteinlerinin tahmini: Aldehit dehidrogenaz inhibitörüyle bir vaka çalışması

Aldehyde dehydrogenases (ALDH) are a large family of enzymes that maintain homeostasis through metabolizing reactive compounds. They play a pivotal role in the regulation of many important biological processes such as detoxification of alcohol and xenobiotics, amino acid salvage pathways, and retinoic acid synthesis. An attractive feature of ALDH is that it is a marker for both normal and cancer stem cells. It is thought to play a role in protection, differentiation and expansion of stem cells. DIMATE, a small molecule inhibitor of ALDH3A1, has been previously shown to display anti-proliferative effects on cancer cells. Drug promiscuity refers to a drug having multiple targets. While multiple targets confer flexibility to a drug, it also creates unintended off-targets. Off-target detection is a crucial step of drug design to prevent unwanted side effects. Reverse docking may be used to identify off-targets, but it is very inefficient and has a high computational cost. An alternative to this is pharmacophore approach. A pharmacophore is a description of the chemical and geometrical features of a compound that is necessary for its interaction with a biological target to initiate or inhibit a biological activity. In this study, we identified the possible off-targets of DIMATE through similarity based approaches employing reverse-docking and through pharmacophore based approaches.

Yusuf Doğuş Doğru
Koç University · Institute of Graduate Studies in Science
2015
00
Master'sOpen AccessEN

Gene2Phen-Fenotipe özgü alt etkileşim ağlarının oluşturulması, görselleştirilmesi ve karşılaştırılması için web tabanlı bir araç

Diseases are commonly the result of dysregulated complex interactions involving large sets of genes and proteins as products of these genes, and their cooperation with other cellular components. Interpreting protein-protein interactions at both network and molecular interaction levels with mutation knowledge requires a comprehensive research process that is fed from different sources. In this thesis, we developed a web-based tool, Gene2Phen, by integrating large-scale protein-protein interaction network, 3D protein structure information and interface mutation knowledge to aid researchers in exploring and comparing the molecular mechanism of different phenotypes. Gene2Phen works as an automatized pipeline tool to build, visualize and compare phenotype specific subnetworks, to examine protein- protein interactions associated with their structure and mutation data. Gene2Phen web tool prioritizes the human protein-protein network based on seed genes specific to a phenotype. From the prioritized-PPI network, users can generate a phenotype specific subnetwork. The phenotype-specific subnetworks can be visualized and compared interactively. Genome annotations and topological properties of each protein are shown in this interactive network representation. A unique feature of Gene2Phen is its ability to display 3D structural models of protein-protein interactions and their predicted protein-protein interfaces. Users can see the list of mutations which are mapped on predicted protein-protein interfaces. This allows users to study mutations altering protein-protein interfaces and their role in the phenotype-specific subnetworks. Gene2Phen, by automating the integration of protein-protein networks, protein structure, and disease - related mutations at large scale, will not only boost the productivity and efficiency, but it may be the leveraging step to the novel solutions/studies.

Bilgesu Erdoğan
Koç University · Institute of Graduate Studies in Science
2018
00
Master'sOpen AccessEN

Protein arayüzlerinin yapısal hizalamaları için hızlı eleme algoritmaları

Protein-protein interactions (PPIs) form the basis of many biological processes in living organisms. The significance of PPIs in mediating biological activity necessitates the identification of novel interactions. Template based structural alignment is one of the computational approaches to predict protein-protein interactions using known protein interfaces. One challenge in template-based prediction is the computational cost due to the one-to-all comparison of the query protein against a database of all known interfaces. In this thesis, two different approaches have been developed a) QuickRet, a hashing based algorithm, b) and a deep learning based algorithm. QuickRet, a fast screening algorithm, ranks interfaces due to their structural similarity to a query protein. It extracts features (angles and distances derived from four atoms) from structures of interfaces and compares them with the features extracted from the query protein. QuickRet is tested with the PIFACE database, a clustered protein-protein interface database, and predictions made by the template interface based PPI prediction algorithm, PRISM. The results indicate that QuickRet is successful in filtering structurally dissimilar interfaces for a given protein. With at least 80% match, 99% (320/43500 interface structures remained) of the database is eliminated and the average RMSD value of the remaining structures is 2.4 Å. With at least 90% match, 99.9% (50/43500 structures remained) of the database is eliminated and the average RMSD value drops to 2.28 Å. In addition, a deep learning based method which predicts, for a given protein complex, if the interface between the proteins of the complex is a true interface or not (based on known interfaces in Protein Data Bank). The model, a 3-dimensional convolutional model, analyzes the given structure and outputs the probability of the given structure being an interface. The accuracy of the model for several interface data sets, including PIFACE, PPI4DOCK, DOCKGROUND is approximately 80%. Both algorithms can be used to reduce the computational cost of template-based PPI predictions.

Ali Tuğrul Balcı
Koç University · Institute of Graduate Studies in Science
2018
00
Master'sOpen AccessEN

HMI-PRED: Konak-mikrop protein etkileşiminin tahmini için web sunucusu tasarımı ve geliştirilmesi

Microbes, commensals and pathogens, control numerous functions in host cells. They can alter host signaling and modulate immune surveillance by interacting with host proteins through mimicking the host protein-protein interfaces. To shed light on the contribution of microbes to health and disease, it is vital to discern how microbial proteins rewire host signaling and through which host proteins. Current host-microbe interaction data is a long way from complete, and experimental methods for large-scale identification of HMIs is challenging. Most of the currently available methods for HMI prediction are based on global sequences or structural similarity. On the other hand, there is only one available webserver for these methods, which limits the usage of these tools by the researches and scientists. To address both issues, we developed Host-Microbe Interaction PREDictor (HMI-PRED), a user-friendly webserver for template-based structural prediction of protein-protein interactions (PPIs) between host (i.e., human) and any microbial species, including bacteria, viruses, fungi, and protozoa. HMI-PRED relies on "interface mimicry" through which the microbial proteins hijack host binding surfaces. HMI-PRED server was optimized by clustering the template interface set, which is used as bases for predictions. Given the 3D structure of a microbial protein of interest, HMI-PRED will return detailed 3D structural models of potential host-microbe interaction (HMI) complexes, the list of host endogenous and exogenous PPIs that can be disrupted by the microbe protein, and the functional annotation and tissue expression of the microbe-targeted host proteins. Also, the server offers 3D visualizations of the predicted structures with highlighted contact residues, as well as visualization of the predicted structure superimposed on the original template interface. The server also allows users to upload homology models of microbial proteins. The prediction results are stored in a repository for the community access. Users can examine and search the accumulated results. HMI-PRED is available for public at https://interactome.ku.edu.tr/hmi. We also introduce a ranking method for the predicted interactions using a deep learning model based on 3D structures which is trained to identify valid protein-protein interfaces.

Asma Omar Hakouz
Koç University · Institute of Graduate Studies in Science
2019
00
DoctorateOpen AccessEN

Protein-protein etkileşim ağlarının alternatif konformasyonlarla zenginleştirilmesi ve farklı arayüzlerin hızlandırılmış filtrelemesi

The common practice in structural protein-protein interaction (PPI) networks is to investigate just one specific conformation for each protein. Yet it is not a comprehensive representation as it neglects the conformational changes of proteins which may lead to different protein interactions, functions, and downstream signaling. In this dissertation, a new representation is proposed for structural PPI networks which inspects the alternative conformations of proteins. This representation uses a method to get all available structures of proteins from protein data bank (PDB) and clusters them based on their sequence and structural similarities. Then, it investigates the alternative conformations of each protein. A large-scale study is done by creating breast cancer lung and brain metastasis sub-networks and equipping them with alternative conformations of the proteins. PPI network analyses showed novel genes and cellular pathways which play important roles in each sub-network. By examining alternative conformations of proteins, the docking results coverage increased from 54% to 76%. In addition, the effects of conformational changes on specific interactions are shown. The conformational changes of KPNB1 directs it to bind to SNAI1 and SNUPN in the open and close conformations, respectively. Exploring the alternative conformations of CXCL12 shows that a point mutation on the binding surface can inhibit CXCL12 homodimerization, as a complex structure, and alter its functions. PRISM is used for protein docking purposes in the sub-networks. The primary limitations of PRISM are its long running time and restricted interface database. These limitations derive mainly from the all-to-all protein surface and interface structure comparison process which is very time-consuming. So, to speed up this process, a new filtering method is proposed which creates 1D descriptors of protein structures to rapidly filter the dissimilar interface. This method uses a library of non-redundant small protein fragments and Geometric hashing technique to create the protein descriptor vectors. Based on the results, thousands of interface comparisons could be done in seconds using these descriptor vectors. 70%-80% of grossly dissimilar interfaces could be filtered rapidly and the remaining candidate similar interfaces would be compared by an alignment method.

Farıdeh Halakou
Koç University · Institute of Graduate Studies in Science
2020
00
Master'sOpen AccessEN

HotRegion v2.0: Protein-protein etkileşim arayüzlerindeki sıcak bölgeleri tahmin etmek için yeni bir yöntem

Proteins interact with each other through their interface to fulfil essential functions in the cell. The study of protein interactions will have a profound effect on understanding various biological pathways. Binding free energies are not uniformly distributed among the residues found in protein-protein interfaces (PPI). Hot regions are tightly packed residue clusters in PPIs that account for the majority of the binding free energy of proteins, and hence, are crucial for the stability of complexes. Providing specificity to binding sites, these regions are of great importance for drug discovery in pharmaceutical research. Experimental discovery of hot regions is time-consuming and requires high effort. Hence, there is a need for computational methods. The existing hot region prediction algorithms perform clustering on computationally predicted hot spots and ignore the potential existence of non-hot spot residues in hot regions. However, experimental studies have demonstrated that non-hot spot residues can also be found in hot regions. In this thesis, using unsupervised learning approaches, we propose a novel method to predict hot regions which may contain both hot spot and non-hot spot residues. We combine affinity propagation (AP) and density-based spatial clustering of applications with noise (DBSCAN) to cluster interface residues and develop a web-based tool to predict and visualize hot regions. Furthermore, we have developed a database that contains hot region information for more than 600.000 protein complexes. In our web server, users can query data from our database or submit a new run to obtain hot regions of a complex that is not found in the database. Our tool demonstrates a 3D structural visualization of the complex with colored hot regions and highlighted hot spot residues. Structural features of interface residues, i.e., accessible surface area values and knowledge-based pair potentials, are also indicated. In order to evaluate our method, we have compared our results with the experimental studies. The precision of our algorithm is 0.68 and the accuracy is 0.62. Additionally, we have conducted a case study to test the significance of our predicted regions for clinical studies. We have investigated the complex of human programmed death-1 (PD-1) and its ligand PD-L1. Our algorithm correctly identified the residues which are known to be significant for PD-L1-antibodies and small-inhibitors. Lastly, we have compared our algorithm with our previous hot region prediction method. The results have shown that our tool outperforms the previous version of HotRegion and may be a leveraging step to novel pharmaceutical studies.

Damla Övek
Koç University · Institute of Graduate Studies in Science
2020
10

Other supervisors