Primer ve sekonder yapılar kullanılarak proteinlerin fold düzeyinde sınıflandırılması ve motif çıkarımı
2015
0 views
0 downloads
Advisor: Prof. Dr. Zümray Dokur Ölmez
Abstract (EN)
Proteins are crucial molecules in biological phenomena because they form much of the functional and structural machinery in every cell in organisms and their function is determined by their spatial structures. Protein structures can be described at various levels in detail, ranging from atomic coordinates, through vector approximations, to secondary structure elements. Protein structure comparison is an important issue that helps biologists understand various aspects of protein function and evolution. It is commonly believed that the 3D fold has a major effect on the ability of a protein to bind other proteins or ligands. The similarity analysis of protein structure is therefore an important process in understanding the protein's role in the machinery of life. Comparison of protein structures is also essential for estimating the evolutionary distances between proteins and protein families. Protein fold classification is also an important problem in bioinformatics and a challenging task for machine-learning algorithms. According to convention a protein could be classified into one of four structural classes based on its secondary structure components; all-α, all-β, α/β, α + β. Structural Classification of Proteins (SCOP) provides a detailed and comprehensive description of the structural and evolutionary relationships among all proteins whose structures are known. According to SCOP four structural classes are divided into folds. Protein fold classification problem is to determine that the query protein belongs to which fold. In this thesis we deal with two problems related to proteins; protein fold classification and structural block comparison (motif retrieval). Proteins are formed by two basic regular 3D structural patterns called secondary structures; helices and strands. A structural motif is a compact 3D protein structure referring to a small specific combination, which appears in a variety of molecules. In this thesis, primarily protein fold classification problem is employed. For the classification of protein folds, neural network based three methods are used; Grow and Learn (GAL) network, Self-Organizing Maps (SOM) and Self-Organizing Maps for Structured Data (SOM-SD). For GAL and SOM primary protein structures are used, on the other hand for SOM-SD secondary protein structures are used. Firstly GAL method is used to classify the protein folds. Here, six attributes which are physicochemical features of amino acids (amino acid composition, predicted secondary structure, hydrophobicity, normalized van der Waals volume, polarity and polarizability) are used as features. A number of proteins are selected from Protein Data Bank (PDB). Then, 27-class protein fold classification problem is tried to be solved with this method. To increase the success rate one-versus-others (OvO) prediction method is used. Secondly SOM is used to classify the protein folds. Features and proteins in the previous method are used also in here. As in the previous method, OvO method is applied for performance evaluation. Thirdly SOM-SD method is used for protein fold classification. While using SOM-SD, Protein Gaussian Image (PGI) representation of proteins is used as feature. PGI is a representation in the Gaussian sphere in which each secondary structure is mapped with a unit vector from the origin of the sphere having the orientation of the secondary structures. The chain sequence of secondary structures is recorded as a list which is mapped on the sphere surface. To test this method the dataset including three folds with 45 proteins (15 proteins in each fold) from PDB is used. To determine the effectiveness of the attributes some tests were made using GAL. Firstly, only C (amino acid composition) attribute was used to be contained in the feature vectors. Then S (predicted secondary structure) attribute was appended to C, so C+S was used to be the elements of the feature vectors, progressively in the last set all six attributes were used and tested by using GAL. The test results showed that the most important attribute is the amino acid composition. This attribute has a good performance even tested alone. Besides in here, for reducing dimension of the feature vector without changing success rate divergence analysis was applied. This analysis calculates divergence values of the features and put them in order according to their importance. After this analysis the most significant 30, 40, 50 and 60 features were determined and they were tested with GAL. The results related to protein fold classification problem showed that proteins are classified according to their folds with a good precision and the results are comparable to the existing methods in the literature. In this thesis after protein fold classification problem, motif retrieval problem is handled. Here, a particular motif is retrieved from a particular protein using structural block comparison. To do this, three methods based on Generalized Hough Transform (GHT) are used. The first method uses single secondary structure, the second one uses secondary structure couple co-occurrences and the third one uses secondary structure triplets. For all three methods the barycenter (geometric mean) of the motif is assigned as Reference Point (RP) and in order to determine this point a mapping rule is figured out. Then, voting process is applied and the point having maximum number of votes is assigned as the candidate RP. For the test, a few proteins selected from PDB are used and the test results showed that the RP is determined with a good precision and the motif is retrieved from the protein with expected number of votes.
Author
Dr. Özlem Polat
Institution

Istanbul Technical University
Elektronik Mühendisliği Bilim Dalı
How to Cite
Özlem Polat (Doctorate thesis). Primer ve sekonder yapılar kullanılarak proteinlerin fold düzeyinde sınıflandırılması ve motif çıkarımı, 2015, Istanbul Technical University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Istanbul Technical University
- Investigation Of Stretching Effect With Mixed Finite Element Formulations For Laminated Beams And Plates(2023)
- Classification of anemia using data mining methods: An application(2015)
- Removal and recovery of platinum group metals through anode slimes of moebius electrolysis(2015)
- A study of design approaches to Istanbul's city halls based on space syntax theory(2015)
- A II. German Empire project: From Kaiser Wilhelm Monument to German fountain(2015)
- Uzaktan algılama verilerinin yersel ölçümlerle entegrasyonu ile toprak tuzluluk haritalaması; Aşağı Seyhan Ovası, Adana, Türkiye(2015)