Middle East Technical University
Discipline

Sağlık Bilişimi Anabilim Dalı

Middle East Technical University

5

Archived Theses

0

DOIs Assigned

0%

DOI Rate

Discipline

5 Theses
Master'sOpen AccessTR

Biyoinformatik teknikleri kullanarak yeni mikro RNA'ların bulunması ve varyant analizlerinin yapılması: Citrus modeli

Amaç: Bu çalışmanın hedefi mevcut YND verilerini kullanarak, biyoinformatik araçlar yardımı ile yeni miRNA'ların doğru olarak tanımlanması, tanımlanan miRNA'ların farklı türler arasındaki korunmuşluğunun incelenmesi, korunmuş miRNA'ların hedef genleri üzerindeki etkilerinin ve metabolik yolaklardaki rollerinin belirlenmesidir. Yöntem: Gerek duyulan tüm genom ve küçük RNA dizileme verileri mevcut olduğu için, Citrus (Narenciye) çalışmamızda model organizma olarak kullanıldı. Ön işlem ve filtreleme basamaklarından geçen veriler, The UEA sRNA Workbench yazılımı kullanılarak miRNA değerlendirmesine tabi tutuldu. tRNAscan-SE ve NOVOMIR yazılımları ile yanlış pozitif olarak değerlendirilen miRNA'lar çalışmadan çıkarıldı. 34 farklı Citrus türüne ait genom dizileme verileri ve TNV verileri kullanılarak, bu türler arasında ortak olan ve dizisi korunmuş miRNA'lar tespit edildi. Yeni Citrus miRNA'larının tespiti için dizisi korunmuş olan miRNA'lar, veri tabanlarından ve literatür taraması sonucunda elde edilen miRNA'lar kullanılarak filtrelendi. Korunmuş miRNA'ların hangi genlerin ifadelenme seviyesini düzenlediği ve bu hedef genler üzerindeki etkileri psRNAtarget yazılımı kullanılarak belirlendi. Korunmuş miRNA'ların metabolik yolaklardaki rollerinin incelenmesinde, hedef genler ve CitrusCyc veri tabanından elde edilen iki farklı Citrus türüne ait metabolik yolaklar kullanıldı. Bulgular: Çalışmamızda toplamda 1504 miRNA tespit edildi. 270 miRNA'nın 34 farklı Citrus türleri arasında ortak ve 111 miRNA dizisinin bu türler arasında korunmuş olduğu tespit edildi. 111 miRNA'nın 36 tanesi ilk kez bu çalışma ile literatüre kazandırıldı. Başka bir araştırma grubu tarafından literatüre kazandırılan Citrus isomir'lerinin, çalışmamızın filtreleme basamaklarında elimine edildiği saptandı. Citrus türleri arasında korunmuş miRNA'ların, farklı türlerde farklı hedef genlere ve dolayısıyla da farklı metabolik mekanizmalara etki ettiği tespit edildi. Sonuç: Bir biyoinformatik uygulaması olan bu çalışmada, miRNA dizi analizine genomik veriler dahil edilerek miRNA alanındaki çalışmalara ve tanımlara farklı bir bakış açısı getirilmiştir. Anahtar Sözcükler: mikro RNA, yeni nesil dizileme, biyoinformatik.

BiyoinformatikBiyoinformatikDizi analizi+7
Cankut Çubuk
Dokuz Eylül University · Institute of Health Sciences
2019
00
Master'sOpen AccessEN

Tekil amino asit mutasyonlarının protein işlevleri üzerindeki etkisinin yapısal ve anotasyon odaklı yaklaşımla tahmini

Whole-genome and exome sequencing studies have indicated that genomic variations may cause deleterious effects on protein functionality via various mechanisms. Single nucleotide variations that alter the protein sequence, and thus, the structure and the function, namely non-synonymous SNPs (nsSNP), are associated with many genetic diseases in human. The current rate of manually annotating the reported nsSNPs cannot catch up with the rate of producing new sequencing data. To aid this process, automated computational approaches are being developed and applied on the unknown data. In this study, we propose a new methodology to collect and organize the information related to the effects of nsSNPs at the amino acid sequence level from various biological databases and to utilize this information in a supervised machine-learning based system to predict the function disrupting capacities of mutations with unknown consequences. For this, 157,138 annotated mutation data points (89,363 deleterious and 67,775 neutral) were collected from multiple resources such as UniProt, ClinVar and Protein Mutant Database. For each mutation data point, a feature vector was constructed using protein 3-D structure information and site-specific feature annotations in the UniProt database. The information about the spatial proximity of the reported mutations to these protein features were also incorporated to the feature vector. The system was trained with these feature vectors and their respective labels in a supervised fashion using random forest, where the ultimate aim was to construct a model that classifies unknown mutations either as deleterious or neutral. The prediction model was evaluated in detail to observe the contribution of different feature types to the prediction success. The finalized model displayed a satisfactory performance (AUROC:0.86, precision: 0.77, recall 0:90, accuracy: 0.78, F1-score: 0.83 and MCC: 0.54) on the independent test dataset. Besides, the performance of the proposed model was compared to the widely used variant effect predictors in the literature, over standard benchmark datasets. As future work, we plan to conduct a case study over interesting prediction examples and to validate our results via literature-based information. Finally, we plan to construct a ready-to-use command line based variant effect prediction tool and to share it with the research community over an open access data repository. We believe that this system will be complementary to the well-known methods in the literature and its incorporation to ensemble-based tools will increase the performance of the state-of-the-art in variant effect prediction.

Fatma Cankara
Middle East Technical University · Enformatik Enstitüsü
2020
00
Master'sOpen AccessEN

Matris factorizasyonu yöntemi ile biyolojik veri entegrasyonu ve ilişki tahmini

The available molecular sequence data has increased greatly in the last decades, thanks to the new technological developments in the field of life-sciences. In order for this data to be useful to the scientific community, it should be characterized. Traditionally, this characterization is done manually, where the experimentally produced molecular data is curated and stored in the biological databases. The huge volume of the currently available data summons the need for the automatic and systematic analysis. A crucial part of this systematic analysis is data integration with the identification of the relationships between the elements from different biological data types. In this study, we propose to integrate large-scale gene/protein annotation data by using non-negative matrix factorization (NMF), which is a frequently used method for recommender systems with successful real-world applications. NMF has also been employed for uniting multi-relational data in many different fields including bioinformatics and cheminformatics. Within the purposes of this study, we first collected protein annotations such as molecular functions, biological processes, sub-cellular localizations and disease relations from different resources such as UniProt-GOA and DisGeNET, and organized them as binary relation matrices. We then applied various NMF-based algorithms to this multi-dimensional relational biomolecular sequence annotation data (i.e. genes/proteins vs. functions, genes/proteins vs. diseases, diseases vs. functions) and evaluated the results of each model in terms of their capacity to learn the intrinsic structure in relational data, via cross-validation. The results indicated that NMF has the capacity to retrieve most of the known protein annotations without using any sequence or structure-based protein features (AUROC: 0.80 – 0.94, accuracy: 0.53 – 0.64, F1-score: 0.06 – 0.40, MCC: 0.13 – 0.38). Using NMF, the ultimate aim here is to predict the unknown binary relationships between these biological entities; and to represent these entities (i.e., proteins, functions and disease entries) as informative and non-redundant quantitative feature vectors (using the low-rank feature matrices generated by the factorization process), which can be used in diverse data mining and machine learning tasks in the future, such as the automated annotations of proteins or the construction of biological knowledge graphs.

Gökçe Abay
Middle East Technical University · Enformatik Enstitüsü
2020
00
DoctorateOpen AccessEN

Motif keşfi için insan uçbirleştirme akseptör bölge sekanslarının genom çapında analizi

For eukaryotic cells, alternative splicing of genes is a vital mechanism that drives protein diversity. Splicing signals on the genomic sequence controls the regulatory factors that orchestrate the alternative splicing. 3' and 5' splice sites and common branchpoint sequences are the primary splicing signals, and changes in these signals can be disease- causing. Nevertheless, an extensive genome-wide analysis of the sequences around these signals is lacking. In this study, we focused on the genome-wide motif analysis of the splice acceptor region. We analyzed 400 nucleotides long sequences (300 nucleotides upstream and 100 nucleotides downstream of 3') to identify motifs with potential functional roles. 207,583 sequences are retrieved from Ensembl Biomart and analyzed with MEME ChIP, resulting in 517 significant splice acceptor region motifs. We identified 457 known motifs and 60 novel motifs. Among the known motifs, 227 mapped to non- human mammalian genomes. Furthermore, proteins binding to the known motifs are mainly annotated for homeoboxes, homeodomains, DNA binding regions, and transcription regulation functions. 17 of the novel motifs comply with RBP binding motifs, and 10 of the novel motifs are computationally identified and supported with experimental evidence from branchpoint studies. Moreover, the acceptor region splice altering or disease-causing variants with experimental evidence are detected to co-locate with novel motifs and known motifs of homo sapiens and other mammalians. Here, we present these novel acceptor region motifs identified for the first time as splice acceptor motif candidates. Furthermore, we provide a set of 76 known mus musculus motifs that are novel to the human genome and highly co-locate with splice-altering SNPs. Experimental validation of the biological roles of novel motifs of this study with further functional studies will increase our understanding of the splicing mechanisms.

Gülşah Karaduman Bahçe
Middle East Technical University · Enformatik Enstitüsü
2020
00
Master'sOpen AccessEN

İlaç-ilaç etkileşimleri analizi için genotip kataloğu

Polypharmacy is an essential practice in today's therapeutics, especially in the care of older population. Most polypharmacy-induced drug-drug interactions (DDIs) are often discovered after drugs are put on the market. Health problems and economic burden due to unpredicted DDIs put the health system in a difficult situation. Therefore, increasing the predictability of DDIs has become one of the most critical concerns towards improving treatment success. Being dependent on several underlying parameters makes DDIs challenging to foresee. One of the most significant determinants of these underlying parameters is genetic variability. Therefore, a more profound knowledge of the DDI-genetic relationship, called drug-drug-gene interactions (DDGI), will improve treatment success. This study aims to design a relational database named DDGICat, which was designed to contribute to the ongoing DDGI research by serving as a guideline to prescribers, researchers, and the pharmaceutical industry. The content of the DDGICat was derived from other knowledge bases, including DrugBank, PharmGKB, Ensembl, KEGG Drug, and ONC High. DDGICat contains drugs, drug target proteins, drug-associated SNPs, DDIs, drug-gene interactions (DGIs), and DDGIs. Additionally, a developed web portal named DDGICat Browser provides the results of mentioned content in both tabular and graphical formats. Furthermore, the end products of this study are tested in a case study on a chosen disease.

Ayşe Özdemir
Middle East Technical University · Enformatik Enstitüsü
2021
00