DoktoraAçık Erişim

New machine learning algorithms and applications to drug design

2008
0 görüntülenme
0 i̇ndirme
Danışman: Prof. Dr. Okan Ersoy ; Prof. Dr. Oya Kalıpsız

Özet (EN)

Machine learning includes the algorithms which aim to find the best-fit model to the data. In this thesis, several machine learning algorithms are developed. Nowadays, the data need to be investigated is growing exponentially. Therefore, machine learning is needed in all sectors. Drug design is selected as the application area of the thesis.Drugs are very useful to maintain good health. This is why drug design very important and necessary. Drug design is also very costly, and requires much effort and time to develop. Because of heavy cost, only some large companies have the capability to work in this area. In a report by TUBITAK, drug design has been declared a research direction of high priority for Turkey during 2003 ? 2023 planning period.One of the important components of drug design and cost is the process of choosing potential drug molecules. Those choosing operations usually contain one or more of classifying, clustering, feature selection, regression problems. By using machine learning methods, that process time and cost involved in these operations, can be minimized.In drug design problems, all subjects in machine learning are almost used. So, in this thesis, a study that contains most such machine learning topics has carried out.For classification problems, a new algorithm family called Cline has been designed. These algorithms are decision tree induction algorithms. Although the algorithms are simple, experiments have shown that they have competitive performance as compared to existing algorithms on UCI and drug datasets.In previous studies, it has been observed that classifier committees have more successful results than single classifiers?. Also, in this study, similar results have been obtained. So Cline decision forests have been added to Cline algorithms family. Cline decision forests have produce better performance than existing algorithms on UCI and drug design. Those Cline decision tree and decision forest algorithms have been serviced to end users via Cline Toolbox application and web site of the thesis owner.For feature selection problems, an approach that uses decision tree and decision forest has been developed but needs further improvement.For clustering problems, an algorithm family called Clusline has been developed and some satisfying results have been achieved as compared to other existing algorithms.Clustering committees have been developed, inspired by the excellent performance of classifying committees. Different decision combination techniques of clustering committees in the literature have been investigated and compared. Our study is one of the most comprehensive studies, and our results can be used as a guideline for researchers.For regression problems, an approach based on data clustering in subspaces has been developed but needs further improvement.For regression committees, the effects of algorithms used in commitee, different generating and decision combination techniques of committees in the literature have been investigated and compared. The experiments on drug design datasets show that the usage of committees for regression problems gives more or less similar performance with single regressors.There is no global algorithm that always gives better result than other algorithms for all datasets. So, trial and error is the method to find the best algorithm for a dataset. To eliminate this lack and provide helper rule series to end users, approaches that aim to estimate algorithms? performance according to the meta features of datasets have been developed. Those approaches are called meta learning. Current meta learning studies are usually for classification problems. Given that most problems in drug data design are regression problems, in this study a new meta regression approach has been developed. By this approach, new meta features have also been extracted in addition to standard features used in meta learning. So, a new model has been developed that can estimate algorithm performance for a dataset by using various dataset features. In this way, the best performable algorithm for a dataset can be estimated in advance. Besides, clustering datasets and algorithms according to similarities have also been investigated.Consequently, in this thesis, many new approaches about various machine learning subjects have been developed and useful results for researchers and end users have been produced. It is our wish that this thesis contribute to studies about both drug design and machine learning research areas in Turkey and the world.

Yazar

Mehmet Fatih Amasyalı

Bu Yayına Nasıl Atıf Yapılır

Mehmet Fatih Amasyalı (Doctorate thesis). New machine learning algorithms and applications to drug design, 2008, Yıldız Technical University, Bilgisayar Mühendisliği Bölümü.

Lisans

Tüm Hakları Saklıdır

Bu eser belirtilen lisans koşulları altında paylaşılmaktadır.

Yıldız Technical University tezlerinden daha fazlası