
171
Archived Theses
0
DOIs Assigned
0%
DOI Rate
Discipline
Meta analizi: Paralel kontrollü çalışmalarda bir uygulama
Bu çalışmada, öncelikle meta analizi kavramı, tarihsel gelişimi, amaçları, avantajları ve dezavantajları, aşamaları, analizde kullanılan kavramlar, modeller ve meta analizi yöntemleri hakkında teorik-kuramsal bilgilere yer verilmiştir. Bununla birlikte, hayvancılık alanında yumurtacı tavuklarda probiyotiğin yumurta ağırlığına etkisi konusunda yapılmış olan çalışmalara yönelik meta analizi gerçekleştirilmiştir. Çalışmada, sürekli verilerde ortalamaları kullanarak etki büyüklüğünün hesaplanması yöntemi kullanılmış olup, ilgili literatür çerçevesinde toplanan deney-kontrol modelli 9 çalışmaya ait etki büyüklükleri, yayın yanlılığı, heterojenlik ve rastgele etki modeli ile genel etki hesaplanmıştır. Çalışma sonucunda, araştırmalar arasında yayın yanlılığı olmadığı tespit edilmiş olup, uygulanan heterojenlik testine göre tüm çalışmaların tek bir yaygın etkiyi paylaşmadığı (heterojen olduğu) belirlenmiştir. Ayrıca, genel etki büyüklüğü, yumurtacı tavuklarda probiyotiğin yumurta ağırlığı üzerinde anlamlı bir etkisinin olduğunu göstermiştir. Anahtar Kelimeler: Meta analizi, Etki büyüklüğü, Sabit etki modeli, Rastgele etki modeli.
Türkiye'deki illerin hayvancılık istatistikleri bakımından çok değişkenli analiz teknikleri ile incelenmesi
Bu araştırmada, kümeleme analizi, faktör analizi ve çok boyutlu ölçekleme analizi gibi çok değişkenli istatistiksel tekniklere teorik anlamda yer verilmiştir. Bununla birlikte, TÜİK'den elde edilen 2018 yılı hayvancılık istatistiklerine ait toplam 29 değişkeni içeren hayvan varlığı ve hayvansal üretim verileri çok değişkenli analiz teknikleri ile incelenmiştir. Çalışmada öncelikle kümeleme analizine yer verilmiş, en uygun küme yapısı hiyerarşik kümeleme yöntemlerinden Ward yöntemi ile elde edilmiş ve iller 2 kümeye ayrılmıştır. Hiyerarşik olmayan kümeleme analizlerinde uygun küme sayısı, bazı parametrelerin yanı sıra küme üyelik kodları kullanılarak diskriminant (ayırma) analizi ile belirlenmiş olup, fuzzy (bulanık) kümeleme yönteminin 5 küme yapısında en etkili yöntem olduğu sonucuna varılmıştır. Ayrıca, çalışmada faktör analizinden elde edilen faktör yükleri kullanılarak kişi başına düşen hayvansal üretim ve hayvan varlığına ait ham veriler ve standartlaştırılmış veriler ile illerin hayvancılık gelişmişlik endeksleri hesaplanmıştır. Her iki sıralamada da ilk sıradaki ilin Ardahan son sıradaki ilin ise İstanbul ili olduğu görülmüştür. Çalışmada son olarak metrik çok boyutlu ölçekleme tekniği kullanılmıştır.
Tip II regresyon analizine yapay sinir ağı yaklaşımı
Bu çalışmada, Tip II Regresyon tekniklerinden olan EKK Açıortay tekniğine ilk defa bu çalışma aracılığıyla Yapay Sinir Ağı (YSA) yaklaşımı yapılmıştır. Yeni oluşturulan bu YSA Açıortay tekniğinin performansının ölçülmesi amacıyla da EKK Açıortay tekniğiyle karşılaştırılmıştır. Öncelikle YSA ve Tip II Regresyon teknikleri konularında literatür bilgilerine yer verilerek, iki tekniğinde özelliklerinden bahsedilmiştir. Bu bilgiler doğrultusunda EKK tabanlı açıortay tekniği ile YSA tabanlı açıortay teknikleri arasında karşılaştırma yapılmıştır. Bu iki tekniğin karşılaştırılması amacıyla farklı dağılımlarda ve farklı örneklem hacimlerinde modellemeler yapılmıştır. Bu modellerin performanslarının karşılaştırılması amacıyla "Ortalama Mutlak Yüzde Hata" (MAPE) kriterinden yararlanılmıştır. Çalışma sonucunda YSA tabanlı Açıortay tekniği EKK tabanlı Açıortay tekniğine göre daha düşük hata ile daha iyi sonuç verdiği görülmüştür. Yapılan bu çalışma ile gelecek dönemde bu alanlarda çalışmak isteyen araştırmacılar için bir örnek temsilinde olduğu ön görülmektedir.
Türkiye'deki enerji santrallerinde doğal gaz tüketiminin destek vektör regresyon ile tahmini
Bu çalışmanın amacı, destek vektör makinelerinde regresyon yöntemini kullanarak Türkiye'deki enerji santrallerinin doğalgaz tüketimi üzerine ön kestirim yapmaktır. Bu amaçla, araştırmada kullanılacak veri seti 2013-2018 yılları arasında Türkiye Enerji Piyasası Düzenleme Kurumu ve Enerji İşleri Genel Müdürlüğünden elde edilmiştir. Bu çalışmada ilk olarak, Türkiye'de doğalgazın enerji piyasasındaki yeri, birincil enerji kaynakları içindeki payı, üretim, tüketim, ithalat ve ihracat değerleri incelenmiştir. Bu değerlerin ölçü birimlerinin farklılığından dolayı, ilgili veri seti istatistiksel analizden önce standartlaştırılmıştır. Enerji santralleri tüketimi (bin sm3) bağımlı değişken iken; sanayi tüketimi (bin sm3), şehir tüketimi (bin sm3), üretim (milyon sm3), ithalat (milyon sm3) ve ihracat (milyon sm3) değerleri bağımsız değişken olarak belirlenmiştir. Destek Vektör Regresyonda kullanılan tüm çekirdek fonksiyonları (Doğrusal, Polinomsal, Radyal Tabanlı Fonksiyon(RTF) ve Sigmoid) test edilmiştir. En küçük Hata Kareler Ortalaması (HKO)'na sahip olan RTF kestirim çekirdek fonksiyonu olarak seçilmiştir. Daha sonra, destek vektörler, ağırlıklar ve karar sabiti belirlenmiştir. Ağırlıklar ve destek vektörler çarpılıp yan eklenerek, son model elde edilmiştir. Son model yardımıyla da Mayıs - Aralık 2018 için Türkiye'deki enerji santrallerinin doğalgaz tüketimlerine ait tahminler yapılmıştır.
Bayes Açıortay Regresyon Tekniği ve bir uygulama
Bu çalışmanın amacı Bayes Tip II regresyon tekniğinin performansını incelemektir. Bu amaçla gerçek bir veri seti üzerinde Bayes yaklaşımı yardımı ile basit doğrusal regresyon ve açıortay regresyon denklemleri hesaplanmıştır. Daha önceki çalışmalarda Tip II regresyon teknikleri arasında en iyi performansı sergileyen tekniğin açıortay tekniği olarak belirtilmesinden dolayı mevcut veri seti için sırasıyla X ve Y değişkenleri bağımlı değişken olarak ele alınarak regresyon denklemleri elde edilmiş, daha sonra elde edilen bu iki regresyon denkleminin açıortayı alınarak Bayes açıortay denklemi hesaplanmıştır. Önsel ve ek bilgiye dayalı sonsal dağılımları elde etmek amacıyla farklı örneklem hacimlerinden yararlanılmış ve Bayes regresyon denklemlerinin performansları HKO ve GE kriterlerine göre karşılaştırılmıştır. Araştırma bulgularına göre n=100, n=75 ve n=50 birimlik örneklemlerde Bayes açıortay tekniğinin performansının en düşük HKO değerine sahip olduğu, dolayısıyla mevcut veri setine ait bu örneklem hacimleri için en iyi performansı sergilediği belirlenmiştir.
Görüntü sınıflandırma için derin öğrenme ile Bayesçi derin öğrenme yöntemlerinin karşılaştırılması
Temel evrişimli derin öğrenmede ağ mimarisi oluşturulurken iç katman sayısını belirleyen ağın derinliğinin ve ağın eğitimi öncesinde öğrenme oranı, momentum ve L2 düzeltmesi gibi öğrenme parametrelerinin başlangıç değerlerinin belirlenmesi gerekir. Bu da çözülmesi gereken ayrı bir optimizasyon problemidir. Bayesçi derin öğrenme ağ mimarisini ve öğrenme parametrelerinin uygun başlangıç değerlerini bulmak için Bayesçi optimizasyon tekniklerini kullanır. Bu tez çalışmasında, temel evrişimli derin öğrenme ile Bayesçi derin öğrenme popüler bir görüntü sınıflandırma problemi üzerinde karşılaştırılmıştır. Her iki yöntemin birbirine göre performansları, avantajları ve dezavantajları ilgili sınıflandırma problemi üzerinde değerlendirilmiştir. Bayesçi derin öğrenme, test veri kümesi üzerindeki sınıflandırma performansını önemli bir derecede arttırmasına rağmen ağın yapısının ve öğrenme parametrelerinin başlangıç değerlerinin optimizasyonu ek bir zaman maliyeti oluşturmuştur.
Büyük veride hiyerarşik kümeleme yöntemlerinin kofenetik korelasyon ile karşılaştırılması
Bu çalışmada, öncelikle büyük verinin tanımı, büyük verinin bilişenleri, büyük veri analitiği ve büyük veri teknolojileri hakkında teorik-kuramsal bilgilere yer verilmiştir. Bununla birlikte kümeleme analizi, kümeleme yöntemleri, kümeleme yöntemi uzaklık ölçütleri ve Kofenetik korelasyon katsayısı hakkında da teorik-kuramsal bilgiler yer almaktadır. Devamında ise büyük veri teknolojilerini kullanarak büyük veride hiyerarşik kümeleme yöntemleri Kofenetik korelasyon katsayısı karşılaştırılmıştır. Veri analizi için açık kaynaklı büyük veri teknolojilerini içeren Amazon bulut sunucusu kullanılmıştır. Sunucu üzerine Python programlama dili kurulmuş ve analiz sürecinde Python için geliştirilmiş kütüphaneler kullanılmıştır. Çalışmada ABD Ulaştırma Bakanlığı tarafından yayınlanan 2015 Hava Seyahat Tüketici Raporundaki veri seti kullanılmıştır. Çalışmanın sonucuna etki etmeyecek veri setindeki değişkenler, analiz süreçlerini uzatabileceğinden özellik seçim işlemi ile çıkartılmıştır. Sonrasında, boş gözlemler temizlenmiş ve veriler standardize edilmiştir. Ardından, veri seti içerisinden ana kütleye temsilen rastgele seçim yöntemiyle 4 farklı veri seti oluşturulmuştur. Bu veri setlerine kümeleme analizi uygulanmıştır. Yapılan analizler sonucunda tüm veri setlerinde Kofenetik korelasyon katsayısının, ortalama bağlantı kümeleme yönteminde en yüksek değeri sağladığı gözlemlenmiştir. 2020, ix + 50 sayfa Anahtar Kelimeler: Kofenetik korelasyon, Büyük veri, Kümeleme analizi.
Investigation of text mining methods on Turkish text
Today, with the widespread use of the internet the size and value of the data have increased. Making meaningful information from large amounts of data is one step ahead for others and for companies. Various mining techniques must be applied to obtain meaningful information from the data. Data mining processes on structured data. Text mining is the sub-study area of data mining that works on texts. Text mining is the process of analyzing text to extract valuable information from text for special purposes. Before the text mining techniques are applied, the data must be prepared and pre-processed. Extracting meaningful information from texts, classifying text, and reaching the desired information in a short time increase the importance of text mining. Text classification is the process of deciding the class of the given text using the training documents for the predefined categories. The aim in thesis is to classify Turkish data as text. The categories were examined in three categories as "Gender Identification", "Author Identification" and "Species Determination". Naive Bayes method and the bit-score weighting k-NN method were used for classification. The accuracy rates of the two methods are compared. The R programming language is used for classification. In this thesis, a dataset consisting of Turkish columns was created to work on text classification.
Local dependence measures, properties and applications
Estimation of the parameters in the balanced incomplete block design
ABSTRACT In this study, the methods which are used for estimating the treatment effects in the balanced incomplete block design and the estimated values of the treatment effects obtained by using these methods have been compared. When it is impossible to make the required number of treatments, which are needed for each block in the randomized complete block design, the experiment is designed in the balanced incomplete block design. In the balanced incomplete block design, it is suggested that other methods should be used rather than the least squares estimators to estimate the treatment effects. Therefore, in order to estimate the treatment effects in the balanced incomplete block design, the intrablock, the interblock and the combined estimates methods are introduced in literature. In this study, the simulation studies were made by using the programs written in statistical software Minitab. The intrablock, the interblock and the combined estimates for the balanced incomplete block design and the least squares estimates for the randomized complete block design of the treatment effects were calculated 2500 times by simulation. The results were compared and the most appropriate estimator of the treatment effects for the balanced incomplete block design was investigated. The means and the standard deviations of the estimates were considered as the criteria of this comparison. After the evaluation of the results, it was observed that each of the three methods gave unbiased results when the block effects were insignificant for the balanced incomplete block design. When the block effects were significant, it was seen that the results of the means of the intrablock and the combined estimates were unbiased. The interblock estimates results were observed as inappropriate. In both of the situations in which the block effects were significant and insignificant, the standardVI deviation of the intrablock estimates is lower than the standard deviations of the interblock and the combined estimates.
Principal components in the problem of multicollineartity
ABSTRACT In this study, principal components regression and ridge regression are examined among the methods used to remedy multicollinearity problem in multiple linear regression model. One of the assumptions in multiple linear regression is that there must be no perfect linear relations among the regressors. The relationship among the regressors is called multicollinearity. In case of multicollinearity, parameter estimations by least square method have large variances and hypothesis tests result in contradictory. There are various methods for dealing with multicollinearity problem. Biased regression methods (BRM) are the ones that can explain the structure of multicollinearity and provide small standard errors among the methods used. In this study two of biased regression methods; principal components regression and ridge regression are examined as theoretically and researched which methods give the best consequence by simulation. In the application, 50 repetitions have been generated for each of the sample sizes of 40, 80 and 120. Least squares, ridge and principal components regression are used for each sample. Regression coefficients for each estimator were computed and the mean and the standard deviation of the estimates were used as statistical comparison criteria. According to comparisons among the estimators the principal components regression has been found to provide better estimates.
Survival models and an application
ABSTRACT The objective of this study is to examine the survival models and describing the data obtained in a particular field by a suitable model. In the first section, the goal of the study was described by comparing the methods used in survival analysis. The second section gives general information about the data types and the functions used in survival analysis. In the third section, Exponential, Gompertz, Makeham, Gamma, Weibull and Lognormal distributions which are the parametric methods used frequently in survival analysis, are examined in order. In the fourth section, the nonparametric methods, which are Kaplan-Meier product limit and Life table methods, were investigated. In the fifth section of the study, the data obtained were grouped in a convenient way and examined by using Kaplan-Meier product limit method. The related data were taken from Ege University Faculty of Medicine, Branch of Radiation Oncology and contains information about survival data and type of tumor belonging to NSCLC patients. The differences in two groups were researched by logrank and Wilcoxon tests. In the last section, the results of the application are discussed. Using Kaplan- Meier product limit method discovers the differences between the levels of variables measured in classification level and the survival times of the patients according to NSCLC are investigated.
Markov karar süreçleri ve uygulama
Markov decision process and dynamic programming which is a solution technique are discussed in this thesis. In application, monthly demands are considered as a random variable, and the process is expressed as a Markov decision process. As a result, according to the data, the optimal production policy is determined. In conclusion, with this work and the model we constructed, we indicate that the solution, which is intuitively optimal, is also technically optimal.
Risk teorisi ve bir uygulama
ABSTRACT Insurance risk models, number of claims, claim sizes, claim frequency rate, and calculation of the risk premium are discussed in this thesis. An application depending on data provided by an insurance company involving with motor vehicle insurance supported our study. In conclusion, we see that the realized values of the five years totals number of claims and number of policies are very close to the values obtained under the planning assumptions. >KÜWiÂNTASYON MERKE2İ
Ortak değişkenlerin varlığı durumunda faktöriyel tasarımlar
Bu çalışmada ortak değişkene sahip faktöriyel tasarımlar ele alınmıştır. Bu tasarımlar için klasik teori, normallik varsayımına dayalı elde edilir. Buna karşın, hata terimleri normal dağılıma sahip değilse model parametrelerinin en küçük kareler (EKK) tahmin edicilerinin etkinliğinin ve test istatistiklerinin gücü ve istatistiksel sağlamlığının azaldığı Monte-Carlo simulasyon çalışması yardımıyla gösterilmiştir. Bu sonuçlar ortak değişkene sahip faktöriyel tasarımlarda normallik varsayımı geçerli olmadığı zaman, EKK ya alternatif olarak, daha etkin tahmin ediciler ile daha güçlü ve istatistiksel olarak sağlam test istatistiklerinin elde edilmesini gerektirir.Bu nedenlerle ortak değişkene sahip faktöriyel tasarımlarda hata terimlerinin bağımsız ve özdeş olarak uzun kuyruklu simetrik (LTS) dağılıma sahip olduğu var\-sayıl\-mıştır. Bu durumda en çok olabilirlik tahmin edicileri analitik olarak elde edilemediğinden uyarlanmış en çok olabilirlik yöntemi kullanılmış ve bu yönteme dayalı olarak parametrelerin tahmin edicileri açık formüllerle ifade edilmiştir. Uyar\-lan\-mış en çok olabilirlik (UEÇO) tahmin edicilerinin EKK tahmin edicilerinden daha etkin olduğu Monte-Carlo simulasyon çalışmasıyla gösterilmiştir. UEÇO tahmin edicilerine dayalı geliştirilen test istatistiklerinin asimptotik olarak F dağılımına sahip olduğunu kanıt\-lanmış ve küçük örneklem hacimleri için de bu test istatistiklerinin F dağılımına sahip olduğu Monte-Carlo simulasyon çalışması yardımıyla gösterilmiştir. Buna ek olarak UEÇO tahmin edicilerine dayalı test istatistiklerinin klasik test istatistiklerinden daha güçlü ve istatistiksel olarak sağlam olduğu Monte-Carlo simulasyon çalışmasıyla gösterilmiştir. Geliştirilen yöntem bir gerçek hayat örneği üzerinde uygulanmıştır.
Bulanık mantık çıkarım sistemi ile tren hızının otomatik kontrolü
Bu tezde, bulanık mantık ve adaptif ağ yapısına dayalı bulanık çıkarım sistemi (Anfis) ile Otomatik Tren Yönetim Sistemi uygulaması yapılmıştır. Öncelikle Otomatik Tren Yönetim Sistemi üzerine yapılan teorik araştırmalardan yola çıkılarak, trenlerin bulanık mantık ile kontrol ve yönetimi sağlanmıştır. Daha sonra mevcut teknik bir adım daha ileriye götürülerek, yapay sinir ağları ve bulanık mantığın bir arada kullanıldığı Anfis ile (dünyada ilk kez denenerek) modellenmiş ve simülasyonu gerçekleştirilmiştir. Bulanık mantık ve Anfis kontrolünün birlikte uygulanması ile Otomatik Tren Yönetim Sistemi gerçek hayatta her bölge ve koşulda çalışabilecek, güvenliği en üst düzeyde olan bir sistem haline getirilmeye çalışılmış ve elde edilen sonuçlar sunulmuştur.
Stokastik diferensiyel denklemlerle modelleme
Bu tez çalışmasında stokastik diferansiyel denklemlerin iki alt denklem sınıfı olan rassal diferansiyel denklemler ve Itô stokastik diferansiyel denklemleri ile somut problemlerin modellenmesi konusu ele alınmıştır. Bu amaçla öncelikle stokastik diferansiyel denklemler teorisi için gerekli olan matematiksel temeller verilmiş, iki somut problem için stokastik diferansiyel denklem modelleri kurulmuş veçözümlenmiştir.İlk olarak, bir radyoaktif bozunma problemi bir rassal diferansiyel denklem ile modellenmiş ve bu modelin çözümü elde edilmiştir. Daha sonra çözüme ilişkin bazı çıkarsamalar yapılmış ve elde edilen sonuçlar, tablo ve şekiller ile sunulmuştur.İkinci olarak ise hisse senetleri fiyatları için Samuelson modeli tanıtıldıktan sonra bu çalışmada verilen modelleme yöntemiyle fiyatlar için bir stokastik diferansiyel denklem modeli kurulmuştur. İki modeli fiyat tahminleri bazında karşılaştırabilmek için MOTOROLA hisse senedinin 20.03.09-11.05.09 tarihleri arası günlük kapanış fiyatları veri seti ele alınmıştır. MATLAB dilinde yazılan programlar yardımıyla her iki modelin parametreleri tahmin edilmiştir. Daha sonra, elde edilen modeller nümerik olarak çözdürülmüş ve hesaplanan tahminler tablolar ve şekiller yardımıyla ortaya konulmuştur. Son olarak her iki model Euclid metriğine göre karşılaştırılmış ve Samuelson modelinin çalışmada verilen yöntemle elde edilen modelden daha iyi sonuçlar verdiği görülmüştür.
Mekansal istatistikte nokta örüntü teknikleri ve bir uygulama
Mekansal istatistik, mekansal olarak düzenlenmiş verileri dikkate alan istatistiksel yöntemlerle ilgilidir. Mekansal verinin içinde mevcut olan mekansal bağımlılık varsayımı yüzünden mekansal istatistik, klasik istatistik yöntem ve tekniklerini kullanma eğiliminde değildir. Klasik istatistikte olduğu gibi, mekansal veri için tanımlayıcı ve çıkarsamalı yaklaşımlara sahip olmakla birlikte kendine özgü yöntem ve tekniklere sahiptir. Mekansal istatistikte kullanılan yöntemler genellikle analiz edilmekte olan mekansal verinin türlerine göre üç kategoriye ayrılmaktadır. Mekansal verinin bu türlerinden biri de mekansal nokta örüntü verileridir. Mekansal nokta örüntü verileri, nokta olayların konumlarından elde edilmiş verilerdir. Birbirleriyle ilişkili konumların, anlamlı bir örüntüyü temsil edip etmediği ile ilgilenilmektedir. Mekansal nokta örüntüler analiz edilirken, temel olarak tam mekansal rassalığa karşılık örüntülerin kümelenme ve düzenlilik gösterip göstermediği ile ilgilenilmektedir. Tam mekansal rassallıktan herhangi bir sapmanın değerlendirilmesine olanak sağlayan dağılım fonksiyonlarının tahminlerinin, tam mekansal rassallık altında bir dağılım ile karşılaştırılmasında bazı simülasyon teknikleri kullanılmaktadır.Bu çalışmada mekansal istatistik yöntem ve tekniklerinin deprem verisi üzerinde uygulanması ele alınmaktadır ve bu amaçla ülkemizin yakın geçmişte büyük bir deprem ile sarsılmış olan Gölcük bölgesi seçilmiştir. Bu bölgede meydana gelmiş depremler yalnızca istatistiksel veri olarak ele alınarak, mekansal istatistik yöntem ve teknikler aracılığıyla bu bölgede olabilecek depremler için benzetim çalışması 1900 yılından 01 Ocak 2010 tarihine kadar elde edilen deprem verileri yardımıyla türetilmiştir.
Ridge ve Liu tahmincilerinin etkinliklerinin ve yanlılıklarının karşılaştırılması
Çoklu regresyon analizinde karşılaşılan sorunlardan birisi de çoklu bağıntı durumudur. Bağımsız değişkenlerden bir veya birkaçının diğer bağımsız değişkenler tarafından iyi açıklandığı zaman, sonuçlarda istenmeyen özellikler oluşturan çoklu bağıntı sorunu meydana gelmektedir. Çoklu bağıntıyı gidermek veya azaltmak için yanlı tahmin yöntemleri kullanılır. Bu çalışmada, yanlı tahmin yöntemleri olarak bilinen Ridge ve Liu tahmincilerinin karşılaştırılmasına yer verilmiştir. İlk olarak bu iki tahminci tanımlanmış ve bununla ilgili 1985?2006 yılları arasında Türkiye'deki turizm geliri fonksiyonunu açıklayan değişkenler olarak; yatak kapasitesi, turist sayısı, seyahat acentelerin sayısı, yabancı sermaye miktarı, Euro cinsi döviz kuru, ABD doları cinsi döviz kuru üzerine bir uygulama yapılmıştır. Sonuç olarak, bu iki yöntem etkinlikleri ve yanlılıkları bakımından karşılaştırılmış ve sonuçlar yorumlanmıştır.Anahtar Kelimeler: Çoklu Bağıntı, Ridge Tahmincisi, Liu Tahmincisi, Yanlı Tahmin Yöntemleri, Turizm Geliri
ARCH modelleriyle bazı ülkelerin döviz kurlarının volatilitesinin incelenmesi
Finansal serilerde, taşıdıkları özellikler nedeniyle doğrusal zaman serisi yerine, doğrusal olmayan koşullu değişen varyans modellerinin kullanılması giderek daha yaygın hale gelmiştir. Öngörü hataları varyansının sabit olmadığı, değişen varyansa sahip olduğu zaman serisinin çözümlenmesinde serilerin bu özelliğini de dikkate alacak modellere gereksinim duyulmuştur. Robert F. Engle (1982), geçerliliği olamayan yukarıda belirtilen varsayımı genelleştirmiş ve Otoregresif Koşullu Değişen Varyans (ARCH) süreçleri olarak adlandırılan stokastik süreçlerin yeni bir sınıfını önermiştir. Bu çalışmada, bazı ARCH (GARCH, GARCH-M ve EGARCH, TGARCH) modellerinin istatistiksel özellikleri ve tahmin yöntemleri incelenmiş, bu modeller farklı gelişmişlik düzeylerindeki rasgele seçilen on ülkenin döviz kuru serilerine uygulanmıştır. Model sonuçları karşılaştırılarak serilere en uygun koşullu varyans modeli belirlenmiştir.
İlköğretim 7. ve 8. sınıf öğrencilerine yönelik negatif tamsayılara ilişkin tutum ölçeği geliştirilmesi ve lojistik regresyonla analizi
Matematik dersine karşı oluşan tutum ile matematik başarısı arasındaki ilişki literatürde üzerinde sıkça durulan önemli araştırma konularından birisidir. Öğrencinin bir konuya olan başarısının artmasında olumlu tutum geliştirmesi oldukça önemli olmaktadır.Negatif tamsayılar konusu ise, özellikle ilköğretim ikinci kademede öğrenilmeye başlanan ve öğrencilerin gerçek hayata uyarlamakta zorlandıkları ve öğrenmekte zorluk yaşadıkları başlıca konuların arasına girmektedir. Bu tez çalışmasında Negatif Tamsayılara karşı tutum ölçeği oluşturulmaya çalışılmıştır. Taslak ölçek 7. ve 8. sınıf 220 öğrenciye uygulanmış, faktör analizi uygulanarak yapı geçerliliği ortaya konulmuş, genel güvenilirlik için Cronbach's Alpha güvenirlilik katsayısı hesaplanmıştır. Sonuçlar % 95 güven düzeyinde değerlendirilmiştir. Sonuç olarak 28 madden oluşan tek bir faktör altında toplanan bir tutum ölçeği geliştirilmiştir.Ölçek geliştirildikten sonra ilk önce alınan veriler ile matematik başarı puanları arasındaki ilişki gözlemlenmiş daha sonra ise öğrencilerin tutum ölçeklerinden aldıkları toplam puanlar eşit aralıkta 4'e bölünerek sıralı hale getirilmiş ve öğrencilerin demografik verileri arasındaki ilişki gözlemlenmeye çalışılmıştır. Bağımlı değişkenler sıralı halde olduğu için sıralı lojistik regresyon tekniği tercih edilmiştir.Oluşturulan sıralı lojistik regresyon modellerinden elde edilen sonuçlarda, bir dönem önceki matematik notları ile bir dönem sonraki matematik notları arasında anlamlı bir ilişki gözlenmiştir. Sadece 7. Sınıf verileri ile oluşturulan diğer bir modelde matematik başarı notları ile tutum ölçeği maddelerinin 24 maddesinin değişik kategorilerinin anlamlı bir ilişkide oldukları sonuçlarına varılmıştır. Ayrıca oluşturulan son modelde negatif tam sayılara karşı 7. Sınıfların 8. Sınıflara göre daha olumlu tutum gösterme eğiliminde oldukları ve okul dışında eğitim yardımı alan öğrencilerinde diğer öğrencilere göre daha olumlu tutum gösterdikleri sonucuna ulaşılmıştır.Anahtar kelimeler: Tutum ölçeği, negatif tamsayılar, sıralı lojistik regresyon
Doğrusal regresyon modeli için m-tahmincilerin incelenmesi
Robust (Sağlam) regresyon tahmincileri, hataların normal dağılıma uymadığı veya veri setinde aykırı değer bulunması durumunda regresyon modelini en güvenilir şekilde tahmin etmek amacı ile geliştirilmiştir.Bu tez çalışmasının amacı, veri setinde aykırı değer olması durumunda en küçük kareler tahmincisine alternatif olarak geliştirilen robust regresyon tahmincilerinden M-tahmincilerin çeşitli açılardan incelenmesidir.İlk olarak M-Tahmincilerin hesaplanmasında kullanılan yeniden ağırlıklandırılmış en küçük kareler algoritmasının başlangıç tahminlerinin seçimine olan duyarlılığı ele alınmış ve M-tahmincilerinin kırılma noktaları, grafikler yardımıyla incelenmiştir. Daha sonra hata teriminin dağılımının normal ve normalden farklı olduğu durumlar için M-tahmincilerinin etkinlik açısından performansı değerlendirilmiş ve son olarak başlangıç ölçek tahmincisinin etkinliğe katkısı araştırılmıştır. Ayrıca reel yaşamdan alınan iki örnek üzerinde M-tahminciler uygulanmış ve elde edilen sonuçlar tartışılmıştır.
Parametrik olmayan bulanık regresyon modelleri analizi
Tez çalışmasında, parametrik olmayan bulanık regresyon modelleri incelenmiştir. k-en yakın komşuluk, çekirdek düzeltme ve yerel polinomiyal düzeltme modelleri bulanık yapıda ifade edilmiştir. Ayrıca bu modeller için, band genişliği seçiminde çapraz geçerlilik ve genelleştirilmiş çapraz geçerlilik kriterleri, şapka matrisi kullanılarak geliştirilmiştir. Çalışmalarda kullanılan verilerde, yanıt değişkenin bulanık değerli ve açıklayıcı değişkenin kesin değerli olduğu durum dikkate alınmıştır. Analizler R programında locpol paketi kullanılarak yapılmıştır. Uygulamalarda özellikle, bulanık yerel polinomiyal regresyon modelinde polinomun derecesinin doğrusal ve kübik alındığı durumlar üzerinde inceleme yapılmıştır. Bu modeller için band genişliği, geliştirilen çapraz geçerlilik ve genelleştirilmiş çapraz geçerlilik kriterleri ile seçilerek, modellerin performansları ortalama karesel hata değerleri kullanılarak karşılaştırılmıştır. Elde edilen sonuçlardan genelleştirme yapabilmek amacıyla farklı türdeki veri setleri üzerinde çalışılmıştır. Uygulamaların pek çoğunda bulanık yerel kübik modele ait performans değerleri daha yüksek olmasına rağmen, özellikle eğriselliği fazla olan modelleri daha pürüzsüz (dalgalanması az olan) bir şekilde ifade eder. Ayrıca, band genişliği değerinin bulanık yerel doğrusal modele göre daha yüksek seçilmesi ile işlem basamaklarını azalttığı göz önüne alındığında, bulanık yerel kübik modellerin kullanımının daha faydalı olduğu sonucuna varılmıştır.
Türkiye'deki istatistik bölümlerinin göreli etkinliklerinin veri zarflama analizi ile belirlenmesi
Bir eğitim sisteminin başarısı, eğitim süreci içinde sunduğu olanaklar ve eğitim sonrası elde edilen yeterlilikler tarafından belirlenebilmektedir. Bu bağlamda ülkemizdeki üniversite performanslarının karşılaştırılmaları genellikle akademisyen performansları, mezunlarının işsizlik oranı ve ülke genelinde yapılan ortak sınav sonuçları üzerinden çıktı yönlü olarak yapıldığı gözlemlenmektedir. Kamu Personeli Seçme Sınavı (KPSS) sonuçları da karşılaştırma unsuru olarak kabul gören bir çıktıdır. Fakat eğitim bilimciler eğitim sitemlerinin performanslarının değerlendirilmesinde sunulan olanakların çıktıya ne kadar verimli olarak dönüştürebildiklerine önem vermektedirler. Bu çalışmada Türkiye?deki devlet üniversitelerinin verilerine ulaşılan 18 adet istatistik bölümü homojen karar verme birimi (KVB) olarak ele alınmış ve bölümlerin göreli etkinlikleri veri zarflama analizi (VZA) yardımı ile hesaplanmıştır. Etkinlik değerleri arasındaki farklılıkların kontrol edilemeyen girdiler tarafından etkilenip etkilenmediği, etkinlik değerleri ve girdi-çıktı değişkenleri arasındaki ilişkiler ve öğretim programlarına göre bölümler arasında etkinlik farkların anlamlılığı ortaya konulmaya çalışılmıştır. Elde edilen sonuçların karar vericilere ve istatistik bölümünde okuyan öğrencilere yol gösterici olması amaçlanmıştır.
Dağılım ve sinir ağı tabanlı bulanık zaman serisi modelleri
Bulanık zaman serisi yaklaşımları genelde bulanıklaştırma, bulanık ilişkiler belirleme ve durulaştırma olmak üzere üç aşamadan oluşur. Bu çalışmada, tek değişkenli birinci dereceden sinir ağı tabanlı bulanık zaman serisi öngörüsü için yeni bir yaklaşım ve yeni bir yöntem önerilmiştir. İlk olarak evrensel küme parçalanmasında yapılan aralık uzunluğu belirleme aşamasında sabit bir aralık uzunluğu almak yerine daha etkili olan dağılım tabanlı uzunluk yaklaşımı kullanılmıştır. Bulanıklaştırma aşamasında yeni bir algoritma oluşturularak işlem kolaylığı sağlanmıştır. Ayrıca bu aşamada önerilen yöntemde ilk defa ağırlıklandırılmış indisler kullanılmıştır. Bulanık ilişki belirlemede bütün üyelik derecelerinin ayarlanması sağlanmıştır. Öngörü performansını geliştirmek için sadece Çok Katmanlı Algılayıcı (ÇKA) değil, ayrıca Genelleştirilmiş Regresyon Sinir Ağları (GRSA), ve Radyal Tabanlı Fonksiyon Sinir Ağı (RTFSA) gibi farklı yapay sinir ağı mimarileri de uygulanmıştır. Bu YSA mimarileri için klasik sinir ağı tabanlı bulanık zaman serisi öngörü yöntemlerinden farklı olarak tek katman ve düğüm sayısı, girdi ve çıktıların toplamları olması yerine en iyi sonucu verecek şekilde farklı sayıda gizli katman ve düğüm sayısı kullanılmıştır. Önerilen yöntem ve yaklaşım, oldukça iyi bilinen ve literatürde sıklıkla kullanılan Alabama Üniversitesi kayıt verileri ve ayrıca büyük bir veri seti olarak İ.M.K.B. (BİST) ulusal 100 endeksi verileri 2006-2010 yılları için kullanılarak literatürde önerilmiş sinir ağı tabanlı veya sinir ağı tabanlı olmayan çeşitli bulanık zaman serisi öngörü yöntemleri ile karşılaştırılmıştır. Sonuçlar; önerilen yeni yöntemin, literatürde verilen diğer yöntemlerden üstün olduğunu göstermiştir.
Sağkalım analizinde COX regresyon ve yapay sinir ağları kullanımı
Sağkalım analizinde en yaygın kullanılan regresyon modeli Cox Regresyon Modelidir. Bu çalışmada, sağkalım analizinde cox regresyon modeli ve yapay sinir ağları modelleri kullanılarak Lösemi hastalarına ait veriler analiz edilmiştir. Çalışmada Yemen Cumhuriyeti'ndeki 2017 yılı Ocak ayı ve 2022 yılı Şubat ayı arasında 1168 Lösemi hastasına ait bilgiler kullanılmıştır. Bu tezin amacı, en yaygın hastalıklardan biri olan kanser hastalığı ile başa çıkmada Yemen Cumhuriyeti'ndeki lösemi hastalarının sağkalım sürelerini etkileyen en önemli faktörleri (risk faktörleri) belirlemektedir. Ayrıca, sağkalım analizinde Yapay Sinir Ağlarını kullanmak ve Yapay Sinir Ağları modeli ile Cox Regresyon Modeli arasındaki performans göstergelerini karşılaştırmaktır. Söz konusu amaçla Cox Regresyon Modeli ve Yapay Sinir Ağları modelleri sağkalım verilerinin analizinde kullanılmıştır. Önerilen Cox modellerinden en iyi model seçilmiş ve ardından aynı veriler yapay sinir ağları kullanılarak analiz edilip modellerden en iyi olanı da seçilmiştir. Cox regresyon modeli sonuçları ile yapay sinir ağlarının kullanımından elde edilen sonuçlar arasında karşılaştırma yapılmıştır. İki yöntem arasındaki karşılaştırma, Hata Kareler Ortalaması (MSE) ve Ortalama mutlak hata (MAE) kriterlerine göre gerçekleştirilmiştir. Cox regresyon modellerinden en iyi model ve yapay sinir ağı modellerinden en iyi model seçilip karşılaştırdıktan sonra yapay sinir ağları modelinin, daha önce belirtilen kriterlere göre sağkalım verilerini analiz etmekte Cox regresyon modelinden daha iyi performans gösterdiği soncuna varılmıştır.
Makine öğrenimi teknikleri ile güneş ışınımı ve güç kestirimi
Son yıllarda iklim değişikliği ve özellikle küresel ısınmanın etkileri görüldükçe, enerji tasarrufu ve yenilenebilir enerjinin önemi giderek artmıştır. Yenilenebilir enerji alanında başta güneş enerjisi ve rüzgar enerjisi olmak üzere jeotermal, biyokütle, dalga ve hidroelektrik enerji gibi pek çok enerji sistemi uygulaması mevcuttur. Yenilenebilir enerji sistemleri kurulmadan önce ekonomik ve çevresel sürdürülebilirliklerinin incelenmesi sistemlerden optimum verim elde etmek amacıyla önem taşımaktadır. Bu doğrultuda yenilenebilir enerji sistemlerinde güvenilirlik ve enerji tahmini çalışmaları sürdürülebilirlik bağlamında etkili çalışma alanları arasına girmektedir. Enerji sistemlerin geleceğe yönelik güvenilirlik ve enerji tahmininde makine öğrenimi metotları kullanılmaktadır. Yapılan çalışmada da kurulması planlanan bir güneş enerjisi sisteminin enerji tahmininde güneş ışınımı tahmini için Yapay Sinir Ağları, LightGBM (Hafif Gradyan Artırma Makineleri) ve LSTM (Uzun Kısa Zamanlı Bellek) gibi makine öğrenimi modelleri kullanılmıştır. Kullanılan makine öğrenimi modellerinin tahmin performansları çeşitli hata metrikleri ile karşılaştırılarak en başarılı sonuçları veren LightGBM modeli seçilmiştir. Seçilen modelin test verisi üzerinde Ortalama Mutlak Hata Yüzdesi (MAPE) değeri %21,29 olmuştur. Seçilen LightGBM modeli üzerinden yapılan güneş ışınımı tahminleri kullanılarak belirli parametrelerdeki bir güneş paneli sisteminin sonraki güne ait saatlik güç ve günlük enerji üretim kestirimi yapılmıştır.
Çamaşır makinelerinde farklı leke grupları için yıkama performansına etki eden parametrelerin denetimli öğrenme algoritmaları ile tahminlenmesi
Ev tipi çamaşır makinelerinin yıkama performansının, Avrupa'daki eko tasarım gereklilikleri gibi belirli sınır değerlere uyması gerekmektedir. Çamaşır makinelerinde lekelerin yıkama performansları için belirlenen sınır değerleri EN 60456 standardına göre elde edilmelidir. Yıkama performansını etkileyen yıkama süresi, su seviyesi, deterjan miktarı, motor yoğunluğu, hacim başına yıkanan yük miktarı ve sıcaklık olarak belirlenen faktörlerin çeşitli düzeyleri için yıkamalar gerçekleştirilmektedir. Bu tez çalışmasında Vestel Beyaz Eşya firmasından sebum, karbon, kan, kakao ve kırmızı şarap lekesi olan kumaşlar üzerinde çeyrek yükle, yarım yükle ve tam yükle elde edilen yıkama performans değerleri temin edilmiştir. 5 leke için yıkama performans değerleri ile tüm lekelerin yıkama performans değeri toplamı ve 3 ayrı yük için toplam 18 veri kümesi elde edilmiştir. Elde edilen veri kümeleri çoklu doğrusal regresyon, regresyon ağacı, torbalama (bagging), rassal ormanlar regresyonu, XGBoost regresyonu, destek vektör regresyonu (doğrusal çekirdek ve polinom çekirdek) ve k-en yakın komşu regresyonu algoritmaları ile modellenmiştir. Tüm modellerin performanslarının karşılaştırılması için hata kareler ortalaması, kök hata kareler ortalaması ve ortalama mutlak hata metrikleri kullanılmıştır. 18 veri kümesinin 8'inde rassal ormanlar regresyonu, 7'sinde bagging ve 3'ünde XGBoost modeli en iyi model olarak tespit edilmiştir. Sonuç olarak ağaç tabanlı algoritmalar kullanan modeller olarak ön plana çıkmıştır.
Sıralı küme örneklemesine dayalı makine öğrenmesi teknikleri
Son yıllardaki hızlı veri artışı ve bu verileri analiz etmenin giderek zorlaşmasıyla birlikte, günümüzde birçok çalışma alanında makine öğrenmesi kullanılmaktadır. Makine öğrenmesi, insan beyninin öğrenme mantığını esas alarak, çeşitli algoritma ve teknikler geliştirmeyi amaçlayan bilimsel çalışma alanıdır. Makine öğrenmesi algoritmaları, olayları inceler ve nasıl meydana geldiklerini anlamaya çalışır. Bu çabaları sonucunda, elde ettikleri sonuçlar ile genelleme yapma yeteneği kazanırlar. Makine öğrenmesi algoritmalarının bilgileri öğrenmeleri ve ne kadar iyi öğrendiklerinin değerlendirilmesi için; veri kümesi, eğitim ve test seti olarak ikiye ayrılır. Literatürde bu işlem, kullanıcının belirlediği bir oranda rastgele olarak yapılır. Bu tez çalışmasında; veri seti bölme işlemi, literatürde kullanılan yönteme ek olarak, Sıralı Küme Örneklemesi (SKÖ), Uç Değer SKÖ (USKÖ), Medyan SKÖ (MSKÖ) ve Yüzdelik SKÖ (YSKÖ) yöntemleriyle yapılarak, elde edilen sonuçların karşılaştırılması amaçlanmıştır. Farklı çalışma alanlarından seçilen gerçek hayat verileri, belirtilen yöntemler ile eğitim ve test setlerine ayrılmıştır. Eğitim setleri ile makine öğrenmesi algoritmaları eğitilerek, test setleri ile öğrenme başarıları sınanmıştır. Karşılaştırma kriteri olarak; regresyon algoritmalarında hata kareler ortalamasının karekökü (HKOK) değerleri, sınıflandırma algoritmalarında ise doğru sınıflandırma oranları kullanılmıştır.
MAUT ve TAOV yöntemlerinin birleştirilmesiyle Türkiye için en faydalı yenilenebilir enerji alternatifinin belirlenmesi
The need for energy in the world has recently increased as a result of the increasing population, global growth, and industrialization. The failure to meet the increasing energy with the existing fossil-resourced reserves, and the increase in environmental awareness and energy supply security shows that the use of renewable energy resources in Turkey is quite important as well as in the world. In the present study, a new combined MAUT and TAOV multi-criteria decision-making approach (TAOV- MAUT) method is submitted by combining the Multi Attribute Utility Theory (MAUT) and Total Area Based on Orthogonal Vectors (TAOV). It is tried to determine the most useful renewable energy resource for Turkey by using proposed method. According to the literature review on the evaluation of renewable energy sources in Turkey, the most important criteria that can affect the decision were determined as follows: efficiency, construction time, cost, government incentives, economic life, foreign dependency, employment, opportunities, social acceptance, space requirement and greenhouse gas emissions. As a result of the study, the most useful renewable energy sources in Turkey were determined as hydroelectric, wind, solar, biomass and geothermal energy sources, respectively.
Akıllı ulaşım sistemleri için ileri istatistiksel yöntemler
The use of machine learning techniques and statistical methods for intelligent transportation systems has gained importance recently. In this thesis, two different methods are proposed for clustering and modeling that can serve the needs of the transportation field. Clustering methods are used to group data points quickly and easily for further analysis of clusters. Also, most clustering methods require complex computations or have an iterative procedure that makes the algorithm time-consuming, especially when the data are relatively large. In this thesis, we propose a new clustering method called spatial adaptive clustering (SAC) based on the idea of adaptive cluster sampling (ACS) design. ACS is a sampling method, which is based on neighborhood search on a grid structure, has an adaptive selection process of units and recursively added units reveal the batched individuals easily and quickly. The SAC algorithm forms clusters based on neighborhood search using grid structures and can detect noise points. The performance of the proposed algorithm is evaluated through comparisons with the results from well-known density-based clustering approaches in the literature using real and artificial data sets. Also, hot spots of accident locations in Birmingham/England and Buca/Turkey are investigated using the SAC algorithm. Additionally, to reduce the number of accidents, hotspot locations are examined in terms of the factors causing the accident. Vehicle headway modeling has also been one of the other important topics for traffic signal optimization and flow modeling. In this thesis, we estimate the parameters of Exponentiated Weibull (EW) distribution using the maximum likelihood method under the assumption that all parameters are unknown. We deal with the performance of ranked set sampling and simple random sampling methods by a simulation study in R-software in terms of mean-squared error. Finally, we illustrate the flexibility and usefulness of EW distribution by analyzing simulated data from a real application study in the transportation field.
Güneş fotovoltaik sistemlerinin performansı için güvenilirlik modellemesi
Solar energy is one of the most widely used renewable energy sources. Photovoltaic systems directly convert solar energy into electricity with no carbon dioxide emission or any other air pollutants. The power generated by a photovoltaic system depends on the characteristics of solar irradiation and weather conditions. This stochastic nature of power systems prompts the researchers to use probabilistic and statistical techniques. In this thesis, we consider the distribution function of the power generated by a photovoltaic system to make predictions about its characteristics. We model the mean power generated by a photovoltaic system. We also consider that photovoltaic modules may have multistate working conditions and different performance levels depending on solar radiation. In this concept, we present a model for solar power systems with PV modules having various levels of operational performance and we develop a reliability model for the system's power regarding the threshold value that is the minimum required total performance level for the system. This model reflects the performance levels and the working probabilities of PV modules. The problem is evaluated under different conditions regarding the dependency of multistate PV modules. In addition, we provide the optimum number of photovoltaic modules that minimizes the total cost based on the level of required total power production. For further analyses, we give real data applications to estimate the characteristics of the power produced by the solar plant for a specific location in Izmir, Turkey. As the software programming tools, R v.1.2.5033 and Mathematica v.11.3 are used for the computations.
ATA öngörü yönteminin ampirik özelliklerinin incelenmesi
Forecasting is important in all scientific fields such as industrial, commercial, medical and economic. There are many forecasting methods in the literature, but exponential smoothing is a very popular method due to its simplicity and accuracy. Simple exponential smoothing is used for data sets randomly distributed around a constant level. Holt's linear trend method is a method that helps to deal with linearly trended data. Despite the fact that exponential smoothing methods are widely used and have been in the literature for a long time, they have some problems that potentially affect the predictive accuracy of models. Ata is a new forecasting method that has been proposed to overcome these problems and to provide better forecasts. In this thesis, the forecasting accuracy of Ata and exponential smoothing will be compared for data sets with no or linear trend. The results given in this study are obtained using simulated data sets with different sample sizes and variances and the forecast accuracy is compared using the mean squared forecast error. In line with these results, the forecast accuracy is calculated for both short and long term forecasting horizons. The results reveal that the proposed approach outperforms exponential smoothing for most types of time series data for both short and long term forecasting horizons.
Farklı bağımlılık varsayımlarına dayalı saklı Markov modelleri
Hidden Markov models are widely used to model the probabilistic structures with latent random variables. The main assumption of hidden Markov models is that; observations are conditionally independent and identically distributed random variables. There may exist some cases where this assumption may not be valid in practice. That is, an observation that occurs in the current state may depend on the previous observation symbol that occurred in the previous state. In this thesis, two types of hidden Markov models are introduced which differ from the classical hidden Markov model based on different first-order Markov dependence assumptions. The introduced models are capable of capturing a possible first-order Markov dependence between the successive observations or successive system informations. They can provide better representation for the appropriate real-life problems where, if the observations have some conditional dependencies among them. The two proposed models are defined with their assumptions and using appropriate notation. Modifications to the algorithms used for parameter estimates and hidden state sequence estimates are explained. In addition, an experimental study is conducted to show the performance of the introduced models compared to the classical hidden Markov model. According to the results of the experimental study, the proposed models outperform in generated observation sequences that have appropriate assumptions. Besides, two different case studies are conducted namely the occurrences of strong earthquakes and daily stock prices. They are modelled with both the classical hidden Markov model and the proposed models, and the results are compared.
Sağkalım verileri için makine öğrenmesi yöntemleri
Survival analysis is a the statistical approach methods used in many fields. The key feature that distinguishes this method from other analysis methods is that it can be used when censored observations are present. In the literature, survival analysis can be categorized as traditional survival analysis methods and machine learning based survival analysis methods. The increasing number of data and variables, the existence of censored observations and the fact that traditional methods require some assumptions make it difficult to analyze survival data with traditional methods. In order to cope with this situation, machine learning methods specific to survival data are used. In this thesis, the traditional survival analysis methods such as the Kaplan Meier, Log rank test and the Cox regression are introduced and machine learning based Survival trees, and Random Survival Forests are studied. In application concordance statistics of Random Survival Forests and Cox regression models were obtained by using both real and simulated data in which different censoring rates were tested. Propher graphical representations were given for the survival curves and variable importance metrics when necessary. As demonstrated by the examples provided in this thesis the traditional approaches are adventageous when assumptions are met and statistical power is high however machine learning based methods work better when censoring is high an assumptions are not satisfied.
Lojistik regresyon ve karar ağacı algoritmalarının tahmin edici performanslarının karşılaştırılması: Yaşam memnuniyeti uygulaması
Decision tree algorithms and regression in machine learning create classes of data. Relationships between variables are modeled. Decision trees create classification rules using training data. They also test these rules on test data. Thus, the decision tree determines the success of the algorithm. With the model created in logistic regression, classification is created and classification performance is found. These methods are easy to interpret. They are easily applied to large data sets. They are used in many different fields due to the lack of assumptions. Satisfaction, which is a part of the concept of life satisfaction, is the fulfillment of needs, desires and wishes. Life satisfaction deals with a person's entire life. Life satisfaction is the whole of processes related to individuals' own life patterns and standards. The aim of this study is to compare the performances of logistic regression method and decision tree algorithms (CART, CHAID, QUEST) estimators using life satisfaction data (n = 8430) obtained by the Turkish Statistical Institute (TURKSTAT) for the year 2017. In this study, performance comparisons (accuracy, sensitivity, selectivity, precision, F-score) were made and it was found that the model that best explains the concept of life satisfaction is the QUEST algorithm.
Makine öğrenmesi algoritmaları ile meteorolojik parametreleri kullanarak toprak radon gazının tahmini
Radon is the natural radiation source with the highest dose of exposure among all radiation sources found on earth. It consists of the degradation of natural uranium and radium elements in rocks. Factors that shape the movement of radon include meteorological factors such as the rate of decaying of radon isotopes, the fluids that fill the pores(air, water and other gases), atmospheric pressure, soil and air temperature, wind speed, and wind direction. The aim of this study is to evaluate the effects of some meteorological factors on Radon gas using Supervised Learning Algorithms and to estimate the radon gas values according to these factors. For the study, in addition to the radon levels obtained from the Seferihisar region in hourly periods between 30 October 2006 and 04 June 2007, the measurements for the parameters of Hourly Actual Pressure(hPa), Hourly 50 cm Soil Temperature(°C), Hourly Relative Humidity(%), Hourly Temperature(°C), Hourly Wind Degree(°), Hourly Wind Speed(m/sec) and Wind Direction were obtained from the Republic of Turkey Ministry of Agriculture and Forestry, General Directorate of Meteorology. To analyze the relationship between Radon and meteorological factors affecting radon with Supervised Learning Algorithms, Multiple Linear Regression, k-Nearest Neighbor, Support Vector Machines, Regression Trees, Bagging, Random Forests, XGBoost methods have been used. To test the success of applied methods K-Fold Cross-Validation(K=5) and verification tests were performed. Specification coefficient(R^2) for comparing the performance of algorithms, Mean Squared Error(MSE), Root Mean Squared Error(RMSE), Mean Absolute Error(MAE) values were used. The best result was random forests regression when performance criteria were taken into account. This method was followed by the XGBoost and k-Nearest Neighbors algorithms, which gave very close results.
Basit doğrusal regresyon modelinde model parametrelerinin dayanıklı tahmini
DANIŞMAN: DOÇ. DR. DEMET HAN AYDIN Bu tez çalışmasında, basit doğrusal regresyon modelinde model parametrelerinin En Küçük Kareler (Least Squares–LS) tahmin edicilerinin performansı farklı senaryolar altında incelenmiştir. İlk senaryoda, hata terimlerinin dağılımı olarak normal dağılıma alternatif olabilecek farklı olasılık dağılımları ve çeşitli aykırı değer modelleri dikkate alınmıştır. İkinci senaryoda ise, hata terimlerinin sağa çarpık bir dağılım olan Gumbel dağılımını izlediği varsayılmıştır. Ayrıca, X-yönünde aykırı gözlemlerin bulunduğu durumlarda, LS tahmin edicilerinin performansları, literatürde yaygın olarak kullanılan En Küçük Mutlak Sapma (Least Absolute Deviation-LAD), Ağırlıklı En Küçük Mutlak Sapma (Weighted Least Absolute Deviation-WLAD) ve En Küçük Medyan Kareler (Least Median of Squares-LMS) gibi dayanıklı (robust) tahmin yöntemleri ile karşılaştırılmıştır. Tahmin yöntemlerinin etkinliğini değerlendirmek amacıyla Monte-Carlo simülasyon tekniği kullanılmış; karşılaştırma ölçütleri olarak ise yan (Bias) ve hata kareler ortalaması (Mean Squared Error-MSE) kriterleri esas alınmıştır. Simülasyon çalışmalarından elde edilen bulguları desteklemek amacıyla, literatürde yer alan gerçek veri seti ile analiz gerçekleştirilmiştir. Hem simülasyon hem de gerçek veri analizi sonuçları, hata terimlerinin normal dağılmadığı veya veri setinde aykırı gözlemlerin bulunduğu durumlarda, LS tahmin edicilerinin performansının önemli ölçüde azaldığını ortaya koymuştur. Buna karşın, dayanıklı tahmin yöntemlerinin bu tür veri bozulmalarına karşı daha kararlı ve güvenilir sonuçlar ürettiği belirlenmiştir. Ayrıca, LMS tahmin edicisinin, model parametrelerinin tahmininde varsayım ihlalleri karşısında diğer yöntemlere kıyasla daha yüksek dayanıklılık gösterdiği sonucuna ulaşılmıştır. WLAD tahmin edicisi ise performans açısından LMS yönteminin ardından en başarılı tahmin yöntemi olarak tespit edilmiş olup, LMS'ye alternatif olarak tercih edilebilecek güvenilir bir seçenek durumundadır. Sonuç olarak, veri setinin aykırı değerler içermesi veya hata terimlerinin normal dağılmaması durumlarında, model parametrelerinin tahmininde klasik LS tahmin yöntemi yerine LMS dayanıklı tahmin yönteminin kullanılması önerilmektedir.
Deney tasarımında optimal blok yapıları
Deney tasarımı teorisinde bloklama, sistematik gürültüyü azaltmak ve etki tahmininin doğruluğunu artırmak amacıyla yaygın olarak kullanılmaktadır. Tasarımların optimal yolla nasıl bloklanacağı ise, uygulamada büyük önem taşımaktadır.Bu çalışmada, çok etkenli ve 2 düzeyli kesirli çok etkenli tasarımlar hakkında bilgi verilmiş; çözüm ve en az sapma kavramları, tanımlayıcı bağıntı alt grupları, blok ve deneme kelime uzunluğu yapıları tanıtılmıştır.Bloklanmış kesirli ve çok etkenli tasarımların seçiminde kullanılan var olan optimallik ölçütleri araştırılmış; tasarımları tanımlayıcı bağıntı alt grupları ve kelime uzunluğu yapılarını kullanmadan karşılaştıran en az moment sapma ölçütü incelenmiştir.Çalışmanın uygulamasında; Sun, Wu ve Chen'in optimal blok yapıları katalogundaki tasarımlar incelenmiş; optimal blok yapısının bulunması için en iyi yöntem seçilmiş ve nedenleri açıklanmıştır.Anahtar Kelimeler: Çok Etkenli Tasarımlar, Kesirli Çok Etkenli Tasarımlar, Kelime Uzunluğu Yapıları, En Az Sapma, Optimal Blok Yapısı.
Türkiye'de orta ölçekli bankaların veri zarflama analizi ile etkinlik ölçümü uygulaması
Bir ülkede ekonomik büyümenin ön koşullarından birisi güçlü ve sağlıklı bankacılık sektörünün olmasıdır. Bu nedenle bankaların etkinliği önem kazanmaktadır. Bu etkinlik gerek yatırımcılar, gerek politikacılar, gerekse de banka üst yönetimleri tarafından takip edilmekte, karar aşamasında da dikkate alınmaktadır. Bunun yanında bankacılık sektörü de kendi içinde rekabet içinde bulunmaktadır. Sektörün rekabet koşulları; sunulan hizmet kalitesi, kaynak ihtiyacı ve kar beklentisinin artışı gibi konuları içermektedir. Son yıllarda Türk Bankacılık Sektöründe nde özellikle; banka sayısı, personel sayısı, teknolojik yatırım, hizmet sayısı ve karlılık konusunda ilerleme kaydedilmiştir. Söz konusu ilerlemenin başlıca sebebi, sektörün geçmiş yıllarda yaşamış olduğu krizlerin sektörü tecrübeli hale getirmiş olmasıdır. Bankaların verimlilik kriterlerini esas alan çalışma ilkesi de bu gelişmede önemli bir rol oynamıştır. Ekonomik sistemde önemli bir yer tutan bankaların etkin çalışmaları büyük önem arz etmektedir. Bu çalışmada Türkiye'de bulunan orta ölçekli ve özel sermayeli (yerli –yabancı) on iki adet bankaya VZA ile etkinlik analizi uygulanmıştır. Analiz sonucunda etkin olan bankalar tespit edilmiş, etkin olmayan bankaların etkin hale gelebilmeleri için etkinleşme önerileri sunulmuştur.
İki değişkenli yaşam verilerinin kopulaya dayalı modellemesi ve analizi
Modelling dependence structure of a bivariate survival data is one of the main issues in biomedical studies. Copulas are key tools to analyze the dependence structures. A bivariate survival function can be expressed as the composition of marginal survival functions and a bivariate copula. Since a survival copula is a great deal of flexibility in modelling bivariate survival data, it provides an effective approach for understanding and modelling the dependent random variables and so the dependence structure. Survival copula deals with a lifetime data and is used for modelling and understanding the distributional structure. In survival studies, the researcher can come across censored survival data. In this study, we consider modelling and analyzing the bivariate survival data in the presence of right censoring using Archimedean copula functions. We use Emura et al. (2010) goodness-of-fit testing procedure for the model selection. Throughout the model selection procedure, we obtain the goodness-of-fit statistics for Gumbel, Frank and Clayton copula models. First, we examine the heart transplant data and model the dependence structure between waiting time for transplant and post-transplant survival time to see the co-movements of these variables. Second, we examine the diabetic retinopathy data and model the dependence between the survival times of the two eyes of the same patient in case of laser photocoagulation treatment. Finally, we use the survival hazard scenario approach to evaluate the probability of exceeding some critical layers. We develop R code to implement the study.
Dayanikli doğrusal olmayan regresyon yöntemleri
Regression analysis is a statistical method for modelling the relationship between two or more variables. Ordinary least squares regression, which is the most commonly used approach to determine relationship between variables may be misleading in the presence of outliers, or when there is heteroscedasticity and non-normality. Even if the underlying assumptions such as normality and homoscedasticity are hold, the relationship between dependent variable and independent variable(s) may be nonlinear. In such cases, using nonparametric regression methods is more appropriate since the shape of the regression function is not needed to be predefined and there is no significant assumptions as in parametric regression situation. The nonparametric regression estimators are called as smoothers and in this thesis four of them are investigated: Kernel smoothing, locally weighted scatter plot smoothing (LOWESS), the running interval smoother (RIS) and constrained b-spline smoothing (COBS). While the running interval smoother predicts the dependent variable by using different location estimators, COBS predicts by using quantiles and the other methods predict the dependent variable using weighted mean. Robust nonlinear regression methods are also blended to create alternative methods. The smoothers and these alternative methods are compared with a simulation study by using theoretical distributions. Furthermore, the methods are examined graphically to understand how the methods can model the relationship between variables. The predicted values of dependent variable corresponding to new observations are calculated as well. COBS and RIS with NO estimator outperformed the other methods in terms of mean squared error (MSE).
Ortalamada kaymalar olduğumda parametrik olmayan CUSUM ve EWMA kontrol kartlarının performanslarının karşılaştırılması
Generally, Shewhart control charts, which require normality hypothesis, are used in monitoring process mean. However, Shewhart control charts may not show adequate performance when the distribution of the process in question does not suit the normal distribution and when there are small shifts in the process mean. For this reason, in cases in which the process distribution is not known, it is more beneficial to use nonparametric control charts that do not require any hypothesis about the distribution. In addition, if there are shifts less than 1.5 sigma, which can be defined as small in the process, preferring the CUSUM (cumulative sum) and EWMA (exponentially weighted moving average) charts, developed as alternatives to Shewhart control charts, would yield more accurate results. In this study, the nonparametric CUSUM control chart and the nonparametric EWMA control chart, designed with the change-point model and the Mann-Whitney Statistic, were introduced and the simulation study was conducted using the R statistical programming language. In this simulation study, data from four different distributions were generated and the average run length (ARL) values for both control charts were calculated. As a result of the calculated ARL values, both control charts were compared in terms of performance and it was observed that the CUSUM chart performed better than the EWMA chart for all distributions applied under the determined conditions.
Denetimsiz anomali tespit algoritmaları
Detection of outliers or anomalies in the data is of great importance in data analysis. Different approaches can be used in anomaly detection according to type of the problem. Unsupervised anomaly detection (UAD) approach is the most challengeable part of these approaches. UAD methods aim to detect anomalies without using a labelled training dataset. UAD algorithms can be considered in three main groups: nearest neighbour based, clustering based and statistical based. In this thesis, UAD approaches is examined and an adjustment, that depends on sample size, is proposed for statistical based algorithm, HBOS. In the first part of the application, performance of the most widely used UAD algorithms, that are k-nearest neighbour (k-NN), local outlier factor (LOF), local density cluster-based outlier factor (LDCOF) and histogram-based outlier score (HBOS), are compared. According to the results, HBOS algorithm is found more successful in terms of accuracy rate and runtime. In the second part of the application, effect of the bin-width determination techniques on to the performance of HBOS algorithm is examined. According to the results of the comparison with multivariate data of different characteristics, there is no superiority between the bin-width determination techniques in terms of accuracy.
Risk ayarlı hastane ölüm tahmin modeli: Bir Türk eğitim ve araştırma hastanesinde uygulama örneği
In today's world, health organizations give much importance to quality and patient safety. To this end, conservation of life and prevent excessive deaths are one of the vital objectives for health services in all countries (Whalley, 2010). Although main function of hospitals is to save lives, there is a little attention to hospital mortality (Champbell et al., 2011). In this context; generating reliable mortality ratio then monitoring them are a prerequisite for improvement in care and development in patient safety. This study aimed to demonstrate the applicability of risk adjusted mortality ratio in Turkey. This is the first study conducted in this field in Turkey. To this end, various risk adjusted hospital mortality prediction models were developed by using some popular data mining techniques; logistic re-gression, decision trees, random forests and artificial neural networks. The data from 30182 inpatients of one of the Turkish training and research hospitals with 1155 beds were used. The data collected from inpatients whose discharge period was January to November in 2014. At the end, the performance of these methods were compared.
Yapısal kırılmalar olduğunda uyarlanmış ve basit üstel düzeltme yöntemlerinin karşılaştırılması
The essential aim of the time series modelling is applied for the forecasting as well as the examination of correlation. One of the most widely used methods in the literature is exponential smoothing (ES) methods. It is a method preferred by many researchers because of its easy application, calculation efficiency, high accuracy and automatic prediction. Such as policy changes, financial crises, natural disasters in the data production processes of the series, permanent structural changes can change affect model parameters as well as analysis results. Having no consideration of such affects, leads parameter estimator to be biased, tests tend to be useless in terms of power and incorrect modelling arise. The main purpose of this study is to compare the predictive performances of the newly developed Modified Exponential Smoothing (MSES) (2016) methods with the simple exponential smoothing (SES) when there are structural breaks in the series. Determining the initial value and misspecification in the selection of the optimum smoothing parameter, as a disadvantage, adversely affect the estimation results. The MSES method gives more weight to the current observations on the series, so that the predictions that are calculated give better performance than the classical method. The MSES method against structural break has not been examined yet. In this study, received from the Central Bank of the Republic of Turkey and traded on the Istanbul Gold Exchange "weighted average price of gold (TL/kg)" data are used. This data set with different break points compares the forecast performance of MSES and SES methods.
Archimedean kopula modelleri için yeni bir uyum iyiliği yaklaşımı ve güç analizi
In this thesis, a new class of bivariate multi-parameter Archimedean copula based on Kendall distribution using Bernstein-Bezier polynomials is introduced. This new class copula has flexible dependence properties depending on the polynomial degree and the control points. Some dependence characteristics such as Kendall's tau, upper tail and lower tail dependence of the this new Archimedean copula class are derived. The simulation procedure based on these desired dependence characteristics is presented. Also, we propose an estimation method for the Archimedean family of copula in a nonparametric setting. Bernstein polynomials and Bezier curve approaches are used to estimate the Kendall distribution function of the Archimedean copula. Also, a new goodness-of-fit test based on Cramer-Von Mises type statistic is constructed using new estimation methods of Kendall distribution function. A Monte Carlo study is performed to measure the performance of the proposed tests. The simulation results show that both the Bernstein polynomial and the B\'ezier curve based tests have better performances than the classical one since they have flexible form according to its order m.
Regresyon ve dağılım tahmininde geometrik modelleme
The contribution of this thesis to inferential statistics by using geometric modeling methods is comprised of three parts. First, we propose a new nonparametric regression model that uses the rational Bézier curves. The main advantages of the proposed regression model are the flexibility over parametric models and having no restriction on the assumptions in regression function estimation. Moreover, rational Bézier curves provide many conveniences such as more control over the shape of a curve and projective invariance. We provide numerical examples to demonstrate superiority of the proposed model and verify it. Secondly, we introduce a new distribution function estimation method using the rational Bernstein polynomials (RBP). The performance of the proposed new method is compared with Bernstein polynomials and empirical distribution function methods using simulation. The new method guarantees monotone nondecreasing function by applying linear constraints on the coefficients of the rational Bernstein basis functions and smooth the empirical distribution function. Furthermore, as a special case, it reduces to Bernstein polynomial estimator method. Some theoretical properties of the new estimator are investigated. Simulation study shows that the proposed estimator is preferable to the Bernstein polynomials and empirical distribution function estimator methods. Thirdly, we introduce a unimodal density estimator based on the Schoenberg's spline operator and thus generalize that of Bernstein polynomials and the beta density. The advantage of this method is the local property. That is, refining the knots while keeping the degree fixed of B-splines yields better estimates. We also give a numerical example to verify our results.
İki bağımsız örneklem için dayanıklı yaklaşımların analizi
Comparing two independent groups has been the subject of many classic studies in applied statistics. Although the most common approach for comparing two independent groups is on the basis of just one reference point, it may provide misleading results and cause missing the important differences among subpopulations of the groups. Determination of the differences in tails of the groups, namely quantiles, can also be of particular concern. This thesis aims to compare two independent groups from a different point of view by utilizing quantiles. Therefore, firstly attention is focused on quantile estimation theory and the newly proposed NO quantile estimator is introduced. NO quantile estimator is computationally practical, has desirable asymptotic properties and is more efficient than some quantile estimators that are commonly used. Following this, NO quantile estimator is used in conjunction with a percentile bootstrap to propose a method that compares two independent groups through multiple quantiles simultaneously. By conducting a series of simulation studies and using some real datasets, the performances of NO quantile estimator and the proposed method for comparing two independent groups are investigated. The proposed method could control the actual type I error rates in almost all cases that are investigated. This method is more preferable than the conventional approaches that considers a single measure of location since it makes a more sensitive comparison by using more than one reference points.
Sıralı küme örneklemesinde sıralama hata modelleri, maliyet ve en uygun küme büyüklüğü
Ranked Set Sampling (RSS) is a sampling method commonly used in recent years. RSS is developed as an alternative to Simple Random Sampling (SRS) in order to estimate population parameters more efficiently where the measurement of sampling units is difficult or costly but the units are easier to rank. There are several factors that make this method useful especially for studies in medicine, agriculture, forestry and ecology. The most important of these factors are the set size and the relative costs of some operations such as sampling, measurement and ranking. Ranking of the units in the set is made on the basis of the visual judgment of the researcher or a concomitant variable which has a strong correlation with the variable of interest. These ranking methods are defined as ranking error models. In this thesis, the widely used cost and ranking error models in RSS literature are investigated. Also, it is aimed to explore the effect of ranking error models on the mean estimator based on RSS and some of its modified methods for different distribution, set and cycle size in infinite population. Besides, it is aimed to examine whether RSS is cost effiective with respect to SRS in terms of mean squared error of the mean estimator considering ranking error models and the N-KPST cost model in infinite population and if so, to determine the optimal set size for RSS. Monte Carlo simulation studies are conducted for these purposes. Additionally, the study is supported by real life data.