Ondokuz Mayıs University
Discipline

Statistics

Ondokuz Mayıs University

108

Archived Theses

0

DOIs Assigned

0%

DOI Rate

Discipline

50 Theses
Master'sOpen AccessTR

Ortak değişkenlerin varlığı durumunda faktöriyel tasarımlar

Bu çalışmada ortak değişkene sahip faktöriyel tasarımlar ele alınmıştır. Bu tasarımlar için klasik teori, normallik varsayımına dayalı elde edilir. Buna karşın, hata terimleri normal dağılıma sahip değilse model parametrelerinin en küçük kareler (EKK) tahmin edicilerinin etkinliğinin ve test istatistiklerinin gücü ve istatistiksel sağlamlığının azaldığı Monte-Carlo simulasyon çalışması yardımıyla gösterilmiştir. Bu sonuçlar ortak değişkene sahip faktöriyel tasarımlarda normallik varsayımı geçerli olmadığı zaman, EKK ya alternatif olarak, daha etkin tahmin ediciler ile daha güçlü ve istatistiksel olarak sağlam test istatistiklerinin elde edilmesini gerektirir.Bu nedenlerle ortak değişkene sahip faktöriyel tasarımlarda hata terimlerinin bağımsız ve özdeş olarak uzun kuyruklu simetrik (LTS) dağılıma sahip olduğu var\-sayıl\-mıştır. Bu durumda en çok olabilirlik tahmin edicileri analitik olarak elde edilemediğinden uyarlanmış en çok olabilirlik yöntemi kullanılmış ve bu yönteme dayalı olarak parametrelerin tahmin edicileri açık formüllerle ifade edilmiştir. Uyar\-lan\-mış en çok olabilirlik (UEÇO) tahmin edicilerinin EKK tahmin edicilerinden daha etkin olduğu Monte-Carlo simulasyon çalışmasıyla gösterilmiştir. UEÇO tahmin edicilerine dayalı geliştirilen test istatistiklerinin asimptotik olarak F dağılımına sahip olduğunu kanıt\-lanmış ve küçük örneklem hacimleri için de bu test istatistiklerinin F dağılımına sahip olduğu Monte-Carlo simulasyon çalışması yardımıyla gösterilmiştir. Buna ek olarak UEÇO tahmin edicilerine dayalı test istatistiklerinin klasik test istatistiklerinden daha güçlü ve istatistiksel olarak sağlam olduğu Monte-Carlo simulasyon çalışmasıyla gösterilmiştir. Geliştirilen yöntem bir gerçek hayat örneği üzerinde uygulanmıştır.

Şükrü Acıtaş
Anadolu University · Institute of Graduate Studies in Science
2010
00
Master'sOpen AccessTR

Bulanık mantık çıkarım sistemi ile tren hızının otomatik kontrolü

Bu tezde, bulanık mantık ve adaptif ağ yapısına dayalı bulanık çıkarım sistemi (Anfis) ile Otomatik Tren Yönetim Sistemi uygulaması yapılmıştır. Öncelikle Otomatik Tren Yönetim Sistemi üzerine yapılan teorik araştırmalardan yola çıkılarak, trenlerin bulanık mantık ile kontrol ve yönetimi sağlanmıştır. Daha sonra mevcut teknik bir adım daha ileriye götürülerek, yapay sinir ağları ve bulanık mantığın bir arada kullanıldığı Anfis ile (dünyada ilk kez denenerek) modellenmiş ve simülasyonu gerçekleştirilmiştir. Bulanık mantık ve Anfis kontrolünün birlikte uygulanması ile Otomatik Tren Yönetim Sistemi gerçek hayatta her bölge ve koşulda çalışabilecek, güvenliği en üst düzeyde olan bir sistem haline getirilmeye çalışılmış ve elde edilen sonuçlar sunulmuştur.

ANFISBulanık mantıkYapay sinir ağları
Erhan Öztok
Anadolu University · Institute of Graduate Studies in Science
2010
00
Master'sOpen AccessTR

Stokastik diferensiyel denklemlerle modelleme

Bu tez çalışmasında stokastik diferansiyel denklemlerin iki alt denklem sınıfı olan rassal diferansiyel denklemler ve Itô stokastik diferansiyel denklemleri ile somut problemlerin modellenmesi konusu ele alınmıştır. Bu amaçla öncelikle stokastik diferansiyel denklemler teorisi için gerekli olan matematiksel temeller verilmiş, iki somut problem için stokastik diferansiyel denklem modelleri kurulmuş veçözümlenmiştir.İlk olarak, bir radyoaktif bozunma problemi bir rassal diferansiyel denklem ile modellenmiş ve bu modelin çözümü elde edilmiştir. Daha sonra çözüme ilişkin bazı çıkarsamalar yapılmış ve elde edilen sonuçlar, tablo ve şekiller ile sunulmuştur.İkinci olarak ise hisse senetleri fiyatları için Samuelson modeli tanıtıldıktan sonra bu çalışmada verilen modelleme yöntemiyle fiyatlar için bir stokastik diferansiyel denklem modeli kurulmuştur. İki modeli fiyat tahminleri bazında karşılaştırabilmek için MOTOROLA hisse senedinin 20.03.09-11.05.09 tarihleri arası günlük kapanış fiyatları veri seti ele alınmıştır. MATLAB dilinde yazılan programlar yardımıyla her iki modelin parametreleri tahmin edilmiştir. Daha sonra, elde edilen modeller nümerik olarak çözdürülmüş ve hesaplanan tahminler tablolar ve şekiller yardımıyla ortaya konulmuştur. Son olarak her iki model Euclid metriğine göre karşılaştırılmış ve Samuelson modelinin çalışmada verilen yöntemle elde edilen modelden daha iyi sonuçlar verdiği görülmüştür.

Stokastik süreçler
Batuhan Bozdağ
Anadolu University · Institute of Graduate Studies in Science
2010
00
DoctorateOpen AccessTR

Mekansal istatistikte nokta örüntü teknikleri ve bir uygulama

Mekansal istatistik, mekansal olarak düzenlenmiş verileri dikkate alan istatistiksel yöntemlerle ilgilidir. Mekansal verinin içinde mevcut olan mekansal bağımlılık varsayımı yüzünden mekansal istatistik, klasik istatistik yöntem ve tekniklerini kullanma eğiliminde değildir. Klasik istatistikte olduğu gibi, mekansal veri için tanımlayıcı ve çıkarsamalı yaklaşımlara sahip olmakla birlikte kendine özgü yöntem ve tekniklere sahiptir. Mekansal istatistikte kullanılan yöntemler genellikle analiz edilmekte olan mekansal verinin türlerine göre üç kategoriye ayrılmaktadır. Mekansal verinin bu türlerinden biri de mekansal nokta örüntü verileridir. Mekansal nokta örüntü verileri, nokta olayların konumlarından elde edilmiş verilerdir. Birbirleriyle ilişkili konumların, anlamlı bir örüntüyü temsil edip etmediği ile ilgilenilmektedir. Mekansal nokta örüntüler analiz edilirken, temel olarak tam mekansal rassalığa karşılık örüntülerin kümelenme ve düzenlilik gösterip göstermediği ile ilgilenilmektedir. Tam mekansal rassallıktan herhangi bir sapmanın değerlendirilmesine olanak sağlayan dağılım fonksiyonlarının tahminlerinin, tam mekansal rassallık altında bir dağılım ile karşılaştırılmasında bazı simülasyon teknikleri kullanılmaktadır.Bu çalışmada mekansal istatistik yöntem ve tekniklerinin deprem verisi üzerinde uygulanması ele alınmaktadır ve bu amaçla ülkemizin yakın geçmişte büyük bir deprem ile sarsılmış olan Gölcük bölgesi seçilmiştir. Bu bölgede meydana gelmiş depremler yalnızca istatistiksel veri olarak ele alınarak, mekansal istatistik yöntem ve teknikler aracılığıyla bu bölgede olabilecek depremler için benzetim çalışması 1900 yılından 01 Ocak 2010 tarihine kadar elde edilen deprem verileri yardımıyla türetilmiştir.

DepremMekansal analizMekansal süreçler
Halil Eryılmaz
Anadolu University · Institute of Graduate Studies in Science
2010
00
Master'sOpen AccessTR

Ridge ve Liu tahmincilerinin etkinliklerinin ve yanlılıklarının karşılaştırılması

Çoklu regresyon analizinde karşılaşılan sorunlardan birisi de çoklu bağıntı durumudur. Bağımsız değişkenlerden bir veya birkaçının diğer bağımsız değişkenler tarafından iyi açıklandığı zaman, sonuçlarda istenmeyen özellikler oluşturan çoklu bağıntı sorunu meydana gelmektedir. Çoklu bağıntıyı gidermek veya azaltmak için yanlı tahmin yöntemleri kullanılır. Bu çalışmada, yanlı tahmin yöntemleri olarak bilinen Ridge ve Liu tahmincilerinin karşılaştırılmasına yer verilmiştir. İlk olarak bu iki tahminci tanımlanmış ve bununla ilgili 1985?2006 yılları arasında Türkiye'deki turizm geliri fonksiyonunu açıklayan değişkenler olarak; yatak kapasitesi, turist sayısı, seyahat acentelerin sayısı, yabancı sermaye miktarı, Euro cinsi döviz kuru, ABD doları cinsi döviz kuru üzerine bir uygulama yapılmıştır. Sonuç olarak, bu iki yöntem etkinlikleri ve yanlılıkları bakımından karşılaştırılmış ve sonuçlar yorumlanmıştır.Anahtar Kelimeler: Çoklu Bağıntı, Ridge Tahmincisi, Liu Tahmincisi, Yanlı Tahmin Yöntemleri, Turizm Geliri

Emine Karakaya
Anadolu University · Institute of Graduate Studies in Science
2011
00
Master'sOpen AccessTR

ARCH modelleriyle bazı ülkelerin döviz kurlarının volatilitesinin incelenmesi

Finansal serilerde, taşıdıkları özellikler nedeniyle doğrusal zaman serisi yerine, doğrusal olmayan koşullu değişen varyans modellerinin kullanılması giderek daha yaygın hale gelmiştir. Öngörü hataları varyansının sabit olmadığı, değişen varyansa sahip olduğu zaman serisinin çözümlenmesinde serilerin bu özelliğini de dikkate alacak modellere gereksinim duyulmuştur. Robert F. Engle (1982), geçerliliği olamayan yukarıda belirtilen varsayımı genelleştirmiş ve Otoregresif Koşullu Değişen Varyans (ARCH) süreçleri olarak adlandırılan stokastik süreçlerin yeni bir sınıfını önermiştir. Bu çalışmada, bazı ARCH (GARCH, GARCH-M ve EGARCH, TGARCH) modellerinin istatistiksel özellikleri ve tahmin yöntemleri incelenmiş, bu modeller farklı gelişmişlik düzeylerindeki rasgele seçilen on ülkenin döviz kuru serilerine uygulanmıştır. Model sonuçları karşılaştırılarak serilere en uygun koşullu varyans modeli belirlenmiştir.

Zeynep Özgün
Anadolu University · Institute of Graduate Studies in Science
2011
00
Master'sOpen AccessTR

İlköğretim 7. ve 8. sınıf öğrencilerine yönelik negatif tamsayılara ilişkin tutum ölçeği geliştirilmesi ve lojistik regresyonla analizi

Matematik dersine karşı oluşan tutum ile matematik başarısı arasındaki ilişki literatürde üzerinde sıkça durulan önemli araştırma konularından birisidir. Öğrencinin bir konuya olan başarısının artmasında olumlu tutum geliştirmesi oldukça önemli olmaktadır.Negatif tamsayılar konusu ise, özellikle ilköğretim ikinci kademede öğrenilmeye başlanan ve öğrencilerin gerçek hayata uyarlamakta zorlandıkları ve öğrenmekte zorluk yaşadıkları başlıca konuların arasına girmektedir. Bu tez çalışmasında Negatif Tamsayılara karşı tutum ölçeği oluşturulmaya çalışılmıştır. Taslak ölçek 7. ve 8. sınıf 220 öğrenciye uygulanmış, faktör analizi uygulanarak yapı geçerliliği ortaya konulmuş, genel güvenilirlik için Cronbach's Alpha güvenirlilik katsayısı hesaplanmıştır. Sonuçlar % 95 güven düzeyinde değerlendirilmiştir. Sonuç olarak 28 madden oluşan tek bir faktör altında toplanan bir tutum ölçeği geliştirilmiştir.Ölçek geliştirildikten sonra ilk önce alınan veriler ile matematik başarı puanları arasındaki ilişki gözlemlenmiş daha sonra ise öğrencilerin tutum ölçeklerinden aldıkları toplam puanlar eşit aralıkta 4'e bölünerek sıralı hale getirilmiş ve öğrencilerin demografik verileri arasındaki ilişki gözlemlenmeye çalışılmıştır. Bağımlı değişkenler sıralı halde olduğu için sıralı lojistik regresyon tekniği tercih edilmiştir.Oluşturulan sıralı lojistik regresyon modellerinden elde edilen sonuçlarda, bir dönem önceki matematik notları ile bir dönem sonraki matematik notları arasında anlamlı bir ilişki gözlenmiştir. Sadece 7. Sınıf verileri ile oluşturulan diğer bir modelde matematik başarı notları ile tutum ölçeği maddelerinin 24 maddesinin değişik kategorilerinin anlamlı bir ilişkide oldukları sonuçlarına varılmıştır. Ayrıca oluşturulan son modelde negatif tam sayılara karşı 7. Sınıfların 8. Sınıflara göre daha olumlu tutum gösterme eğiliminde oldukları ve okul dışında eğitim yardımı alan öğrencilerinde diğer öğrencilere göre daha olumlu tutum gösterdikleri sonucuna ulaşılmıştır.Anahtar kelimeler: Tutum ölçeği, negatif tamsayılar, sıralı lojistik regresyon

Ölçekler
Yasin Memiş
Anadolu University · Institute of Graduate Studies in Science
2012
00
Master'sOpen AccessTR

Doğrusal regresyon modeli için m-tahmincilerin incelenmesi

Robust (Sağlam) regresyon tahmincileri, hataların normal dağılıma uymadığı veya veri setinde aykırı değer bulunması durumunda regresyon modelini en güvenilir şekilde tahmin etmek amacı ile geliştirilmiştir.Bu tez çalışmasının amacı, veri setinde aykırı değer olması durumunda en küçük kareler tahmincisine alternatif olarak geliştirilen robust regresyon tahmincilerinden M-tahmincilerin çeşitli açılardan incelenmesidir.İlk olarak M-Tahmincilerin hesaplanmasında kullanılan yeniden ağırlıklandırılmış en küçük kareler algoritmasının başlangıç tahminlerinin seçimine olan duyarlılığı ele alınmış ve M-tahmincilerinin kırılma noktaları, grafikler yardımıyla incelenmiştir. Daha sonra hata teriminin dağılımının normal ve normalden farklı olduğu durumlar için M-tahmincilerinin etkinlik açısından performansı değerlendirilmiş ve son olarak başlangıç ölçek tahmincisinin etkinliğe katkısı araştırılmıştır. Ayrıca reel yaşamdan alınan iki örnek üzerinde M-tahminciler uygulanmış ve elde edilen sonuçlar tartışılmıştır.

Aykırı değerlerDoğrusal regresyonEn küçük kareler yöntemi+1
Vural Yıldırım
Anadolu University · Institute of Graduate Studies in Science
2012
00
DoctorateOpen AccessTR

Parametrik olmayan bulanık regresyon modelleri analizi

Tez çalışmasında, parametrik olmayan bulanık regresyon modelleri incelenmiştir. k-en yakın komşuluk, çekirdek düzeltme ve yerel polinomiyal düzeltme modelleri bulanık yapıda ifade edilmiştir. Ayrıca bu modeller için, band genişliği seçiminde çapraz geçerlilik ve genelleştirilmiş çapraz geçerlilik kriterleri, şapka matrisi kullanılarak geliştirilmiştir. Çalışmalarda kullanılan verilerde, yanıt değişkenin bulanık değerli ve açıklayıcı değişkenin kesin değerli olduğu durum dikkate alınmıştır. Analizler R programında locpol paketi kullanılarak yapılmıştır. Uygulamalarda özellikle, bulanık yerel polinomiyal regresyon modelinde polinomun derecesinin doğrusal ve kübik alındığı durumlar üzerinde inceleme yapılmıştır. Bu modeller için band genişliği, geliştirilen çapraz geçerlilik ve genelleştirilmiş çapraz geçerlilik kriterleri ile seçilerek, modellerin performansları ortalama karesel hata değerleri kullanılarak karşılaştırılmıştır. Elde edilen sonuçlardan genelleştirme yapabilmek amacıyla farklı türdeki veri setleri üzerinde çalışılmıştır. Uygulamaların pek çoğunda bulanık yerel kübik modele ait performans değerleri daha yüksek olmasına rağmen, özellikle eğriselliği fazla olan modelleri daha pürüzsüz (dalgalanması az olan) bir şekilde ifade eder. Ayrıca, band genişliği değerinin bulanık yerel doğrusal modele göre daha yüksek seçilmesi ile işlem basamaklarını azalttığı göz önüne alındığında, bulanık yerel kübik modellerin kullanımının daha faydalı olduğu sonucuna varılmıştır.

Münevvere Yıldız
Anadolu University · Institute of Graduate Studies in Science
2013
00
Master'sOpen AccessTR

Türkiye'deki istatistik bölümlerinin göreli etkinliklerinin veri zarflama analizi ile belirlenmesi

Bir eğitim sisteminin başarısı, eğitim süreci içinde sunduğu olanaklar ve eğitim sonrası elde edilen yeterlilikler tarafından belirlenebilmektedir. Bu bağlamda ülkemizdeki üniversite performanslarının karşılaştırılmaları genellikle akademisyen performansları, mezunlarının işsizlik oranı ve ülke genelinde yapılan ortak sınav sonuçları üzerinden çıktı yönlü olarak yapıldığı gözlemlenmektedir. Kamu Personeli Seçme Sınavı (KPSS) sonuçları da karşılaştırma unsuru olarak kabul gören bir çıktıdır. Fakat eğitim bilimciler eğitim sitemlerinin performanslarının değerlendirilmesinde sunulan olanakların çıktıya ne kadar verimli olarak dönüştürebildiklerine önem vermektedirler. Bu çalışmada Türkiye?deki devlet üniversitelerinin verilerine ulaşılan 18 adet istatistik bölümü homojen karar verme birimi (KVB) olarak ele alınmış ve bölümlerin göreli etkinlikleri veri zarflama analizi (VZA) yardımı ile hesaplanmıştır. Etkinlik değerleri arasındaki farklılıkların kontrol edilemeyen girdiler tarafından etkilenip etkilenmediği, etkinlik değerleri ve girdi-çıktı değişkenleri arasındaki ilişkiler ve öğretim programlarına göre bölümler arasında etkinlik farkların anlamlılığı ortaya konulmaya çalışılmıştır. Elde edilen sonuçların karar vericilere ve istatistik bölümünde okuyan öğrencilere yol gösterici olması amaçlanmıştır.

Cenk İçöz
Anadolu University · Institute of Graduate Studies in Science
2013
00
DoctorateOpen AccessTR

Dağılım ve sinir ağı tabanlı bulanık zaman serisi modelleri

Bulanık zaman serisi yaklaşımları genelde bulanıklaştırma, bulanık ilişkiler belirleme ve durulaştırma olmak üzere üç aşamadan oluşur. Bu çalışmada, tek değişkenli birinci dereceden sinir ağı tabanlı bulanık zaman serisi öngörüsü için yeni bir yaklaşım ve yeni bir yöntem önerilmiştir. İlk olarak evrensel küme parçalanmasında yapılan aralık uzunluğu belirleme aşamasında sabit bir aralık uzunluğu almak yerine daha etkili olan dağılım tabanlı uzunluk yaklaşımı kullanılmıştır. Bulanıklaştırma aşamasında yeni bir algoritma oluşturularak işlem kolaylığı sağlanmıştır. Ayrıca bu aşamada önerilen yöntemde ilk defa ağırlıklandırılmış indisler kullanılmıştır. Bulanık ilişki belirlemede bütün üyelik derecelerinin ayarlanması sağlanmıştır. Öngörü performansını geliştirmek için sadece Çok Katmanlı Algılayıcı (ÇKA) değil, ayrıca Genelleştirilmiş Regresyon Sinir Ağları (GRSA), ve Radyal Tabanlı Fonksiyon Sinir Ağı (RTFSA) gibi farklı yapay sinir ağı mimarileri de uygulanmıştır. Bu YSA mimarileri için klasik sinir ağı tabanlı bulanık zaman serisi öngörü yöntemlerinden farklı olarak tek katman ve düğüm sayısı, girdi ve çıktıların toplamları olması yerine en iyi sonucu verecek şekilde farklı sayıda gizli katman ve düğüm sayısı kullanılmıştır. Önerilen yöntem ve yaklaşım, oldukça iyi bilinen ve literatürde sıklıkla kullanılan Alabama Üniversitesi kayıt verileri ve ayrıca büyük bir veri seti olarak İ.M.K.B. (BİST) ulusal 100 endeksi verileri 2006-2010 yılları için kullanılarak literatürde önerilmiş sinir ağı tabanlı veya sinir ağı tabanlı olmayan çeşitli bulanık zaman serisi öngörü yöntemleri ile karşılaştırılmıştır. Sonuçlar; önerilen yeni yöntemin, literatürde verilen diğer yöntemlerden üstün olduğunu göstermiştir.

Bulanık kümelerYapay sinir ağları
Özer Özdemir
Anadolu University · Institute of Graduate Studies in Science
2013
00
DoctorateOpen AccessTR

Sağkalım analizinde COX regresyon ve yapay sinir ağları kullanımı

Sağkalım analizinde en yaygın kullanılan regresyon modeli Cox Regresyon Modelidir. Bu çalışmada, sağkalım analizinde cox regresyon modeli ve yapay sinir ağları modelleri kullanılarak Lösemi hastalarına ait veriler analiz edilmiştir. Çalışmada Yemen Cumhuriyeti'ndeki 2017 yılı Ocak ayı ve 2022 yılı Şubat ayı arasında 1168 Lösemi hastasına ait bilgiler kullanılmıştır. Bu tezin amacı, en yaygın hastalıklardan biri olan kanser hastalığı ile başa çıkmada Yemen Cumhuriyeti'ndeki lösemi hastalarının sağkalım sürelerini etkileyen en önemli faktörleri (risk faktörleri) belirlemektedir. Ayrıca, sağkalım analizinde Yapay Sinir Ağlarını kullanmak ve Yapay Sinir Ağları modeli ile Cox Regresyon Modeli arasındaki performans göstergelerini karşılaştırmaktır. Söz konusu amaçla Cox Regresyon Modeli ve Yapay Sinir Ağları modelleri sağkalım verilerinin analizinde kullanılmıştır. Önerilen Cox modellerinden en iyi model seçilmiş ve ardından aynı veriler yapay sinir ağları kullanılarak analiz edilip modellerden en iyi olanı da seçilmiştir. Cox regresyon modeli sonuçları ile yapay sinir ağlarının kullanımından elde edilen sonuçlar arasında karşılaştırma yapılmıştır. İki yöntem arasındaki karşılaştırma, Hata Kareler Ortalaması (MSE) ve Ortalama mutlak hata (MAE) kriterlerine göre gerçekleştirilmiştir. Cox regresyon modellerinden en iyi model ve yapay sinir ağı modellerinden en iyi model seçilip karşılaştırdıktan sonra yapay sinir ağları modelinin, daha önce belirtilen kriterlere göre sağkalım verilerini analiz etmekte Cox regresyon modelinden daha iyi performans gösterdiği soncuna varılmıştır.

Elıas Abdullah Al-samaı
Anadolu University · Institute of Graduate Studies in Science
2024
00
Master'sOpen AccessTR

Makine öğrenimi teknikleri ile güneş ışınımı ve güç kestirimi

Son yıllarda iklim değişikliği ve özellikle küresel ısınmanın etkileri görüldükçe, enerji tasarrufu ve yenilenebilir enerjinin önemi giderek artmıştır. Yenilenebilir enerji alanında başta güneş enerjisi ve rüzgar enerjisi olmak üzere jeotermal, biyokütle, dalga ve hidroelektrik enerji gibi pek çok enerji sistemi uygulaması mevcuttur. Yenilenebilir enerji sistemleri kurulmadan önce ekonomik ve çevresel sürdürülebilirliklerinin incelenmesi sistemlerden optimum verim elde etmek amacıyla önem taşımaktadır. Bu doğrultuda yenilenebilir enerji sistemlerinde güvenilirlik ve enerji tahmini çalışmaları sürdürülebilirlik bağlamında etkili çalışma alanları arasına girmektedir. Enerji sistemlerin geleceğe yönelik güvenilirlik ve enerji tahmininde makine öğrenimi metotları kullanılmaktadır. Yapılan çalışmada da kurulması planlanan bir güneş enerjisi sisteminin enerji tahmininde güneş ışınımı tahmini için Yapay Sinir Ağları, LightGBM (Hafif Gradyan Artırma Makineleri) ve LSTM (Uzun Kısa Zamanlı Bellek) gibi makine öğrenimi modelleri kullanılmıştır. Kullanılan makine öğrenimi modellerinin tahmin performansları çeşitli hata metrikleri ile karşılaştırılarak en başarılı sonuçları veren LightGBM modeli seçilmiştir. Seçilen modelin test verisi üzerinde Ortalama Mutlak Hata Yüzdesi (MAPE) değeri %21,29 olmuştur. Seçilen LightGBM modeli üzerinden yapılan güneş ışınımı tahminleri kullanılarak belirli parametrelerdeki bir güneş paneli sisteminin sonraki güne ait saatlik güç ve günlük enerji üretim kestirimi yapılmıştır.

Enerji sistemleri
Can Yunus Erözden
Dokuz Eylül University · Institute of Graduate Studies in Science
2023
00
Master'sOpen AccessTR

Çamaşır makinelerinde farklı leke grupları için yıkama performansına etki eden parametrelerin denetimli öğrenme algoritmaları ile tahminlenmesi

Ev tipi çamaşır makinelerinin yıkama performansının, Avrupa'daki eko tasarım gereklilikleri gibi belirli sınır değerlere uyması gerekmektedir. Çamaşır makinelerinde lekelerin yıkama performansları için belirlenen sınır değerleri EN 60456 standardına göre elde edilmelidir. Yıkama performansını etkileyen yıkama süresi, su seviyesi, deterjan miktarı, motor yoğunluğu, hacim başına yıkanan yük miktarı ve sıcaklık olarak belirlenen faktörlerin çeşitli düzeyleri için yıkamalar gerçekleştirilmektedir. Bu tez çalışmasında Vestel Beyaz Eşya firmasından sebum, karbon, kan, kakao ve kırmızı şarap lekesi olan kumaşlar üzerinde çeyrek yükle, yarım yükle ve tam yükle elde edilen yıkama performans değerleri temin edilmiştir. 5 leke için yıkama performans değerleri ile tüm lekelerin yıkama performans değeri toplamı ve 3 ayrı yük için toplam 18 veri kümesi elde edilmiştir. Elde edilen veri kümeleri çoklu doğrusal regresyon, regresyon ağacı, torbalama (bagging), rassal ormanlar regresyonu, XGBoost regresyonu, destek vektör regresyonu (doğrusal çekirdek ve polinom çekirdek) ve k-en yakın komşu regresyonu algoritmaları ile modellenmiştir. Tüm modellerin performanslarının karşılaştırılması için hata kareler ortalaması, kök hata kareler ortalaması ve ortalama mutlak hata metrikleri kullanılmıştır. 18 veri kümesinin 8'inde rassal ormanlar regresyonu, 7'sinde bagging ve 3'ünde XGBoost modeli en iyi model olarak tespit edilmiştir. Sonuç olarak ağaç tabanlı algoritmalar kullanan modeller olarak ön plana çıkmıştır.

Merih Şükrü Akgün
Dokuz Eylül University · Institute of Graduate Studies in Science
2023
00
Master'sOpen AccessTR

Sıralı küme örneklemesine dayalı makine öğrenmesi teknikleri

Son yıllardaki hızlı veri artışı ve bu verileri analiz etmenin giderek zorlaşmasıyla birlikte, günümüzde birçok çalışma alanında makine öğrenmesi kullanılmaktadır. Makine öğrenmesi, insan beyninin öğrenme mantığını esas alarak, çeşitli algoritma ve teknikler geliştirmeyi amaçlayan bilimsel çalışma alanıdır. Makine öğrenmesi algoritmaları, olayları inceler ve nasıl meydana geldiklerini anlamaya çalışır. Bu çabaları sonucunda, elde ettikleri sonuçlar ile genelleme yapma yeteneği kazanırlar. Makine öğrenmesi algoritmalarının bilgileri öğrenmeleri ve ne kadar iyi öğrendiklerinin değerlendirilmesi için; veri kümesi, eğitim ve test seti olarak ikiye ayrılır. Literatürde bu işlem, kullanıcının belirlediği bir oranda rastgele olarak yapılır. Bu tez çalışmasında; veri seti bölme işlemi, literatürde kullanılan yönteme ek olarak, Sıralı Küme Örneklemesi (SKÖ), Uç Değer SKÖ (USKÖ), Medyan SKÖ (MSKÖ) ve Yüzdelik SKÖ (YSKÖ) yöntemleriyle yapılarak, elde edilen sonuçların karşılaştırılması amaçlanmıştır. Farklı çalışma alanlarından seçilen gerçek hayat verileri, belirtilen yöntemler ile eğitim ve test setlerine ayrılmıştır. Eğitim setleri ile makine öğrenmesi algoritmaları eğitilerek, test setleri ile öğrenme başarıları sınanmıştır. Karşılaştırma kriteri olarak; regresyon algoritmalarında hata kareler ortalamasının karekökü (HKOK) değerleri, sınıflandırma algoritmalarında ise doğru sınıflandırma oranları kullanılmıştır.

Makine öğrenmesiSıralı küme örneklemesi
Sena Aslan
Dokuz Eylül University · Institute of Graduate Studies in Science
2022
00
Master'sOpen AccessEN

MAUT ve TAOV yöntemlerinin birleştirilmesiyle Türkiye için en faydalı yenilenebilir enerji alternatifinin belirlenmesi

The need for energy in the world has recently increased as a result of the increasing population, global growth, and industrialization. The failure to meet the increasing energy with the existing fossil-resourced reserves, and the increase in environmental awareness and energy supply security shows that the use of renewable energy resources in Turkey is quite important as well as in the world. In the present study, a new combined MAUT and TAOV multi-criteria decision-making approach (TAOV- MAUT) method is submitted by combining the Multi Attribute Utility Theory (MAUT) and Total Area Based on Orthogonal Vectors (TAOV). It is tried to determine the most useful renewable energy resource for Turkey by using proposed method. According to the literature review on the evaluation of renewable energy sources in Turkey, the most important criteria that can affect the decision were determined as follows: efficiency, construction time, cost, government incentives, economic life, foreign dependency, employment, opportunities, social acceptance, space requirement and greenhouse gas emissions. As a result of the study, the most useful renewable energy sources in Turkey were determined as hydroelectric, wind, solar, biomass and geothermal energy sources, respectively.

Utility theoryRenewable energy resourcesMulti criteria decision making
Burhan Aydın
Dokuz Eylül University · Institute of Graduate Studies in Science
2022
00
DoctorateOpen AccessEN

Akıllı ulaşım sistemleri için ileri istatistiksel yöntemler

The use of machine learning techniques and statistical methods for intelligent transportation systems has gained importance recently. In this thesis, two different methods are proposed for clustering and modeling that can serve the needs of the transportation field. Clustering methods are used to group data points quickly and easily for further analysis of clusters. Also, most clustering methods require complex computations or have an iterative procedure that makes the algorithm time-consuming, especially when the data are relatively large. In this thesis, we propose a new clustering method called spatial adaptive clustering (SAC) based on the idea of adaptive cluster sampling (ACS) design. ACS is a sampling method, which is based on neighborhood search on a grid structure, has an adaptive selection process of units and recursively added units reveal the batched individuals easily and quickly. The SAC algorithm forms clusters based on neighborhood search using grid structures and can detect noise points. The performance of the proposed algorithm is evaluated through comparisons with the results from well-known density-based clustering approaches in the literature using real and artificial data sets. Also, hot spots of accident locations in Birmingham/England and Buca/Turkey are investigated using the SAC algorithm. Additionally, to reduce the number of accidents, hotspot locations are examined in terms of the factors causing the accident. Vehicle headway modeling has also been one of the other important topics for traffic signal optimization and flow modeling. In this thesis, we estimate the parameters of Exponentiated Weibull (EW) distribution using the maximum likelihood method under the assumption that all parameters are unknown. We deal with the performance of ranked set sampling and simple random sampling methods by a simulation study in R-software in terms of mean-squared error. Finally, we illustrate the flexibility and usefulness of EW distribution by analyzing simulated data from a real application study in the transportation field.

Intelligent transportation systemsCluster methodPoint analysis+4
Büşra Güngör
Dokuz Eylül University · Institute of Graduate Studies in Science
2022
00
DoctorateOpen AccessEN

Güneş fotovoltaik sistemlerinin performansı için güvenilirlik modellemesi

Solar energy is one of the most widely used renewable energy sources. Photovoltaic systems directly convert solar energy into electricity with no carbon dioxide emission or any other air pollutants. The power generated by a photovoltaic system depends on the characteristics of solar irradiation and weather conditions. This stochastic nature of power systems prompts the researchers to use probabilistic and statistical techniques. In this thesis, we consider the distribution function of the power generated by a photovoltaic system to make predictions about its characteristics. We model the mean power generated by a photovoltaic system. We also consider that photovoltaic modules may have multistate working conditions and different performance levels depending on solar radiation. In this concept, we present a model for solar power systems with PV modules having various levels of operational performance and we develop a reliability model for the system's power regarding the threshold value that is the minimum required total performance level for the system. This model reflects the performance levels and the working probabilities of PV modules. The problem is evaluated under different conditions regarding the dependency of multistate PV modules. In addition, we provide the optimum number of photovoltaic modules that minimizes the total cost based on the level of required total power production. For further analyses, we give real data applications to estimate the characteristics of the power produced by the solar plant for a specific location in Izmir, Turkey. As the software programming tools, R v.1.2.5033 and Mathematica v.11.3 are used for the computations.

Photovoltaic systemsSolar energy
Melek Esemen
Dokuz Eylül University · Institute of Graduate Studies in Science
2022
00
Master'sOpen AccessEN

ATA öngörü yönteminin ampirik özelliklerinin incelenmesi

Forecasting is important in all scientific fields such as industrial, commercial, medical and economic. There are many forecasting methods in the literature, but exponential smoothing is a very popular method due to its simplicity and accuracy. Simple exponential smoothing is used for data sets randomly distributed around a constant level. Holt's linear trend method is a method that helps to deal with linearly trended data. Despite the fact that exponential smoothing methods are widely used and have been in the literature for a long time, they have some problems that potentially affect the predictive accuracy of models. Ata is a new forecasting method that has been proposed to overcome these problems and to provide better forecasts. In this thesis, the forecasting accuracy of Ata and exponential smoothing will be compared for data sets with no or linear trend. The results given in this study are obtained using simulated data sets with different sample sizes and variances and the forecast accuracy is compared using the mean squared forecast error. In line with these results, the forecast accuracy is calculated for both short and long term forecasting horizons. The results reveal that the proposed approach outperforms exponential smoothing for most types of time series data for both short and long term forecasting horizons.

Beyza Çetin
Dokuz Eylül University · Institute of Graduate Studies in Science
2021
00
DoctorateOpen AccessEN

Farklı bağımlılık varsayımlarına dayalı saklı Markov modelleri

Hidden Markov models are widely used to model the probabilistic structures with latent random variables. The main assumption of hidden Markov models is that; observations are conditionally independent and identically distributed random variables. There may exist some cases where this assumption may not be valid in practice. That is, an observation that occurs in the current state may depend on the previous observation symbol that occurred in the previous state. In this thesis, two types of hidden Markov models are introduced which differ from the classical hidden Markov model based on different first-order Markov dependence assumptions. The introduced models are capable of capturing a possible first-order Markov dependence between the successive observations or successive system informations. They can provide better representation for the appropriate real-life problems where, if the observations have some conditional dependencies among them. The two proposed models are defined with their assumptions and using appropriate notation. Modifications to the algorithms used for parameter estimates and hidden state sequence estimates are explained. In addition, an experimental study is conducted to show the performance of the introduced models compared to the classical hidden Markov model. According to the results of the experimental study, the proposed models outperform in generated observation sequences that have appropriate assumptions. Besides, two different case studies are conducted namely the occurrences of strong earthquakes and daily stock prices. They are modelled with both the classical hidden Markov model and the proposed models, and the results are compared.

Özgür Danışman
Dokuz Eylül University · Institute of Graduate Studies in Science
2021
00
Master'sOpen AccessEN

Sağkalım verileri için makine öğrenmesi yöntemleri

Survival analysis is a the statistical approach methods used in many fields. The key feature that distinguishes this method from other analysis methods is that it can be used when censored observations are present. In the literature, survival analysis can be categorized as traditional survival analysis methods and machine learning based survival analysis methods. The increasing number of data and variables, the existence of censored observations and the fact that traditional methods require some assumptions make it difficult to analyze survival data with traditional methods. In order to cope with this situation, machine learning methods specific to survival data are used. In this thesis, the traditional survival analysis methods such as the Kaplan Meier, Log rank test and the Cox regression are introduced and machine learning based Survival trees, and Random Survival Forests are studied. In application concordance statistics of Random Survival Forests and Cox regression models were obtained by using both real and simulated data in which different censoring rates were tested. Propher graphical representations were given for the survival curves and variable importance metrics when necessary. As demonstrated by the examples provided in this thesis the traditional approaches are adventageous when assumptions are met and statistical power is high however machine learning based methods work better when censoring is high an assumptions are not satisfied.

Cox regression modelSurvival analysis
Tuğçe Paksoy
Dokuz Eylül University · Institute of Graduate Studies in Science
2021
00
Master'sOpen AccessEN

Lojistik regresyon ve karar ağacı algoritmalarının tahmin edici performanslarının karşılaştırılması: Yaşam memnuniyeti uygulaması

Decision tree algorithms and regression in machine learning create classes of data. Relationships between variables are modeled. Decision trees create classification rules using training data. They also test these rules on test data. Thus, the decision tree determines the success of the algorithm. With the model created in logistic regression, classification is created and classification performance is found. These methods are easy to interpret. They are easily applied to large data sets. They are used in many different fields due to the lack of assumptions. Satisfaction, which is a part of the concept of life satisfaction, is the fulfillment of needs, desires and wishes. Life satisfaction deals with a person's entire life. Life satisfaction is the whole of processes related to individuals' own life patterns and standards. The aim of this study is to compare the performances of logistic regression method and decision tree algorithms (CART, CHAID, QUEST) estimators using life satisfaction data (n = 8430) obtained by the Turkish Statistical Institute (TURKSTAT) for the year 2017. In this study, performance comparisons (accuracy, sensitivity, selectivity, precision, F-score) were made and it was found that the model that best explains the concept of life satisfaction is the QUEST algorithm.

Arzu Yavuz
Dokuz Eylül University · Institute of Graduate Studies in Science
2021
00
Master'sOpen AccessEN

Makine öğrenmesi algoritmaları ile meteorolojik parametreleri kullanarak toprak radon gazının tahmini

Radon is the natural radiation source with the highest dose of exposure among all radiation sources found on earth. It consists of the degradation of natural uranium and radium elements in rocks. Factors that shape the movement of radon include meteorological factors such as the rate of decaying of radon isotopes, the fluids that fill the pores(air, water and other gases), atmospheric pressure, soil and air temperature, wind speed, and wind direction. The aim of this study is to evaluate the effects of some meteorological factors on Radon gas using Supervised Learning Algorithms and to estimate the radon gas values according to these factors. For the study, in addition to the radon levels obtained from the Seferihisar region in hourly periods between 30 October 2006 and 04 June 2007, the measurements for the parameters of Hourly Actual Pressure(hPa), Hourly 50 cm Soil Temperature(°C), Hourly Relative Humidity(%), Hourly Temperature(°C), Hourly Wind Degree(°), Hourly Wind Speed(m/sec) and Wind Direction were obtained from the Republic of Turkey Ministry of Agriculture and Forestry, General Directorate of Meteorology. To analyze the relationship between Radon and meteorological factors affecting radon with Supervised Learning Algorithms, Multiple Linear Regression, k-Nearest Neighbor, Support Vector Machines, Regression Trees, Bagging, Random Forests, XGBoost methods have been used. To test the success of applied methods K-Fold Cross-Validation(K=5) and verification tests were performed. Specification coefficient(R^2) for comparing the performance of algorithms, Mean Squared Error(MSE), Root Mean Squared Error(RMSE), Mean Absolute Error(MAE) values were used. The best result was random forests regression when performance criteria were taken into account. This method was followed by the XGBoost and k-Nearest Neighbors algorithms, which gave very close results.

Çağla Öztürk Zan
Dokuz Eylül University · Institute of Graduate Studies in Science
2021
00
Master'sOpen AccessTR

Basit doğrusal regresyon modelinde model parametrelerinin dayanıklı tahmini

DANIŞMAN: DOÇ. DR. DEMET HAN AYDIN Bu tez çalışmasında, basit doğrusal regresyon modelinde model parametrelerinin En Küçük Kareler (Least Squares–LS) tahmin edicilerinin performansı farklı senaryolar altında incelenmiştir. İlk senaryoda, hata terimlerinin dağılımı olarak normal dağılıma alternatif olabilecek farklı olasılık dağılımları ve çeşitli aykırı değer modelleri dikkate alınmıştır. İkinci senaryoda ise, hata terimlerinin sağa çarpık bir dağılım olan Gumbel dağılımını izlediği varsayılmıştır. Ayrıca, X-yönünde aykırı gözlemlerin bulunduğu durumlarda, LS tahmin edicilerinin performansları, literatürde yaygın olarak kullanılan En Küçük Mutlak Sapma (Least Absolute Deviation-LAD), Ağırlıklı En Küçük Mutlak Sapma (Weighted Least Absolute Deviation-WLAD) ve En Küçük Medyan Kareler (Least Median of Squares-LMS) gibi dayanıklı (robust) tahmin yöntemleri ile karşılaştırılmıştır. Tahmin yöntemlerinin etkinliğini değerlendirmek amacıyla Monte-Carlo simülasyon tekniği kullanılmış; karşılaştırma ölçütleri olarak ise yan (Bias) ve hata kareler ortalaması (Mean Squared Error-MSE) kriterleri esas alınmıştır. Simülasyon çalışmalarından elde edilen bulguları desteklemek amacıyla, literatürde yer alan gerçek veri seti ile analiz gerçekleştirilmiştir. Hem simülasyon hem de gerçek veri analizi sonuçları, hata terimlerinin normal dağılmadığı veya veri setinde aykırı gözlemlerin bulunduğu durumlarda, LS tahmin edicilerinin performansının önemli ölçüde azaldığını ortaya koymuştur. Buna karşın, dayanıklı tahmin yöntemlerinin bu tür veri bozulmalarına karşı daha kararlı ve güvenilir sonuçlar ürettiği belirlenmiştir. Ayrıca, LMS tahmin edicisinin, model parametrelerinin tahmininde varsayım ihlalleri karşısında diğer yöntemlere kıyasla daha yüksek dayanıklılık gösterdiği sonucuna ulaşılmıştır. WLAD tahmin edicisi ise performans açısından LMS yönteminin ardından en başarılı tahmin yöntemi olarak tespit edilmiş olup, LMS'ye alternatif olarak tercih edilebilecek güvenilir bir seçenek durumundadır. Sonuç olarak, veri setinin aykırı değerler içermesi veya hata terimlerinin normal dağılmaması durumlarında, model parametrelerinin tahmininde klasik LS tahmin yöntemi yerine LMS dayanıklı tahmin yönteminin kullanılması önerilmektedir.

Ömer Bayraktar
Sinop University · Institute of Graduate Studies
2025
00
Master'sOpen AccessTR

Deney tasarımında optimal blok yapıları

Deney tasarımı teorisinde bloklama, sistematik gürültüyü azaltmak ve etki tahmininin doğruluğunu artırmak amacıyla yaygın olarak kullanılmaktadır. Tasarımların optimal yolla nasıl bloklanacağı ise, uygulamada büyük önem taşımaktadır.Bu çalışmada, çok etkenli ve 2 düzeyli kesirli çok etkenli tasarımlar hakkında bilgi verilmiş; çözüm ve en az sapma kavramları, tanımlayıcı bağıntı alt grupları, blok ve deneme kelime uzunluğu yapıları tanıtılmıştır.Bloklanmış kesirli ve çok etkenli tasarımların seçiminde kullanılan var olan optimallik ölçütleri araştırılmış; tasarımları tanımlayıcı bağıntı alt grupları ve kelime uzunluğu yapılarını kullanmadan karşılaştıran en az moment sapma ölçütü incelenmiştir.Çalışmanın uygulamasında; Sun, Wu ve Chen'in optimal blok yapıları katalogundaki tasarımlar incelenmiş; optimal blok yapısının bulunması için en iyi yöntem seçilmiş ve nedenleri açıklanmıştır.Anahtar Kelimeler: Çok Etkenli Tasarımlar, Kesirli Çok Etkenli Tasarımlar, Kelime Uzunluğu Yapıları, En Az Sapma, Optimal Blok Yapısı.

Erdinç Kolay
Sinop University · Institute of Graduate Studies in Science
2011
00
Master'sOpen AccessTR

Türkiye'de orta ölçekli bankaların veri zarflama analizi ile etkinlik ölçümü uygulaması

Bir ülkede ekonomik büyümenin ön koşullarından birisi güçlü ve sağlıklı bankacılık sektörünün olmasıdır. Bu nedenle bankaların etkinliği önem kazanmaktadır. Bu etkinlik gerek yatırımcılar, gerek politikacılar, gerekse de banka üst yönetimleri tarafından takip edilmekte, karar aşamasında da dikkate alınmaktadır. Bunun yanında bankacılık sektörü de kendi içinde rekabet içinde bulunmaktadır. Sektörün rekabet koşulları; sunulan hizmet kalitesi, kaynak ihtiyacı ve kar beklentisinin artışı gibi konuları içermektedir. Son yıllarda Türk Bankacılık Sektöründe nde özellikle; banka sayısı, personel sayısı, teknolojik yatırım, hizmet sayısı ve karlılık konusunda ilerleme kaydedilmiştir. Söz konusu ilerlemenin başlıca sebebi, sektörün geçmiş yıllarda yaşamış olduğu krizlerin sektörü tecrübeli hale getirmiş olmasıdır. Bankaların verimlilik kriterlerini esas alan çalışma ilkesi de bu gelişmede önemli bir rol oynamıştır. Ekonomik sistemde önemli bir yer tutan bankaların etkin çalışmaları büyük önem arz etmektedir. Bu çalışmada Türkiye'de bulunan orta ölçekli ve özel sermayeli (yerli –yabancı) on iki adet bankaya VZA ile etkinlik analizi uygulanmıştır. Analiz sonucunda etkin olan bankalar tespit edilmiş, etkin olmayan bankaların etkin hale gelebilmeleri için etkinleşme önerileri sunulmuştur.

Nurettin Gökmen
İstanbul University · Institute of Graduate Studies in Social Sciences
2019
00
Master'sOpen AccessEN

İki değişkenli yaşam verilerinin kopulaya dayalı modellemesi ve analizi

Modelling dependence structure of a bivariate survival data is one of the main issues in biomedical studies. Copulas are key tools to analyze the dependence structures. A bivariate survival function can be expressed as the composition of marginal survival functions and a bivariate copula. Since a survival copula is a great deal of flexibility in modelling bivariate survival data, it provides an effective approach for understanding and modelling the dependent random variables and so the dependence structure. Survival copula deals with a lifetime data and is used for modelling and understanding the distributional structure. In survival studies, the researcher can come across censored survival data. In this study, we consider modelling and analyzing the bivariate survival data in the presence of right censoring using Archimedean copula functions. We use Emura et al. (2010) goodness-of-fit testing procedure for the model selection. Throughout the model selection procedure, we obtain the goodness-of-fit statistics for Gumbel, Frank and Clayton copula models. First, we examine the heart transplant data and model the dependence structure between waiting time for transplant and post-transplant survival time to see the co-movements of these variables. Second, we examine the diabetic retinopathy data and model the dependence between the survival times of the two eyes of the same patient in case of laser photocoagulation treatment. Finally, we use the survival hazard scenario approach to evaluate the probability of exceeding some critical layers. We develop R code to implement the study.

Ece Gorceğiz
Dokuz Eylül University · Institute of Graduate Studies in Science
2020
00
Master'sOpen AccessEN

Dayanikli doğrusal olmayan regresyon yöntemleri

Regression analysis is a statistical method for modelling the relationship between two or more variables. Ordinary least squares regression, which is the most commonly used approach to determine relationship between variables may be misleading in the presence of outliers, or when there is heteroscedasticity and non-normality. Even if the underlying assumptions such as normality and homoscedasticity are hold, the relationship between dependent variable and independent variable(s) may be nonlinear. In such cases, using nonparametric regression methods is more appropriate since the shape of the regression function is not needed to be predefined and there is no significant assumptions as in parametric regression situation. The nonparametric regression estimators are called as smoothers and in this thesis four of them are investigated: Kernel smoothing, locally weighted scatter plot smoothing (LOWESS), the running interval smoother (RIS) and constrained b-spline smoothing (COBS). While the running interval smoother predicts the dependent variable by using different location estimators, COBS predicts by using quantiles and the other methods predict the dependent variable using weighted mean. Robust nonlinear regression methods are also blended to create alternative methods. The smoothers and these alternative methods are compared with a simulation study by using theoretical distributions. Furthermore, the methods are examined graphically to understand how the methods can model the relationship between variables. The predicted values of dependent variable corresponding to new observations are calculated as well. COBS and RIS with NO estimator outperformed the other methods in terms of mean squared error (MSE).

Burak Dilber
Dokuz Eylül University · Institute of Graduate Studies in Science
2019
00
Master'sOpen AccessEN

Ortalamada kaymalar olduğumda parametrik olmayan CUSUM ve EWMA kontrol kartlarının performanslarının karşılaştırılması

Generally, Shewhart control charts, which require normality hypothesis, are used in monitoring process mean. However, Shewhart control charts may not show adequate performance when the distribution of the process in question does not suit the normal distribution and when there are small shifts in the process mean. For this reason, in cases in which the process distribution is not known, it is more beneficial to use nonparametric control charts that do not require any hypothesis about the distribution. In addition, if there are shifts less than 1.5 sigma, which can be defined as small in the process, preferring the CUSUM (cumulative sum) and EWMA (exponentially weighted moving average) charts, developed as alternatives to Shewhart control charts, would yield more accurate results. In this study, the nonparametric CUSUM control chart and the nonparametric EWMA control chart, designed with the change-point model and the Mann-Whitney Statistic, were introduced and the simulation study was conducted using the R statistical programming language. In this simulation study, data from four different distributions were generated and the average run length (ARL) values for both control charts were calculated. As a result of the calculated ARL values, both control charts were compared in terms of performance and it was observed that the CUSUM chart performed better than the EWMA chart for all distributions applied under the determined conditions.

CUSUM control chartsEWMA control chartsStatistical process control
Fatma Kaymakamtorunları Deniz
Dokuz Eylül University · Institute of Graduate Studies in Science
2019
00
Master'sOpen AccessEN

Denetimsiz anomali tespit algoritmaları

Detection of outliers or anomalies in the data is of great importance in data analysis. Different approaches can be used in anomaly detection according to type of the problem. Unsupervised anomaly detection (UAD) approach is the most challengeable part of these approaches. UAD methods aim to detect anomalies without using a labelled training dataset. UAD algorithms can be considered in three main groups: nearest neighbour based, clustering based and statistical based. In this thesis, UAD approaches is examined and an adjustment, that depends on sample size, is proposed for statistical based algorithm, HBOS. In the first part of the application, performance of the most widely used UAD algorithms, that are k-nearest neighbour (k-NN), local outlier factor (LOF), local density cluster-based outlier factor (LDCOF) and histogram-based outlier score (HBOS), are compared. According to the results, HBOS algorithm is found more successful in terms of accuracy rate and runtime. In the second part of the application, effect of the bin-width determination techniques on to the performance of HBOS algorithm is examined. According to the results of the comparison with multivariate data of different characteristics, there is no superiority between the bin-width determination techniques in terms of accuracy.

Beyza Kızılkaya
Dokuz Eylül University · Institute of Graduate Studies in Science
2019
00
DoctorateOpen AccessEN

Risk ayarlı hastane ölüm tahmin modeli: Bir Türk eğitim ve araştırma hastanesinde uygulama örneği

In today's world, health organizations give much importance to quality and patient safety. To this end, conservation of life and prevent excessive deaths are one of the vital objectives for health services in all countries (Whalley, 2010). Although main function of hospitals is to save lives, there is a little attention to hospital mortality (Champbell et al., 2011). In this context; generating reliable mortality ratio then monitoring them are a prerequisite for improvement in care and development in patient safety. This study aimed to demonstrate the applicability of risk adjusted mortality ratio in Turkey. This is the first study conducted in this field in Turkey. To this end, various risk adjusted hospital mortality prediction models were developed by using some popular data mining techniques; logistic re-gression, decision trees, random forests and artificial neural networks. The data from 30182 inpatients of one of the Turkish training and research hospitals with 1155 beds were used. The data collected from inpatients whose discharge period was January to November in 2014. At the end, the performance of these methods were compared.

Fatma Güntürkün
Dokuz Eylül University · Institute of Graduate Studies in Science
2019
00
Master'sOpen AccessEN

Yapısal kırılmalar olduğunda uyarlanmış ve basit üstel düzeltme yöntemlerinin karşılaştırılması

The essential aim of the time series modelling is applied for the forecasting as well as the examination of correlation. One of the most widely used methods in the literature is exponential smoothing (ES) methods. It is a method preferred by many researchers because of its easy application, calculation efficiency, high accuracy and automatic prediction. Such as policy changes, financial crises, natural disasters in the data production processes of the series, permanent structural changes can change affect model parameters as well as analysis results. Having no consideration of such affects, leads parameter estimator to be biased, tests tend to be useless in terms of power and incorrect modelling arise. The main purpose of this study is to compare the predictive performances of the newly developed Modified Exponential Smoothing (MSES) (2016) methods with the simple exponential smoothing (SES) when there are structural breaks in the series. Determining the initial value and misspecification in the selection of the optimum smoothing parameter, as a disadvantage, adversely affect the estimation results. The MSES method gives more weight to the current observations on the series, so that the predictions that are calculated give better performance than the classical method. The MSES method against structural break has not been examined yet. In this study, received from the Central Bank of the Republic of Turkey and traded on the Istanbul Gold Exchange "weighted average price of gold (TL/kg)" data are used. This data set with different break points compares the forecast performance of MSES and SES methods.

İrem Efe
Dokuz Eylül University · Institute of Graduate Studies in Science
2019
00
DoctorateOpen AccessEN

Archimedean kopula modelleri için yeni bir uyum iyiliği yaklaşımı ve güç analizi

In this thesis, a new class of bivariate multi-parameter Archimedean copula based on Kendall distribution using Bernstein-Bezier polynomials is introduced. This new class copula has flexible dependence properties depending on the polynomial degree and the control points. Some dependence characteristics such as Kendall's tau, upper tail and lower tail dependence of the this new Archimedean copula class are derived. The simulation procedure based on these desired dependence characteristics is presented. Also, we propose an estimation method for the Archimedean family of copula in a nonparametric setting. Bernstein polynomials and Bezier curve approaches are used to estimate the Kendall distribution function of the Archimedean copula. Also, a new goodness-of-fit test based on Cramer-Von Mises type statistic is constructed using new estimation methods of Kendall distribution function. A Monte Carlo study is performed to measure the performance of the proposed tests. The simulation results show that both the Bernstein polynomial and the B\'ezier curve based tests have better performances than the classical one since they have flexible form according to its order m.

Selim Orhun Susam
Dokuz Eylül University · Institute of Graduate Studies in Science
2019
00
DoctorateOpen AccessEN

Regresyon ve dağılım tahmininde geometrik modelleme

The contribution of this thesis to inferential statistics by using geometric modeling methods is comprised of three parts. First, we propose a new nonparametric regression model that uses the rational Bézier curves. The main advantages of the proposed regression model are the flexibility over parametric models and having no restriction on the assumptions in regression function estimation. Moreover, rational Bézier curves provide many conveniences such as more control over the shape of a curve and projective invariance. We provide numerical examples to demonstrate superiority of the proposed model and verify it. Secondly, we introduce a new distribution function estimation method using the rational Bernstein polynomials (RBP). The performance of the proposed new method is compared with Bernstein polynomials and empirical distribution function methods using simulation. The new method guarantees monotone nondecreasing function by applying linear constraints on the coefficients of the rational Bernstein basis functions and smooth the empirical distribution function. Furthermore, as a special case, it reduces to Bernstein polynomial estimator method. Some theoretical properties of the new estimator are investigated. Simulation study shows that the proposed estimator is preferable to the Bernstein polynomials and empirical distribution function estimator methods. Thirdly, we introduce a unimodal density estimator based on the Schoenberg's spline operator and thus generalize that of Bernstein polynomials and the beta density. The advantage of this method is the local property. That is, refining the knots while keeping the degree fixed of B-splines yields better estimates. We also give a numerical example to verify our results.

Geometric modellingNonparametric regression
Mahmut Sami Erdoğan
Dokuz Eylül University · Institute of Graduate Studies in Science
2019
00
DoctorateOpen AccessEN

İki bağımsız örneklem için dayanıklı yaklaşımların analizi

Comparing two independent groups has been the subject of many classic studies in applied statistics. Although the most common approach for comparing two independent groups is on the basis of just one reference point, it may provide misleading results and cause missing the important differences among subpopulations of the groups. Determination of the differences in tails of the groups, namely quantiles, can also be of particular concern. This thesis aims to compare two independent groups from a different point of view by utilizing quantiles. Therefore, firstly attention is focused on quantile estimation theory and the newly proposed NO quantile estimator is introduced. NO quantile estimator is computationally practical, has desirable asymptotic properties and is more efficient than some quantile estimators that are commonly used. Following this, NO quantile estimator is used in conjunction with a percentile bootstrap to propose a method that compares two independent groups through multiple quantiles simultaneously. By conducting a series of simulation studies and using some real datasets, the performances of NO quantile estimator and the proposed method for comparing two independent groups are investigated. The proposed method could control the actual type I error rates in almost all cases that are investigated. This method is more preferable than the conventional approaches that considers a single measure of location since it makes a more sensitive comparison by using more than one reference points.

Gözde Navruz
Dokuz Eylül University · Institute of Graduate Studies in Science
2019
00
Master'sOpen AccessEN

Sıralı küme örneklemesinde sıralama hata modelleri, maliyet ve en uygun küme büyüklüğü

Ranked Set Sampling (RSS) is a sampling method commonly used in recent years. RSS is developed as an alternative to Simple Random Sampling (SRS) in order to estimate population parameters more efficiently where the measurement of sampling units is difficult or costly but the units are easier to rank. There are several factors that make this method useful especially for studies in medicine, agriculture, forestry and ecology. The most important of these factors are the set size and the relative costs of some operations such as sampling, measurement and ranking. Ranking of the units in the set is made on the basis of the visual judgment of the researcher or a concomitant variable which has a strong correlation with the variable of interest. These ranking methods are defined as ranking error models. In this thesis, the widely used cost and ranking error models in RSS literature are investigated. Also, it is aimed to explore the effect of ranking error models on the mean estimator based on RSS and some of its modified methods for different distribution, set and cycle size in infinite population. Besides, it is aimed to examine whether RSS is cost effiective with respect to SRS in terms of mean squared error of the mean estimator considering ranking error models and the N-KPST cost model in infinite population and if so, to determine the optimal set size for RSS. Monte Carlo simulation studies are conducted for these purposes. Additionally, the study is supported by real life data.

Sami Akdeniz
Dokuz Eylül University · Institute of Graduate Studies in Science
2019
00
Master'sOpen AccessEN

Modifiye sıralı küme örneklemesi yöntemlerinde regresyon kestiricilerinin etkinliklerinin incelenmesi

Ranked Set Sampling (RSS) has been a popular sampling method in recent years. RSS is used when visual ranking can be done easily while the variable of interest is difficult or expensive to measure. In the last years, different modifications of the RSS which are Pair RSS, Extreme RSS, Median RSS, Double RSS, LRSS and Truncation Based RSS methods have been used in wide applications and suggested by many researchers. The aim of this study is to estimate the regression estimators and to compare relative efficiencies of mean square errors of the regression models, population mean and regression coefficients for Simple Random Sampling (SRS), RSS and the different modified RSS methods. The Monte Carlo simulation study is performed via R Project with 10,000 repetitions. The performances of the estimators are compared based on bias, mean square error and relative efficiency for different levels of correlation coefficient, set and cycle sizes under Bivariate Normal Distribution (BVN) with different parameters for normal and outlier cases. The results indicate that the regression estimators under the modified RSS methods performs better than the regression estimators under SRS.

Ranked set sampling
Eda Davaslıoğlu
Dokuz Eylül University · Institute of Graduate Studies in Science
2019
00
Master'sOpen AccessEN

Spor veri madenciliği tekniklerinin incelenmesi

Sports data mining is the usage of data collected on players, teams, and games for performance evaluation, player selection, score-outcome prediction, and strategy development with data mining tools and techniques. Special performance measures developed for each sports branch have an important role in sports data mining and sports statistics. Performance measures calculated for the team sports and the players can be used to predict the expectation of winning. The Pythagorean Expectation developed for this purpose was first used in baseball games. The Pythagorean Expectation, which attracts the attention of other sports branches, is adapted to other team sports with two possible outcomes such as basketball. However, Pythagorean Expectation studies are limited for sports which have three possible outcomes such as football. In this thesis, it is aimed to investigate sports data mining studies related to various sports branches. In this context, performance measurements and sports data mining studies related to many sports branches are examined and a detailed literature review is performed. Also, it is aimed to suggest a new approach to computation of Pythagorean Expectation for football. In the application section, for the teams competing in fifteen different European football leagues, 2017/2018 season-end rankings and points are predicted using the proposed approach. The data of the last five seasons of the selected European football leagues is used as training data set. All calculations are performed in R.

Sezer Baysal
Dokuz Eylül University · Institute of Graduate Studies in Science
2019
00
Master'sOpen AccessEN

Üssel düzleştirme ve ATA metodu ile finansal veri analizi

Time series occurs by collecting the data in a particular category in a given time period. Accurate analysis of financial data, which is a sort of time series, has a great importance for financial institutions to make predictions for the future. Exponential smoothing method is one of the most used method in time series analysis. Exponential smoothing methods have been used widely for many years due to their simplicity and success in prediction results. The success of the method has proved many times in the famous M-competitions. However, the selection of initial value and smoothing constant according to subjective choices for exponential smoothing method adversely affect the accuracy of this method. The ATA method, which is a new method developed as an alternative to the exponential smoothing method, eliminates these disadvantages of the exponential smoothing method. In this study, the M4 results of exponential smoothing method and ATA method will be compared, especially their performance in financial data will be evaluated.

Financial datasTime seriesExponential smoothing
Selma Şalk
Dokuz Eylül University · Institute of Graduate Studies in Science
2019
00
Master'sOpen AccessTR

Türkiye ve AB ülkelerinin COVID-19 'a karşı etkinliğinin veri zarflama analizi ile karşılaştırılması

2020 yılının başında, tüm halk sağlığı camiasının onlarca yıldır korktuğu senaryo gerçek olmuş ve hayvanlardan insanlara bulaşarak hastalığa neden olan korona virüslerin yeni bir tipi Çin' den tüm dünyaya hızlıca yayılmaya başlamıştır. Dünya Sağlık Örgütü tarafınca "Covid-19" olarak adlandırılan virüs 24 Ocak 2020 tarihinde Avrupa kıtasına ulaşmış ülkemizde ise ilk vaka 11 Mart 2020 tarihinde görülmüştür. Bu çalışmanın amacı Aralık 2019`da Çin`in Wuhan şehrinde ortaya çıkan ve kısa sürede dünyayı saran COVID-19 salgını ile mücadelede Avrupa Birliği ülkelerinin ve Türkiye'nin etkinlik düzeyinin Veri Zarflama Analizi ile araştırılmasıdır. Sağlık sistemlerinin performans ölçümünde sıklıkla kullanılan Veri Zarflama Analizi, pek çok girdi ve çıktının kullanılabildiği parametrik olmayan etkinlik ölçüm yöntemidir. Analizde girdi değişkenleri olarak; onbin kişi başına doktor, hemşire, hastane ve yoğun bakım yatak sayıları ile sağlık harcamalarının Gayri Safi Yurtiçi Hâsıla (GSYH) içindeki oranı, çıktı olarak; ülkelerin 10 Haziran 2021 tarihine ait milyon kişi başına Covid-19 test sayısı, vaka sayısı ve ölüm sayısı ele alınmıştır. Etkinlik analizinde sabit ölçekli Charnes, Cooper, Rhodes (CCR) ve değişken ölçekli Banker, Charnes, Cooper (BCC) yöntemleri kullanılmış, etkin olan ülkelerin etkinlik skorları belirlenmiş ve etkin olmayan ülkeler için potansiyel iyileştirme önerilerinin geliştirilmesi amaçlanmıştır.Daha sonra doldurulacaktır.

Gonca Çetinkaya Eroğlu
Gazi University · Institute of Graduate Studies in Science
2021
00
DoctorateOpen AccessTR

Dengeli IXJ tasarımları için senkronize permütasyon testlerinde yeni bir yaklaşım

IxJ faktöriyel tasarımlar iki faktör ve düzeylerinin çeşitli kombinasyonlarının yanıt değişkeni üzerindeki etkilerinin araştırıldığı deneydir. Başta mühendislik alanlarında olmak üzere, sağlık bilimleri, zootekni ve tarım araştırmalarında yaygın olarak kullanılan bu tasarımın en önemli avantajı, incelenen özelliği etkileyebileceği düşünülen faktörlerin hem ana etkilerini hem de etkileşimlerini en küçük hatayla tahmin edebilmesidir. İki yönlü ANOVA (Varyans Analizi) metodolojisi birçok araştırma alanlarında kullanılan deney tasarım teorisindeki önemli modellerden biridir . Yaygın olarak faktöriyel tasarımlarda ana etki ve etkileşim etkisi İki yönlü ANOVA ile test edilir. İki yönlü ANOVA için gerekli varsayımlar sağlanmadığı durumlarda varsayımlardan etkilenmeyen permütasyon testleri parametrik olmayan bir yöntem olarak uygulanabilmektedir. Bu yöntemlerden biri olan Senkronize permütasyon testleri faktöriyel tasarımlarında iki ana etki ve etkileşim etkisi için birbiri ile ilişkisiz üç ayrı tam test üretmektedir. Bu çalışmada, senkronize permütasyon testlerinin ana faktörlerin etkisi ve etkileşim etkisi için test istatistiklerinin sahip olduğu karesel form mutlak değer ile değiştirilrilerek yenilenmiştir. IXJ dengeli faktöriyel tasarımlar için hata terimlerinin normal dağılıma sahip olmadığı durumlarda karesel form içeren test istatistiklerine sahip Senkronize permütasyon testi, mutlak değer fonksiyonu içeren test istatistiklerine sahip Senkronize permütasyon testi ve iki yönlü ANOVA I.tip hata ve testin gücü bakımından karşılaştırılmıştır. Senkronize permütasyon testinde permütasyon prosedürleri incelenmiş olup bu prosedürlerden kısıtlı senkronize permütasyonlara (CSPs ) alternatif olarak tekrar sayısı n>7 olduğu durumlar için rasgele seçimli permütasyon prosedürü önerilmiş ve bu yöntem senkronize permütasyon testi altında I.tip hata ve testin gücü bakımından kısıtlı senkronize permütasyonlar ile karşılaştırılmıştır.

Dilşad Yıldız Kaçar
Gazi University · Institute of Graduate Studies in Science
2021
00
Master'sOpen AccessTR

Veri zarflama analizi ve makine öğrenmesi metodlarından oluşan hibrit bir yöntem

Çin'in Wuhan şehrinde 2019 yılının sonunda ortaya çıkan ve hava yolu ile bulaşan bir hastalık olan COVID – 19 dünya genelinde hızla yayılmıştır. Salgının yayılma miktarı ve hızından dolayı 11 Mart 2020 tarihinde Dünya Sağlık Örgütü tarafından pandemi olarak ilan edilmiştir. Ülkeler hem bulaşıcılığı kontrol altına almak hem de hastalığa yakalanan kişilere uygulanan tıbbi tedavilerin etkinliğini arttırmak amacıyla çeşitli stratejiler uygulamışlardır. Bu stratejilerin sonuçlarının değerlendirilmesi ise birçok çalışmanın ana konusu olmuştur. Veri Zarflama Analizi karar verme birimlerinin performanslarının karşılaştırılmasında güçlü bir teknik olduğundan bu çalışmalarda yaygın olarak kullanılmıştır. Ancak bu çalışmalarda karar verme birimlerinin homojenliğini sağlamak amacıyla değişkenlere nüfusa göre ölçeklendirme yapılması sonuçların tutarlılığını tehlikeye atmaktadır. Oluşan tutarsızlık, her ülkenin nüfus yoğunluğunun farklı olmasından kaynaklanmaktadır. Bu çalışmanın amacı nüfus yoğunluğu artışının bulaş ve ölüm oranına etkisini doğru olarak konumlandırarak Veri Zarflama Analizi'ndeki heterojenlik sorununun ortadan kaldırılması ve ülkelerin COVID – 19 salgınında bulaşıcılığı kontrol altına alma amacıyla uyguladıkları stratejilerde ve hastalığa yakalanan kişilere uygulanan tıbbi tedavilerin etkinlik performanslarının doğru şekilde hesaplanmasıdır. Çalışmada yer alan 85 ülkeye homojenliğini sağlamak için öncelikle nüfusa göre ölçeklendirme yapmak yerine ülkelerin nüfus yoğunlukları kullanılarak kümeleme analizi yapılmıştır. Kümeleme analizi sonucunda elde edilen homojen kümelerin her birine ülkelerin bulaşıcılık ve tıbbi tedavi performanslarının ölçülmesini amaçlayan iki farklı senaryodan oluşan Seri Hiyerarşik Veri Zarflama Analizi uygulanmıştır. Kümeleme yapılmadan önce ve yapıldıktan sonra uygulanan Seri Hiyerarşik Veri Zarflama Analizleri sonucunda elde edilen etkinlik skorlarının birbirlerinden farklı olduğu tespit edilmiştir.

Canberk Arslan
Bolu Abant Izzet Baysal University · Institute of Graduate Studies in Science
2021
00
DoctorateOpen AccessEN

Gönüllü web anketlerinde yanlılığı azaltmak için ağırlıklı düzeltme tekniklerinin karşılaştırılması

Web surveys have become one of the most widely utilized and popular survey methods in recent years. A web survey is performed over the World Wide Web by inviting individuals to complete the questionnaire by themselves. Internet usage has increased among inhabitants of developed countries as well as developing countries. Since the advent of the smartphone, web surveys have become even more viable and reasonable data collection method. Although, in the last years, internet-based data had been collected only for marketing researches. Data collection on the internet is faster than other methods such as paper-and-pencil, computer-assisted telephone interviews and personal interviews, because it is simple, cheap and provides quick access to the desired large group of respondents. However, in web surveys, bias may arise mainly due to limited coverage and self-selection. This study appraises characteristics of web surveys, their importance and trend, as well as problems to identify their bias. Some weighting adjustment techniques for reducing the bias are illustrated, and non-probability estimates from volunteer panel web surveys are compared to random sample estimates. Post-stratification weighting, generalized regression modeling, raking ratio estimation and propensity score adjustment techniques are used for reducing these biases. In the application of this study, population estimates are compared to those from a volunteer panel web survey. The data are from the survey "Using Social Networking Sites in the Education of Students of Open Education System, Anadolu University". A random sample based on simulation is created by a stratified probability sampling design, and estimates based on this sample are compared with estimates from a non-probability-based volunteer panel web sample. It is shown that the weighting adjustment has reduced the bias substantially and the random sample estimates provide better results than those from the volunteer panel web survey. However, when it is necessary to use volunteer panel web surveys, it is recommended to adjust at least one of the weighting adjustment techniques described in this study. Keywords: Web surveys; Volunteer panel web surveys; Probability and non-probability sampling design; Random sample; Bias; Weighting adjustment techniques.

Md Musa Khan
Anadolu University · Institute of Graduate Studies in Science
2018
00
Master'sOpen AccessTR

Bulanık kümeleme analizinde bulanık kümeleme algoritmalarının karşılaştırılması

Kümeleme analizi, örneklerin özelliklerine dayanarak veri noktalarını gruplamak için kullanılan denetimsiz bir öğrenme tekniğidir. Son dönemlerde bu teknikler özellikle biyoinformatik, biyoistatistik, genetik gibi alanlarda kullanımı artmıştır. Kümeleme, keskin ve bulanık olarak iki modda gerçekleştirilebilir. Keskin kümeleme yöntemlerinde, her birim kesinlikle bir kümeye atanmalıdır. Bulanık kümeleme algoritmaları ise, her bir nesnenin tek bir kümeye atanma kısıtını ortadan kaldırarak, belirli üyelik dereceleriyle tüm kümelere ait olmasına olanak kılar. Bu tezin amacı; bulanık kümeleme sürecinin araştırmacılara detaylı olarak aktarılmasını sağlamaktır. Ayrıca biyoinformatik, genetik gibi alanlarda bulanık küme algoritmalarının kümeleme tekniklerine alternatif bir yöntem olarak kullanılabileceğini göstermektir. Bu bağlamda, ilk olarak bu teknikler için önemli parametre olan optimal küme sayısının bulunması için bir uygulama gerçekleştirilmiştir. Bu uygulamada yaygın kullanılan genetik veri setinde hem geçerlilik indeksleri hem de dirsek yöntemi kullanılarak kapsamlı karşılaştırmalar yapılmıştır. Daha sonra, gen ekspresyon modellerine bulanık ve klasik kümeleme algoritmaları uygulanmış ve algoritmaların performansları karşılaştırılmıştır. Son olarak, bulanık kümelemede bir dezavantaj olan aykırı değer sorununu üstesinden gelmek için geliştirilmiş bulanık kümeleme algoritmaları karşılaştırılmıştır. Uygulama sonuçları basit bir şekilde analiz edilmiştir.

Aslı Kaya
Anadolu University · Institute of Graduate Studies in Science
2018
00
Master'sOpen AccessEN

Veri madenciliği ve Anadolu Üniversitesi açıköğretim sisteminde bir uygulama

In Data mining, the discovery of frequent sets of items that occur together in a dataset is a fundamental task, particularly in transaction datasets. This thesis focus on closed frequent itemsets, a compressed representation of frequent patterns. Since mining all frequent itemsets generates duplicates and subsets, and with dense dataset, it gives a huge number of generated frequent itemsets which may be not understandable, then to avoid generate duplicates and subsets, and to get useful and compressed information about data under study, closed frequent itemset technique was adopted in this study. In this thesis, the CHARM algorithm to discover closed frequent itemsets is proposed to identify frequent itemsets for Anadolu University Open Education System data to find out the collection of books taken together. The reason of choosing CHARM algorithm in this study is that it employs a vertical layout which facilitates the process of computing support count, moreover, it employs a novel way to traverse the search space. The solution is implemented in R software environment for statistical computing and graphics. The result obtained shows that the encoded CHARM algorithm can get an efficient performance. Moreover, the result obtained could be useful to help decision makers to package the books in packets to achieve the aim of the study.

Bılal Al-rubaıee
Anadolu University · Institute of Graduate Studies in Science
2018
00
DoctorateOpen AccessTR

Kuzey anadolu fay hattı üzerinde gerçekleşen depremlerin mekânsal ve mekân- zamansal olarak incelenmesi

Depremler gerçekleşme sebepleri açısından incelendiğinde birçok faktör tarafından etkilenen karmaşık olaylardır. Gerek ilgilenilen bölgede bulunan fay hattı, gerekse depremlerin öncü şok, ana şok ve artçı şok gibi kendi aralarında yer alan ilişkiler deprem oluşumlarındaki karmaşıklığa sebep olabilmektedir. Deprem riskinin tahmin edilmesi veya deprem yoğunluklarının modellenmesi ileriye dönük yaşanabilecek maddi ve manevi kayıpların en aza indirgenmesi açısından hayati bir önem taşımaktadır. Bu çalışmada Kuzey Anadolu Fay Hattını kapsayan dikdörtgensel bir bölge çalışma alanı olarak seçilmiş ve bu bölgedeki deprem örüntülerinin analizi amaçlanmıştır. Öncelikle, deprem kataloğunda belirli bir zaman aralığında ve çalışma aralığı içerisinde yer alan büyük depremler (M>5) açıklayıcı veri analizi yardımıyla ilgili öznitelikleri açısından görselleştirilmiştir. Ayrıca mekânsal ve mekân-zamansal süreçler yardımı ile mekânsal ve mekân-zamansal örüntü türleri tespit edilmiş ve ilgili çalışma alanı ve ele alınan zaman aralığı içerisinde mekânsal ve mekân-zamansal deprem benzetimleri gerçekleştirilmiştir. Son olarak ise deprem yoğunluklarını arka plan ve tetiklenen olay yoğunlukları olmak üzere iki açıdan inceleyen ve literatürde oldukça sık kullanılan mekân-zamansal epidemik-tip şok sonrası modeli ile ilgili depremler modellenmeye çalışılmıştır. Bu yoğunlukların bölgelere göre farklılık gösterdiği tespit edilmiştir. Depremlerin mekânsal olarak kümelenmiş bir örüntü, mekân-zamansal olarak da düzenli bir örüntüye sahip olduğu sonucuna ulaşılmıştır.

Cenk İçöz
Anadolu University · Institute of Graduate Studies in Science
2018
00
DoctorateOpen AccessEN

Longitudinal veri analizinde eksik gözlem, aykırı değer ve modelleme üzerine çalışma

Longitudinal data consists in gathering several observations of the same subjects intermittently over time. Attrition, outliers and complexity of modeling are common issues in longitudinal data. Therefore, this dissertation outlines those issues and proposes different approaches to overcome them by following three main pillars. First pillar emphasises the prominence of missingness mechanisms and suggests a novel algorithm to treat missing data via Multilayer Perceptron (MLP) with comparison to the ad hoc methods and Expectation Maximum (EM) algorithm. Second pillar consists in presenting outliers as a friendly subject in statistical data not a misleading dilemma, via proposing two novel algorithms using wavelet decomposition within subjects and across subjects along with applying the winsorisation approach within subjects. Last pillar concentrates on modeling via constructing a semiparametric model that combines parametric and nonparametric features. For the nonparametric part of the model, smoothing approaches are required. This research proposes wavelet analysis to smooth data. To examine the efficiency of the proposed algorithms, a real longitudinal dataset and a generated one, are utilized. The results revealed that wavelet decomposition has an impressive capacity as a smoothing approach and as a microscope figuring out the outliers and handling them without losing the originality of the data features. Also, the novel algorithm related to missing data imputation via the output predictions of MLP showed valuable results better than the ad hoc imputation methods and with very slight difference from the EM algorithm. Keywords: Longitudinal data, Semiparametric model, Missing data, Missingness mechanisms, Outliers, Wavelet analysis, Neural network

OutliersWavelet analysisMissing data+3
Maroua Ben Ghoul
Anadolu University · Institute of Graduate Studies
2019
00
Master'sOpen AccessTR

Orta Karadeniz bölgesinde rüzgar hızının panel verili regresyon modeli yardımıyla tahmin edilmesi

Panel veri analizi özellikle ekonomik ilişkilerin tahmininde kullanılan oldukça yaygın bir yöntemdir. Panel veri analizinde kullanılan modeller sabit etkili model ve rastgele etkili model olarak iki başlık altında incelenebilir. Bu modeller birime, zamana, birime ve zamana göre ayrı ayrı incelenebilir. Böylelikle birimler arası farklılıklar, zaman dönemlerine ve birime ve zamana göre farklılıklar göz önünde bulundurularak tahmin yapılır.Bu çalışmada panel veri modelleri kullanılarak 2009-2010 yılları arasındaki rüzgar hızı tahmin edilmeye çalışılmıştır. Bu amaçla rüzgar hızı açıklanan değişken, ortalama maksimum sıcaklık, ortalama minimum nem, ortalama nem açıklayıcı değişkenler olarak alınmıştır. Uygulamada sabit etkili model ve rastgele etkili model karşılaştırılarak, kullanılan veriler için en uygun modelin bulunması amaçlanmıştır. Sonuç olarak uygulamada kullanılan veriler için sabit etkili modelin daha uygun olduğu görülmüştür.

Panel veri modelleri
Zeynep Atlı
Sinop University · Institute of Graduate Studies in Science
2012
00
DoctorateOpen AccessEN

Biyolojik ağların inferansında grafik modeller

In recent years, particularly, on the studies about the complex system's diseases, better understanding the biological systems and observing how the system's behaviors, which are affected by the treatment or similar conditions, accelerate with the help of the explanation of these systems via the mathematical modeling. Gaussian Graphical Models (GGM) is a model that describes the relationship between the system's elements via the regression and represents the states of the system via the multivariate Gaussian (normal) distribution. This distribution also explains the structure of biological systems by means of its "conditional independence" feature. Therefore, in the inverse of the covariance matrix of the multivariate normal distribution, the "zero" value implies no functional interaction, and the "non-zero" value stands for the interaction between the proteins in the estimate of the system's structure. In this study, as the novelty, we use the Copula Gaussian Graphical Models (CGGM) in modeling the steady-state activation of the biological networks and make the inference of the model parameters under the Bayesian setting. We suggest the reversible jump Markov chain Monte Carlo (RJMCMC) algorithm to estimate the plausible interactions (conditional dependence) between the systems' elements which are proteins or genes. Several data sets are used to illustrate the out-performance of the proposed RJMCMC in comparison with most of its alternatives. Also, we used some semi- Bayesian RJMCMC method to estimate the autoregressive coefficient matrix where GGM repeated through time. We improved the model by full-Bayesian approach and followingly, by a tuning parameter to increase the accuracy of the estimated matrices. Some simulated data sets are used to show the accuracy of the different proposed methods. Finally, we suggested a method to discover the relationships between variables through copula which is more flexible and it is more appropriate for the nonsymmetric or tail dependent cases. We applied the suggested ways in four real data set and we saw that copula can discover the joint density structure in addition to the available relationships in terms of the shape of the joint distribution to see whether it is symmetric or non-symmetric or even tail dependent or not.

Hajar Farnoudkıa
Middle East Technical University · Institute of Graduate Studies in Science
2020
00
Master'sOpen AccessEN

Zaman serilerinin tahminlenmesinde klasik ve makine öğrenmesi yaklaşımlarına yönelik karşılaştırmalı bir çalışma: Türkiye'nin ihracatı üzerine deneysel bir analiz

Exports has become one of the main economic indicator for countries. Accordingly, an accurate forecasting for exports is an important step for decision making and finding the most appropriate forecasting model constitutes the main subject of many studies. By taking the popularity and success of the machine learning (ML) methods on time series forecasting tasks into consideration, they are utilized also in this study to observe their predictive performances on Turkish exports. In this respect, Long Short Term Memory (LSTM), Support Vector Machines (SVM) and Random Forest (RF) are applied and the results are compared with the most commonly used classical time series models such as Autoregressive Integrated Moving Average (ARIMA) and Exponential Smoothing (ETS) models. The analysis is conducted on Turkish monthly exports data taken from Turkish Statistical Institute (TURKSTAT) within the time interval of January 1997 – September 2019 and the main steps of the analysis are anomaly detection and cleaning, data preprocessing, model development, hyperparameter tuning and model selection and model comparison. The main findings can be summarized as follows; the anomaly detection and cleaning process improves the forecasting ability of the models, ETS is the best forecasting model and SVM model is the most promising among the ML models and the most competitive with the leading one. Besides, ARIMA has the poorest generalization ability among the others.

Eda Günel
Middle East Technical University · Institute of Graduate Studies in Science
2020
00