Theses supervised by Doç. Dr. Hüseyin Polat
11 theses · Anadolu University, Aksaray University, Gazi University
On the robustness of privacy-preserving collaborative filtering schemes
Privacy-preserving collaborative filtering has been receiving increasing attention. There are various algorithms providing accurate recommendations while preserving privacy. Like collaborative filtering algorithms, privacy-preserving collaborative filtering methods might be subjected to shilling attacks. Such attacks are employed by malicious users to increase/decrease the popularity of some target items. They might affect the overall performance of recommendation systems. Therefore, it is imperative to design such attacks with privacy concerns, determine how robust the privacy-preserving collaborative filtering schemes are, how to find out fake profiles, and analyze them. In this dissertation, designing shilling attacks with privacy concerns is studied. Also, robustness analysis of various privacy-preserving collaborative filtering schemes (memory-based, model-based, and hybrid methods) is performed. Determining fake or shilling profiles from perturbed databases is scrutinized. Besides employing the modified existing detection methods, a new shilling attack detection algorithm is proposed. Real data-based experiments are conducted for assessing the overall performance. Empirical outcomes show that designing effective shilling attacks with privacy concerns is possible. Also, existing detection methods can be effectively used to determine fake profiles from masked data. In addition, the novel detection method is successful on filtering out shilling profiles. Compared to memory-based and hybrid schemes, privacy-preserving model-based recommendation algorithms are very robust against shilling attacks.
Privacy-preserving geostatistics
Geo-statistics deals with spatial data and tries to find out relationship between locations and measured data. Methods used in geo-statistics interpolations rely on the principle that things are closer to each other more alike than the things are farther apart. Inverse distance weighting and kriging are most well-known and applied methods in geo-statistics. It is important to perform such methods without violating data confidentiality due to privacy reasons. Also, their accuracy depends on the total number of sample points. If there are insufficient sample points due to financial or privacy reasons, accuracy of the predictions produced by these methods may become unconvincing. There are cases in which institutions obtain measurements for the same or neighbor region. To create more accurate models, they may want to collaborate. However, they do not want to share their private data. In this thesis, privacy-preserving methods are proposed to provide inverse distance weighting- or kriging-based predictions for different data partitioning schemas including central server-based case. The proposed solutions are analyzed with respect to privacy, performance, and accuracy. Different sets of experiments are conducted using real data sets to analyze the proposed methods. Empirical outcomes show that the methods are able to provide accurate predictions while preserving privacy.
Privacy-preserving distributed collaborative filtering
In order to provide accurate and dependable recommendations, online vendors need to have adequate data; however, due to the nature of online shopping and increasing amount of e-commerce sites, data collected for collaborative filtering purposes might be distributed among various companies, even competing ones. Those online vendors holding distributed data might want to offer predictions based on integrated data collaboratively. However, concerns regarding protecting private data, financial fears due to revealing valuable assets, and legal regulations imposed by various organizations prevent them from alliance.In this dissertation, various solutions are proposed to enable online vendors? collaboration for estimating recommendations on vertically or horizontally distributed data while preserving their confidentiality. The proposed solutions mainly employ randomized and cryptographic techniques for protecting privacy. To improve online performance, which may become worse due to collaboration, preprocessing methods such as clustering, dimensionality reduction, and trust are utilized. The recommended methods are analyzed in terms of privacy. Also, superfluous loads caused by privacy concerns are examined. Finally, real data-based trials are performed for evaluating the proposed schemes in terms of the quality of predictions. The analyses and experimental outcomes demonstrate that the methods preserve confidentiality, cause insignificant overheads, and offer accurate recommendations.
Privacy-preserving collaborative filtering on arbitrarily partitioned data
Data collected for collaborative filtering purposes might be arbitrarily partitioned between two parties, even rival companies. Online vendors might have insufficient user ratings. Scarce data then might cause offering inaccurate and unreliable recommendations. In order to supply trustworthy and dependable predictions, one solution for such companies might be cooperation on partitioned user preference data. However, it is still a challenge to convince e-commerce sites cooperate on partitioned data so that they can provide richer collaborative filtering services, due to privacy concerns. Unless confidentiality is protected, such companies are expected to face with serious legal and financial deadlocks in managerial operations.This study aims to scrutinize how to estimate predictions based on arbitrarily partitioned data configurations between two e-commerce companies without deeply jeopardizing their privacy. Privacy-preserving schemes are proposed to offer numerical or binary recommendations using item-based, trust-based, and naïve Bayesian classifier-based prediction algorithms on arbitrarily partitioned data. Along the study, how two parties ended up with cross partitioned data can provide CF services using hybrid CF algorithm is also investigated. It is shown that each proposed method does not intensely violate data owners? confidentiality. The proposed schemes are also investigated in terms of supplementary computation, communication, and storage overheads. Experimental trials are conducted using real data sets to show how the quality of the predictions improves due to collaboration and privacy measures affect accuracy. All appraisements demonstrate that the proposed solutions are preferable for estimating higher quality predictions efficiently on partitioned data while preserving data holders? privacy.
Effects of binary similarity measures on collaborative filtering
With increasing popularity of the Internet, shopping over the Internet through several online vendors is also receiving increasing attention. Customers want to purchase the appropriate products. In other words, they try to select those products that they might like. In order to help their customers, many online companies utilize collaborative filtering systems. Such systems provide two services, namely prediction and top-N recommendations. Quality of these two services mainly depends on similarity measures that collaborative filtering algorithms use in order to determine the most similar entities. Data collected for collaborative filtering purposes might include either numeric or binary ratings. Several studies have been conducted to compare different similarity measures proposed for numeric data. Although there are various binary ratings-based similarity metrics, their effects on accuracy and performance in collaborative filtering systems have not been deeply studied.In this thesis, we investigate seven binary ratings-based similarity metrics in terms of both accuracy and online performance while providing predictions for single items and top-N lists. Although there are more than seven measures, we consider the most widely used ones in various data mining applications. To compare them in terms of correctness and efficiency, we perform several experiments based on two well-known real data sets. We produce both predictions and top-N lists while using different similarity metrics, where we propose to modify prediction and top-N recommendation algorithms in such a way so that the most similar users? data are involved in collaborative filtering process. We also study how varying controlling parameters affect overall performance with different similarity metrics. We analyze our empirical results in terms of preciseness and performance.Keywords: Similarity measures, prediction, top-N recommendation, accuracy, performance.
Improving performance of privacy-preserving collaborative filtering schemes
Privacy-preserving collaborative filtering methods offer useful filtering skills without deeply jeopardizing individual privacy. However, they mostly suffer from accuracy, scalability, and sparseness problems. Applying privacy measures to conceal confidential data in recommendation systems causes a bias in collected data, which might make accuracy worse. As the content in recommendation domain proliferates, the size of collected data expands rapidly, which aggravates scalability challenge of those systems. In addition, since users are typically able to rate a small fraction of existing products, sparseness of collected data becomes an issue.In this dissertation, various preprocessing methods are proposed to overcome accuracy, scalability, and sparseness challenges faced by various privacy-preserving collaborative filtering systems. Through application of the proposed preprocessing techniques like item ordering and elimination, clustering, dimensionality reduction, user profiling, profile cloning, and son on, novel privacy-preserving collaborative filtering schemes are cultivated. Essentially, the proposed enhanced systems focus on producing accurate predictions while coping with constantly growing nature of collections without jeopardizing individual privacy. The proposed schemes are analyzed in terms of privacy and overhead costs. Also, real data-based experiments are performed to scrutinize their effects on accuracy, scalability, and privacy. The analysis and experimental outcomes demonstrate that the methods preserve individual privacy and offer adequately accurate recommendations in scalable amount of time.Keywords: Preprocessing; Privacy; Scalability; Accuracy; Sparsity; Collaborative ?ltering.
Lipopolisakkarit ile indüklenen hipotermik ratlarda kan serumu aldosteron hormonu ve bazı mineral düzeylerindeki değişim
Bu tez çalışmasında, sıçanlara LPS (Serotip E.coli: O111:B4; 250 µg/kg, ip) uygulanarak hipotermi oluşturduk. Oluşturulan hipoterminin girişinde, en derin noktasında ve çıkışta intrakardiyak delinme ile kan örnekleri alındı. Yine aynı çalışmada kontrol grubu sıçanlara tuzlu su (% 0.09'luk NaCl, 0.5 ml/kg ip) uygulanarak hipotermi giriş, en derin noktasına ve çıkışına dek gelen zaman dilimlerinde kan örnekleri alındı. Uygulamalar sonucunda hipoterminin en derin (ΔT: -1.48±0.09 ºC) noktasına ortalama 104.46±4.32 dakikada ulaşıldı. Alınan kan serumu örneklerinin analizleri sonucunda aldosteron, Na, K, Cl, Ca ve P düzeyleri kontrol grubuna göre LPS indüklü hipotermi grupları (giriş, dip ve çıkış) arasındaki farklar istatiksel olarak önemsiz olduğu belirlenmiştir. Elde etiğimiz sonuçlar, rodentlerde LPS ile termoregülatuvar anlamda ateşle birlikte hipotermininde gelişebileceği göstermiştir. Bu çalışmadaki LPS dozunun endotoksik etkileşim süresinin aldosteron seviyesini değiştirmediğini göstermiştir. Yine bu araştırma sonucunda elde edilen Na, K, Cl, Ca ve P düzeylerinin değişmemesinin LPS indüklü hipotermi süresince aldosteron hormon düzeylerinin değişmemesi sonucu olduğunu düşündürdü.
Dişi akkeçilerin hematolojik parametrelerinin yıllık değişimi
Bu tez çalışmasında, dişi Akkeçilerde (1,5 yaşlı n: 5; 2,5 yaşlı n: 5) yıllık hematolojik parametrelerdeki değişim araştırıldı. Araştırma sonucunda eritrosit (RBC), hematokrit (Htc) ve hemoglobin (Hb) seviyeleri, sıcaklığın yüksek olduğu aylarda azalma, sıcaklığın düşük olduğu aylarda artış istatistiksel olarak önemlidir (P<0.01). Dişi Akkeçilerde lökosit (WBC), lenfosit (Lym), granülosit (Gra) ve monosit (Mon) aylık seviyelerindeki değişim istatistiksel olarak önemli (P<0.01) olduğu tespit edilmiştir. Yine aynı araştırmada WBC, Gra ve Mon yüksek sıcaklık gösteren aylarda artış, Lym düzeylerinde azalma tespit edilmiştir. Trombosit (THR) değerleri üzerinde mevsimsel değişimin (sıcaklık) önemli (P<0.01) etkisi olup yüksek sıcaklık gösteren aylarda artış belirlenmiştir. Yüksek sıcaklık gösteren aylarda RBC, Htc, Hb ve THR değerlerindeki azalmanın yaz aylarında kan plazma hacminin genişlemesinin sonucu olabileceğini düşündürdü. Araştırmada WBC, Gra ve Mon değerlerindeki artışın yüksek sıcaklık gösteren aylarda immünolojik düzenleme sonucu ve farklı fizyolojik dönemlerin etkisinin olabileceğini düşündürdü.
Optimizasyon temelli öznitelik seçme yöntemleri ile desteklenen topluluk öğrenme yaklaşımına dayalı yazar tanıma
İnternet ve özellikle sosyal medya aracılığıyla veri arama, kopyalama ve yayma fırsatlarının artması doğru bilgiye ulaşımı azaltmıştır. Veriden doğru bilgiye ulaşma konusunda metin madenciliği alanında yapılan çalışmalardan biri metin yazarı tahminidir. Bir metin, onu yazan kişinin karakteristik özelliklerini taşır ve bu özellikler metnin yazarını tanımlamak için kullanılabilir. Bu çalışmada 54 adet yazar, değişken sayıda ve değişken uzunlukta toplam 46.837 adet köşe yazısı ile bir derlem oluşturulmuştur. Yazarların karakteristik özelliklerini çıkarmak için iki farklı analiz ve iki analizin birleştirilmesi ile oluşturulmuş karma analiz hazırlanmıştır. Analiz sonuçlarının verimliliğini artırmaya yönelik Rastgele Orman Algoritması ile Genetik Algoritma ve Tavuk Sürü Optimizasyon Algoritması yaklaşımlarına dayalı toplamda iki farklı öznitelik seçim yöntemi sunuldu. Yapılan analizler sonucunda en verimli yazar tahmini çözüm önerisi, Topluluk Öğrenimi Algoritmaları ile desteklendi. En iyi performans, Karma Analiz ve sonrasında Genetik Optimizasyon algoritması ile gerçekleştirilen işlemler sonucu oluşturulan on yazarlı veri kümesi üzerinde Torbalama Algoritmasında sınıflandırıcı metot olarak Karar Ağacı kullanımında % 95,74 olarak elde edilmiştir.
Derin öğrenme mimarileri kullanılarak ayrık video görüntüleri üzerinden işaret dili tanıma
İşaret dilleri, işitme ve konuşma engelli bireylerin günlük yaşamda kullandıkları ana iletişim ortamları olan görsel dillerdir. Çok sayıda kanal üzerinden aktarılan işaretlerin bilgisayarlı tanınması sayesinde, işitme ve konuşma engelli bireyler hem diğer bireylerle hem de makineler ile iletişimlerini doğal şekilde yapabileceklerdir. Bu tez çalışmasında, derin öğrenme kullanılarak ayrık işaret dili videoları üzerinden işaret dili tanıma gerçekleştirilmiştir. BosphorusSign veri kümesinin "genel" isimli alt kümesi kullanılarak yapılan çalışmada, öncelikle veri artırma ve önişleme parametrelerinin belirlenmesi için çalışmalar yürütülmüştür. Ardından çeşitli derin öğrenme modelleri kullanılarak yapılan deneyler sonucunda işaret dili tanıma için kullanılabilecek uygun bir model belirlenmiştir. Daha sonra işaret dilindeki çeşitli kanalları ifade etmek üzere çıkarılan farklı veri kiplerinin tek başlarına ve çeşitli birleşimlerle başarımları değerlendirilmiştir. Bu sayede çok kipli bir işaret dili tanıma için kullanılacak en uygun veri kipi kombinasyonu elde edilmiştir. Son olarak, deneyler sonucunda elde edilen parametreler ve veri kiplerini kullanan çok kipli bir işaret dili tanıma modeli önerilmiştir. Önerilen model, RGB, eklem ve optik akış kiplerinde toplamda 6 farklı veri akışını bir arada girdi olarak almaktadır. Model bünyesindeki birleştirme mekanizması ile veri akışlarından çıkarılan öznitelikler birleştirilmiş ve derin öğrenme tabanlı sınıflandırıcı katmanlara aktarılmıştır. Uçtan uca bir yöntemle eğitilen bütünsel işaret dili tanıma modeli, kullanılan veri setinde görülen en yüksek başarım olan %89,3 doğruluk sunmuştur. Önerilen çok kipli işaret dili tanıma modelinin işaret dili tanıma başarımını iyileştirme konusunda geliştirilebilir bir potansiyeli vardır.
Shilling attack design and detection on masked binary data
Privacy-preserving collaborative filtering methods are effectual ways of coping with information overload problem while protecting confidential data. Their success depends mainly on the quality of the collected data for filtering purposes. Malicious entities might create fake profiles (noise data) and insert user-item matrices of such filtering schemes. Hence, shilling attacks play an important role on the quality of data. Designing effective shilling attacks, developing methods to detect them, and performing robustness analysis of privacy-preserving collaborative filtering methods are receiving increasing attention. In this thesis, six well-known shilling attack models are modified in order to attack binary masked databases in privacy-preserving collaborative filtering methods. Three attack design approaches are proposed. The attack profiles, generated by such schemes, are applied to naïve Bayesian classifier-based collaborating filtering scheme with privacy. A novel shilling attack detection scheme based on classification is proposed to detect fake profiles. Attributes derived from user profiles are utilized for detecting shill profiles. Empirical results show that designing effective shilling attacks is still possible on binary masked data. The proposed detection method is able to successfully detect fake profiles. Keywords: Shilling Attack, Collaborative Filtering, Privacy, Binary Data, Detection, Robustness