
19
Archived Theses
0
DOIs Assigned
0%
DOI Rate
Discipline
Detection of misinformation related to pandemic diseases using machine learning techniques on social media platforms
COVİD-19 salgınının ortaya çıkışı beraberinde getirdi. Bu sadece küresel bir sağlık krizi değil, aynı zamanda bir bilgi salgınıdır. Sosyal medya platformlarında yanlış bilgiler hızla yayılıyor. İçeri Etkili yanlış bilgi tespitine yönelik acil ihtiyaca yanıt olarak, bu çalışma, makine öğreniminden yararlanan kapsamlı bir yaklaşım sunuyor ve topluluk yöntemleriyle sonuçlanan derin öğrenme teknikleri, Facebook ve Twitter'da COVID-19 ile ilgili yanlış bilgilerin yayılmasıyla mücadele etmek,Instagram ve YouTube. Kullanıcıyı içeren zengin bir veri kümesinden çizim yapın COVID-19 ile ilgili çeşitli konuları kapsayan bu platformlara ilişkin yorumlar tartışmalar, araştırmamız Destek Vektör Makinesi (SVM), Karar ağ ağacı, lojistik regresyon ve sinir ağlarının derinlemesine gerçekleştirilmesi Yorumların analizi ve sınıflandırılması iki kategoriye ayrılır: olumlu ve olumsuz bilgi. Yaklaşımımızın yeniliği finalde yatıyor Güçlü yönleri güçlendirmek için topluluk yöntemlerini kullandığımız aşama Çeşitli makine öğrenmesi ve derin öğrenme algoritmalarının kullanımı. bu topluluk Yaklaşım, modelin genel doğruluğunu ve uyarlanabilirliğini önemli ölçüde artırır.yetenek. Deneysel sonuçlar metodolojimizin etkinliğini vurgulamaktadır. ile karşılaştırıldığında algılama performansında önemli gelişmeler sergiler.bireysel modeller. Topluluk öğrenimini uyguladıktan sonra Facebook verilerinin %91'i müstehcen, Instagram verilerinin %76'u, Twitter'ın %81'i veriler ve %93'i YouTube verileri için şeklinde bir sonuca varıyoruz. Sistemimiz yalnızca önlemeye yardımcı olmakla kalmıyor. COVİD-19 ile ilgili yanlış bilgilerin yayılmasını önlerken aynı zamanda Çeşitli bağlamlarda yanlış bilgilerin ele alınmasına yönelik daraltılmış çerçeve.sosyal medya platformları.
Point in time probability of default modeling for international financial standards - a Turkish bank case study
Point in time probability of default modeling plays a crucial role in the context of International Financial Reporting Standards 9 (IFRS 9), which expects the measurement and recognition of expected credit losses (ECL) for financial instruments. IFRS 9 introduces a forward-looking approach, necessitating the estimation of one-year and lifetime PDs to capture credit risk over the entire expected life of financial assets. This thesis presents an in-depth analysis of PiT PD modeling in the framework of IFRS 9, highlighting its significance, methodologies, and implications for financial institutions. A case study of a Turkish Bank examined and established an autoregressive macroeconomic model to forecast the default rate (DR) of small and medium enterprise loan segments using the autoregressive linear model (ARLM) method. Results indicate that interest rates positively affect DR, and the USD-TRY exchange rate negatively affects DR. A basic quantitative validation of the DR model is implemented, and the forecast power of the DR model is examined by using out of time (OOT) period of DR. Finally, forward PDs for 10 years are calculated using adjusted Weibull distribution. Forward PDs are calibrated with the result of the ARLM model forecast under different scenarios. Thus, an applicable PiT PD term structure for SME portfolio is created.
Fine-tuning and hyperparameter optimization; a path to superior model performance and accurate predictions
This thesis explores the use of data mining tools and techniques in business intelligence, focusing on how they can enhance the process of obtaining insights and knowledge from massive datasets, enabling more effective decision-making in businesses. The research is divided into sections that explain the core principles of business intelligence and data mining, as well as how they can be combined to improve business analytics. A practical application and case study focusing on forecasting song popularity on the Spotify platform showcases the usefulness and potential of data mining techniques in a real-world context. The case study describes data collecting, preprocessing, and feature engineering methodologies, as well as data mining techniques used to find patterns and trends in music data. The study adds to the existing literature by demonstrating the actual implementation of data mining tools and methodologies for business intelligence. This study adds to the current literature by demonstrating the actual implementation of data mining tools and methodologies for business intelligence. This thesis highlights the potential benefits and problems of integrating data mining techniques in real-world contexts by concentrating on the specific instance of forecasting song popularity on Spotify. The insights provided will help music business experts, marketing teams, and artists make informed decisions about song releases, promotional methods, and target audience selection. Overall, this thesis emphasizes the importance of data mining tools and methodologies in improving business intelligence operations and bridges the theoretical and practical divides by providing a real application and case study that demonstrates the power of data mining approaches in deriving actionable insights from massive datasets.
Fraud detection using big data tools and machine learning in banking
The use of big data analytics applications is prevalent in the banking industry, owing to the abundance and quality of customer information and transaction records available through online and offline channels. Processing such data through machine learning algorithms can greatly benefit decision-making processes. Big data applications can be employed by banks to identify fraudulent money transfer transactions, which pose a substantial risk to their financials and reputation. This study presents information on rule-based systems and big data applications for fraud detection in banks. Digital money transfer data obtained from a private bank was subjected to different supervised and unsupervised classification models, including extreme gradient boosting and isolation forests, and their results were compared. The extreme gradient boosting model displayed superior performance, while the unsupervised isolation forest algorithm provided notable outcomes. It was also concluded that the application of big data analytics and machine learning significantly contributes to fraud detection.
Does the mood affect a player's performance? Sentiment analysis using football data
Football teams attempt to get as much consistent information about players as they can, which is like what currently happens in other types of activities in the sports' world. Today, it's critical to manage a variety of factors in addition to physiology, diet, and health indicators. Goals scored and assists are just a few examples of how a player's performance may be judged objectively. It's a good approach to compare and rank the top players in each category. The relationship between a player's mood and their performance on the field is a topic of interest in the field of sports psychology. Sentiment analysis, a method used to extract and analyze emotions and opinions expressed in text data, has been increasingly used in the field of sports to study player and team performance. This thesis aims to study the relationship between a player's mood and their performance on the field by using senti-ment analysis on football data. The main research question of this study is: Does the mood affect a player's performance? By analyzing the sentiment of social media posts, specifically tweets of individual football players, the sentiment of these tweets is ana-lyzed in relation to the performance of the players on the field. The study uses natural language processing techniques to extract sentiment from tweets. The findings of this study have the potential to inform the practice of sports psychology and improve the performance of players on the field.
Predicting participant risk profiles in private pension funds using machine learning techniques
Individuals in the private pension system can voluntarily participate in the risk profile assessment survey before they decide to direct which funds are eligible for their savings. Thus, the system proposes suitable funds in accordance with the result of the questionnaire. Nevertheless, participants who do not fill out this questionnaire are in the majority within the private pension system. The aim of this thesis is to predict the risk profile level of participants who do not fill out the survey by utilizing the information of others ones who have already filled out the survey. In this scope, the model has been built by using machine learning techniques through data including financial, demographics, and other features which belong to the customers who completed or did not complete the questionnaire. It has been shown that the XGBoost algorithm (F! Score: 59%, Accuracy: 60%) which has been applied to four risk categories distributing almost balanced is the best one among the machine learning model for prediction. The built model has proven its usability as a supplementary tool by testing both the private pension company and the fund advisors.
Use of a deep learning CNN architecture in product image quality assessment for improving e-commerce customer experience
In the dynamic world of e-commerce, the visual impact of product images holds incredible sway over consumer perceptions and choices. This study delves into the intriguing interplay of artificial intelligence (AI) and e-commerce, proposing a groundbreaking AI-driven model that precisely assesses image quality from the customer's perspective. Through the careful curation of a diverse dataset, we crafted a sophisticated convolutional neural network (CNN) architecture. Impressively, our model achieved a remarkable 98% accuracy on the test dataset, demonstrating its prowess in categorizing images accurately. Moreover, it is imperative to highlight that the dataset itself was meticulously created from scratch, with the images designed and integrated directly into the model. This bespoke dataset served as the cornerstone for training the CNN architecture. While this accomplishment is noteworthy, it's important to acknowledge the limited size of our training data. This raises important considerations about the model's adaptability to a broader range of visual inputs. To address this, we applied innovative techniques such as data augmentation and model regularization, fortifying the model's ability to handle new, unseen data. Furthermore, our research extends beyond model development. We have created an interactive website (Visual Analysis Platform) that allows users to experience firsthand the capabilities of our model in assessing image quality based on the categories established in our research. This platform serves not only for analysis and user engagement but also for collecting user-generated images. These images contribute to our dataset enrichment, aiming for continuous improvement in the model's ability.
Big data characteristics and decision-making:The mediating role of knowledge management
This work discovers the influence of big data characteristics on decision-making through knowledge management orientation in companies. Quantitative methods were employed in collecting data through a questionnaire for a random sample of 331 employees in the information technology sector. Findings demonstrate a statistically significant correlation between big data and decision-making through knowledge management orientation. These revelations have significant implications for companies that must navigate the confluence of knowledge management, big data, and decision-making.
Introducing forward looking information to expected credit loss under IFRS 9 : Different time series approaches for Turkish banking system
Credit risk is the most critical type of risk to manage for banks, which are the most important part of the financial system. It is expected that the capital of banks, whose most important activity is to provide loans, will be well structured and will be at a level that will protect the bank against risks. Credit risk is an element that must be managed not only against its realization, but also against the possibility of its occurrence. Credit risk should be managed against both internal factors (such as bank portfolio characteristics, characteristics of bank customers) and external factors (such as economic conjuncture, political and political outlook, natural disasters). When considering the credit risk management against the external factors, finding a relationship between the risk factors such as non-performing loans ratios of the banks and key macroeconomic indicators and modelling them become an issue for the banks. It has become critical to develop time series models by establishing a relationship between credit risk factors and macroeconomic indicators, in order to include forward-looking information in expected credit loss calculations in accordance with the International Financial Reporting Standard (IFRS 9), which entered into force at the beginning of 2018 and both as a stress test. In this study, it has been tried to measure the effectiveness of different time series algorithms comparatively, especially since a certain method is not imposed within the scope of IFRS 9 standard. Different time series models were established between the non-performing loan rates published by the Banks Association of Turkey and all potential macroeconomic indicators officially published, and diagnostic tests were applied to measure the robustness of the models. As algorithms, Stepwise Regression, Autoregressive Distributed Lag Model (ARDL) and Extreme Gradient Boosting Model (XGBoost) were chosen. As a result of this study, Autoregressive Distributed Lag Model showed the best performance. As expected credit loss models are regulative models, they are subject to the examination of many regulatory authorities. It has been demonstrated that ARLD model which also has high explainability and interpretability, is a method that can be used frequently by banks.
Developing a life insurance recommendation system using machine learning methods
In the last 10 years, the use of advanced technologies and big data handling methods in the field of artificial intelligence has led to an increase in the number of machine learning-based projects in many sectors and domain such as personalized product offerings that enhance customer loyalty and business value. Algorithm based development has been ongoing for 70 years and continues to grow. The use of machine learning techniques in the insurance industry has the potential to greatly improve customer satisfaction and increase company profitability. In this study, by collecting and analyzing data on the portfolio movements, payment behavior, and demographic characteristics of existing product owners, predictive models were conducted to identify potential customers for cross-selling. This study followed data preprocessing steps, including handling missing data, detecting, and repairing outliers, and preprocessing categorical data for use in the model. The prediction problem was treated as a classification problem, and explanatory data analysis and correlation analysis were performed to gain a deeper understanding of the data. The results of this study could be used to inform future efforts to personalize product offerings and increase sales in the insurance industry. The prediction problem was addressed using supervised learning methods, including Decision trees, Logistic regression, Random forest algorithms, Naive Bayes and Gradient boosting algorithms. The performance of the models was optimized through scenario-based experiments, and the effects of various data preprocessing steps, such as normalization and dimensionality reduction, on model performance were observed. The performance of the models was evaluated using a range of metrics, including accuracy, AUC, and F-1 scores. The results of this study suggest that hyperparameter tuning can play a significant role in improving the performance of machine learning models in this context. Overall, the use of machine learning techniques has the potential to greatly enhance the accuracy of predictions and improve decision-making in the insurance industry.
Big data analytics: Using big data analytics in tracking player performance and scouting in football
In an era characterized by emerging tech and digital transformation, the sports industry is leveraging big data analytics which has given rise to fields like sport an-alytics (SA). In the game of football (soccer), the exponential rise in available data has led to lots of innovation and research on how data can be maximized in performance anal-ysis, scouting, injury prevention, management etc. This research provides of overview knowledge into big data analytics in sports with emphasis in football while paying attention to player performance tracking and scouting. Scouting is reliant on player performance tracking and as such a data driven decision approach will be beneficial in a system still dominated by bilateral relations and intuitive comments of scouting teams. A model is presented in this research for tracking performance, rating and rec-ommending players using data acquired. The data consist of five national European competitions in the 2017/2018, Euro 2016 and 2018 World cup. The data was used with permission from Luca Papalardo et al as it was first used in his research paper "A public data set of spatio-temporal match events in soccer competitions."
Stock variation in optimum portfolios with mean-variance and mean-semivariance approach: A comparative analysis for emerging markets
The main objective of this thesis is to determine the optimal number of equities required to be included within portfolio for each single country by applying both mean-variance and mean-semivariance approaches and then compare them. All the analysis has been conducted in 8 emerging countries which are Brazil, China, India, Mexico, South Africa, Turkey, South Korea, and Pakistan. Data from Bloomberg spanning Sep'16 to Sep'21 was used. Emerging countries are chosen within the scope of this thesis due to their nature of skewed and non-normal equity return distributions, which make downside risk analysis more appropriate. Findings show that mean-semivariance optimization generally results in portfolios with fewer equities, except for Brazil. The recommended number of equities typically ranges from 12 to 18. The results reveal that optimal portfolios of the mean-semivariance model contain significantly fewer equities compared to those determined with the mean-variance model. This suggests reduced turnover and transaction costs and thus increased net returns.
Predicting students' academic achievement with educational data mining: Using interaction data from learning management system
The purpose of this study is to predict the end-of-term passing grades of the students enrolled in the GEP0322-Digital Literacy course in the fall semester of the 2020-2021 academic year, using only learning management system interaction data. Different feature selection methods were used to examine the effect of feature selection techniques on model performance. In all cases, different algorithms worked efficiently, but their prediction results were not the same. While there were some minor differences between model values, it wasn't enough to say one was better than the other. CatBoost regressor (RMSE=14.27), Random Forest regressor (RMSE=14.39), and Extra Trees regressor (RMSE=15.84) were the best models in the conducted cases respectively.
Big data applications in banking credit risk audits: Sampling through machine learning
Banking is one of the industries where big data analytics applications are widely used. Banks have precious data in quantity and quality, thanks to customer information and transaction records obtained through online and offline channels. It will be very beneficial to process those data via machine learning algorithms and exploit them effectively in decision-making processes. Banks can use big data applications to audit credit risks, one of the most significant banking risks. In this study, information about big data applications in credit risk audits of banks is provided, and how these can be used in audit sampling is explained through machine learning models. Different classification models, such as decision trees and random forests, were applied to corporate customer data obtained from a private bank, and their results were compared. As a result of the study, the random forest model showed the best performance. In addition, it has been concluded that big data analytics and machine learning applications significantly contribute to the sample selection of credit risk audits.
Web scraping in ecommerce and designing a blocking prevention method for web scraping
The development of new technologies in the modern world has great consequences for the development of numerous areas in Internet and in big data environment. With this affect e-commerce has become a great online source that is widely used and has huge amount of big data, which can be analysed for future decisions. Web scraping technology is used to scrape the data from online platforms. In this study, a web scraper was created. One of the most popular and leading e-commerce web site in Turkey was taken for testing the web scraper. With the help of web scraping, you can track business processes, customer behaviour, their preferences and demand for the goods. However, online platforms have a blocking mechanism that blocks web scraping programs. That is why, the created scraper also faced blocking. Therefore, contribution to this research work was creating a method to prevent blocking, which consists of three algorithms and called "Hybrid humanoid scraper". The first algorithm was introduced, which was not blocked, but did not have a review part. After that, the second low and high frequency algorithm was introduced with a review part. By using the low frequency, the web scraper retrieved data successfully. However, with high frequency scraper was blocked. Finally, it was possible to use a high frequency by adding the third algorithm and bypass the blocking of web scraping by changing IP addresses on the e-commerce platform. By using web scraping and algorithms to prevent web scraping blocking on the web platform, ethical and legal standards were not violated.
A comparative study on node classification methods for undirected social networks
Traditional Neural Networks can solve problems with normal 1-dimensional and 2- dimensional Euclidean data such as image and text classification. However, most real- life problems are relationship-based and the corresponding real-world data has a non- Euclidean nature. This paper presents a comparative study between Graph Neural Networks (GNNs) and traditional approaches (K-Nearest Neighbour (KNN)) for solving real-life problems involving non-Euclidean data with complex relationships. The results demonstrate that GNN models outperform KNN in terms of accuracy while maintaining low runtime on graph structured data. The study highlights the strengths and limitations of Graph Convolutional Networks (GCNs) and GraphSAGE, with GraphSAGE offering flexibility and scalability but potentially introducing aggregation bias. Ongoing experiments include investigating the sensitivity of the algorithms to hyperparameter values. Overall, this research contributes to the understanding and practical application of GNN methodologies in various domains.
Analysis of the Turkish job market and ict industry labor demands using web scraping
This research provides a comprehensive analysis of Turkish job market trends using web scraping. Utilizing Octoparse8, we extracted a wealth of job postings data from online job platforms like LinkedIn and Kariyer.net. The data was then prepared and cleaned using Python's Pandas library, involving processes such as importing, formatting, sampling, handling missing values, removing duplicates, data type conversions, and outlier treatment. Our analysis tested several hypotheses and explored various relationships within the job market, including those between company size, job advertisement frequency, and in-demand programming skills. We found a positive correlation between larger corporations and the number of job ads. Furthermore, our study identified the most sought-after programming skills in the Turkish market, such as Javascript, SQL, HTML, C#, and Python, and highlighted high-demand job roles and sectors with significant job growth rates. v We employed various data visualization techniques, including bar plots, histograms, scatter plots, heatmaps, box plots, and pie charts, to present the data in a graphical format, facilitating a better understanding of patterns, trends, and insights in the job market. The findings from this study not only contribute to the existing literature on the application of web scraping and data aggregation techniques in job market analysis, but also provide valuable insights for various stakeholders in the job market ecosystem, including job seekers, employers, educational institutions, and policymakers. The study demonstrates the power of Python and Jupyter Notebook in conducting such an analysis, highlighting the utility of libraries like Pandas, Matplotlib, and Seaborn. This research lays the groundwork for future predictive job market models.
A machine learning-driven sales forecastingapproach for newly launched products
Sales forecasting is one of the most challenging problems for companies operating in the fashion industry. The effects of demand volatility, trend, constrained lead times, huge product variety, and seasonality are vital to demand forecasting. Most importantly, fashion retailers are struggling with forecasting future demand because of the seasonality effect and short life cycles of the products. Even though it is a generally accepted opinion, forecasting future demand of newly launched products is even more challenging, because of the fact that new products do not have any historical sales data. When the number of organizations operating in the fashion industry is considered, satisfying customer demand becomes one of the most significant features for success and profit improvements. This study introduces a machine learning-driven sales forecasting approach for newly launched products. The proposed model applied on a fashion retailer's censured historical sales data. On the other hand, the proposed algorithm combines product-level predictions with chain-level predictions to provide better sales forecasting accuracy for fashion retailers and for newly launched products. Once the predictions are obtained from the model, they are measured with traditional accuracy metrics.
Predicting chronic kidney disease using ML by using two different datasets: A comparison study
The kidneys, filter waste products from the blood, produce urine, and a hormone that is important for the red blood cells, blood pressure regulation, and calcium metabolism. When the functioning of the kidneys gradually declines over time, Chronic Kidney Disease emerges. Glomerular Filtration Rate is a strong indicator of this disease which can be effectively predicted by using Machine Learning techniques. In this study, the performances of eight Machine Learning models and an ensemble of them are compared using two different public datasets and a combined dataset of these two. One of the datasets is from India with 400 records, and 26 attributes from the year 2015, and the other dataset is from Bangladesh with 200 records and 29 attributes from the year 2021. Cross-checking is performed using the eight models and an ensemble of them, such that the nine models trained on one dataset are tested on the other dataset which is unseen by the models. The nine models are trained and tested on a combined dataset of the two. Also, the models are used on a reduced set of features of the combined dataset.