Theses supervised by Yrd. Doç. Dr. Abdül Kadir Görür

24 theses · Çankaya University

Master'sOpen AccessEN

Sentiment analysis and gender prediction in twitter data

In this thesis, tweets from Twitter that have been sent by users will be considered on a preferential basis in accordance with determined or requested specific key word(s). Also the interpretation of these tweets, by the computer, will be examined in a way as they are "Positive", "Negative" or "Neutral". In this context, under the heading 'Twitter Sentiment Analysis', studies were conducted and the success rates of achieved results were compared. In addition to this, on the basis of the usernames(of users) who send tweets, tweets was compared with Turkish Special Names which is shared by the Turkish Language Association (TDK) and also achieved results and gender determinations of users in terms of "Female", "Male" or "Not Determined," were examined. Under the heading of 'Gender Prediction in Twitter' studies were conducted and the success rates of achieved results were compared. On the basis of this study, the related topics of 'Sentiment Analysis' and 'Gender Prediction' were examined for Turkish Language and all of these studies were carried out through Turkish language.

Ertuğrul Balaban
Çankaya University · Institute of Graduate Studies in Science
2015
00
Master'sOpen AccessEN

Investigation in MYSQLdatabase and NEO4J database

Currently, there are two major of database management systems which are used to deal with data, the first one called Relational Database Management System (RDBMS) which is the traditional relational databases, it deals with structured data and have been popular since decades from 1970, while the second one called Not only Structure Query Language databases (NoSQL), they have been dealing with semi-structured and unstructured data; the NoSQL term was introduced for the first time in 1998 by Carlo Strozzi and Eric Evans reintroduced the term NoSQL in early 2009, and now the NoSQL types are gaining their popularity with the development of the internet and the social media. NoSQL are intending to override the cons of RDBMS, such as fixed schemas, JOIN operations and handling the scalability problems. With the appearance of Big Data, there was clearly a need for more flexible databases. In this study, a theoretical study of investigating in two types of databases MySQL one of the traditional relational databases, and Neo4j one of the graph databases. First of all, choosing these two types of the databases according to their features, both of them depending on the replication and sharding in their systems. Secondly, a brief background with literature review will introduce the opinions for some researchers with some concepts. Thirdly, mentioning relational database and NoSQL in generally and MySQL and Neo4j in specifically, then try to make a comparison between the features for both of them. Moreover, all the websites, the researches, and the journals are the resources for this study. In addition, this study is like mentioned before is a theoretical study so there was no implementation work. However, the comparison presented the result in this study. All in all, we try to provide an understanding for MySQL and Neo4j, that leads us to find both of MySQL and Neo4j have their features, pros and cons that differ from each other. Keywords: RDBMS, NoSQL, MySQL, Neo4j, Replication, Sharding.

Zahraa Mustafa Abdulrahman Al-ani
Çankaya University · Institute of Graduate Studies in Science
2015
00
Master'sOpen AccessEN

Özellik belirleme matriksinin metin siniflandirma sisteminin performansi üzerindeki etkisi

Text Categorization (TC) is an important intelligence information processing technology. This technology has high value in information retrieval, Electronic Governments, information filtering, text databases, digital libraries, and other aspects, but the problem of feature selection is equally or more important than text-categorization. In this thesis, we did our experiments with the help of standard Reuters-21578 dataset, and we discussed many important topics ranging from collecting data, to organizing data and ultimately using the organized data to efficiently conduct tests using the feature selection metrics.The general idea of any feature selection metric is to determine importance of words using some measure that can keep informative words, and remove non-informative words, which can then help the text-categorization engine categorize a document, D, into some category, C. The feature selection metrics that will be discussed in this thesis are: Term frequency-Inverse Document Frequency (TF-IDF), Document Frequency (DF), Mutual Information- Explanation (MI), Chi-square Statistics (CHI), GSS (Galavotti-Sebastiani-Simi) Coefficient – Explanation. It will combine Term frequency-inverse document frequency (TF-IDF) and Documents Frequency (DF) metrics to prepare the texts in a perfect way. After that, those texts will be used by classification process in Weka to get the best learning machines algorithms and the best performance of system, by computing performance measures such as (accuracy, error rate, recall, precision and F-measure). We compare the reusability of popular active learning algorithms for text classification and identify the best classifiers to use in active learning for text classification. All these mentioned measures were computed and plotted.

Text categorization
Asmsaa Al-gartanee
Çankaya University · Institute of Graduate Studies in Science
2015
00
Master'sOpen AccessEN

Combined JPEG with DWT or image compression

The main idea for this research is apply JPEG technique with enhanced quantization, on Discrete Wavelet Transformation (DWT). The DWT minimize the size of the image into convert the image into LL, HL, LH and HH sub-bands. The LL sub-band represents approximately original image, and then applied JPEG technique on the LL sub-band, then applied quantization to increase number of zeros. Finally apply Sequential Search Coding on the JPEG matrix for coding. The Sequential Search Coding work to convert group of data into single floating point value by using Key.

Omar Nabeel Dara
Çankaya University · Institute of Graduate Studies in Science
2012
00
Master'sOpen AccessEN

Statistical based detection of DOS: (Denial of Service) attacks

Automated detection of anomalies in network traffic is an important and challenging task. Anomalies and intrusion are the main factors the affects the network performance. Detecting intruders that are aiming in degrading and preventing access to certain web pages or network address is the aim of many research works. In this work we propose an automated system to detect Denial of Service (DoS) attacks by using statistical methods. Related research work in this field will be surveyed and the aim will be the design of an enhanced method. Visual Basic .Net will used in designing the programs.The proposed program will be tested on real environment and its effectiveness will be verified.

Tariq Abed Mohamad
Çankaya University · Institute of Graduate Studies in Science
2012
00
Master'sOpen AccessEN

Evaluation of taxonomy based concept extraction system cosmix case for text categorization

The aim of this study is creating a Document Classification system using Vector Space Model as baseline classifier. Cosine similarity is used to calculate similarity between Training Set and Test Set. Finally similar files are used to suggest topics for test files. Same method is used to create Kosmix Training and Test Sets and suggest topics. Results are compared and comparison results whon that Cosine Similarity method is more successful.

Vector spaces
Umut Ergün
Çankaya University · Institute of Graduate Studies in Science
2011
00
Master'sOpen AccessEN

Sentiment analysis and opinion mining via microblogging in social media like: Twitter

This research is a study of microblogging on social websites such as Twitter and shows the techniques of emotion detection and sentiment analysis for the same. This research has three objectives. The first objective is a discussion about how to extract and classify emotions in tweets using the unigram feature extractor with word presence or word frequency as a factor of extraction. High accuracy of classification is obtained when considering the word presence as a factor of extraction. Moreover, one can obtain high accuracy also by using word frequency as a factor of extraction when supplying the test data on training corpora of tweets in the case of multi-domain tweets. The second objective is the extraction and classification of the emotions of tweets using n-gram (1

Mustafa Salman Abd Al-bndi
Çankaya University · Institute of Graduate Studies in Science
2015
00
Master'sOpen AccessEN

Development of tool for managing semantic text content

The aim of this study is creating multi-document summaries using latent semantic analysis and centroid based approach. First, key-terms are extracted using latent semantic analysis (LSA). Key-terms are used to filter the redundant sentences before sentence extraction. Then summary sentences are extracted from the sentences containing the key-terms using latent semantic indexing (LSI) and centroid-based method with clustering consecutively.

Samet Karakaynak
Çankaya University · Institute of Graduate Studies in Science
2009
00
Master'sOpen AccessEN

3D visualization using data received from the processes of object recognition and object reconstruction

This thesis presents the demonstration of what two images or two video sequences can tell us about the situation and model of a third video sequence or image. The method bears ideas from projective geometry as it?s basis.The main purpose of the thesis is to be able to form a base line for tracking an object in a 3D environment not only by using two stereo cameras but also by using other cameras that may be located in various points of the environment. The method visualizes the object and gives the information to a third camera. This way it can be possible to track a moving object, along with it?s visualized model, in an environment without losing sight of it and without having to move the other two stereo cameras which we received data from.

Three dimensional imaging
Seher Pelin Güvenç
Çankaya University · Institute of Graduate Studies in Science
2008
00
Master'sOpen AccessEN

Stereo video karelerinde görüntü izleme

Tracking in video refers to the process of locating moving objects in the following video frames. It is an important topic and has various application fields such as robotics, military, etc. In this thesis, moving object tracking in stereo video sequences is studied. Tracking is done by several processes. First, the object, that is going to be tracked, is detected, and extracted from its background. Then, a noise elimination process follows that, where we detect false candidates and eliminate them. After key features of the target object are gathered, we can track the object by searching these features in the following frame sequences.

Image viewingStereo systemsVideo
Serkan Kefel
Çankaya University · Institute of Graduate Studies in Science
2008
00
Master'sOpen AccessEN

3D reconstruction of a scene using stereo images

Two-dimensional photographs do not have depth-information. One solution to determine the location of an object in three-dimensional environment is to use more than one photograph as exposed by the nature. Extracting the depth information using stereo images is purposed in this thesis.The thesis analyzes the steps and encountered problems in three-dimensional reconstruction process, explains the solutions exposed with the aid of epipolar geometry using some of the feature-based matching techniques. Stereo images which are taken from two calibrated cameras viewing the same scene are used to obtain estimated three-dimensional data. Pinhole camera model, epipolar geometry and its recovery are discussed; common stereo triangulation methods are explained in the chapters of the thesis. Besides, feature extraction and matching topics which are used for the reconstruction process are examined. Some of the methods used in the thesis are presented by algorithmic solutions and mathematical notations. Significant advantages and disadvantages of the methods are briefly discussed and encountered problems are tried to be challenged by fundamental approaches.

Epipolar geometryStereoscopy
Faris Serdar Taşel
Çankaya University · Institute of Graduate Studies in Science
2008
00
Master'sOpen AccessEN

Vehicle number plate detection: Iraqi plates

This thesis presents a system for Iraqi vehicle number plate detection. It is one of the most interesting and challenging research topics in the past few years. The system is designed to perform detection for Iraqi vehicle number plates under any environmental conditions, it is shown that the number plates have different size and shape and also have different colors. Vehicle number plate detection is mainly used to detect an object for traffic management. There is a need for intelligent system. Vehicle number plate detection is widely used for security control, detecting speeding cars, electronic toll collection and traffic law enforcement. In Iraq the most common vehicle number plates uses white color as background and black color as character color, also there are different colors like yellow color which is used as background and black color as for character color for the trucks. In this thesis we propose a system for the detection of number plates mainly for the vehicles in Iraq. This thesis presents an approach based on simple and efficient horizontal and vertical histogram operation and Robert edge detection method. After reducing noise from the input image, we try to enhance the contrast of the binarized image. In this thesis, examples of correct and incorrect results, as well as possible, practical applications of proposed method are presented, and the system was implemented on 50 car plate, and the system has successfully detected all car plates.

Edge detection methodsExtraction
Othman Subhi Sıddık
Çankaya University · Institute of Graduate Studies in Science
2014
00
Master'sOpen AccessEN

Text categorization based on semantic similarity with word2vector

With an increase in online information, which is mostly in the form of a text document, there was a need to organize it so that management and retrieval by the search engine became easier. It is difficult to manually organize these documents, therefore, machine-learning algorithms can be used to classify and organize them. Mostly, they are faster, more accurate and less expensive than manual classification. Most traditional approaches of machine learning algorithms depend on the term frequency in determining the importance of the term within a document and neglect semantically similar words. For this reason, we proposed to build a classifier based on semantically similar words in text classification by using the Word2Vector model as a tool to compute the similarity between documents and capture the correct topic. So we built two models by applying three phases: the first phase, we applied preprocessing steps and the second phase, we created a dictionary for top ten categories of Reuters 21578 datasets and the final phase we trained Word2Vector model on the Wikipedia English dataset and use it to compute similarity v between documents. Depending on the results of our study, we found that the second model (the most similar predicted topic) is better than the first model (average based predicted topic) in all categories. When we compare the results of our study with other studies, we found that result of our study is a parallel to the results of other studies, but not overcome them, although these studies use feature selection in the improvement of their results while we use feature extraction in explaining of our results.

Ather Abdulrahem Mohammedsaed Alsamuraı
Çankaya University · Institute of Graduate Studies in Science
2017
00
Master'sOpen AccessEN

The challenges to apply electronic government to gain the passport services

It is well known that Iraq after 2003 goes to investment in e-government to provide services to citizen, and one of the important service is the passport issuance service, while the Iraqi passport is the most important document for Iraqi citizens. This study prepared to address the challenges in the process of obtaining a passport in Iraq. The importance of this study is that it involves field work, and data obtained from Iraq. In the beginning, the researcher visited the General Directorate for nationality in order to better understand the process and procedures for obtaining this document. After that, a questionnaire form was prepared and analyzed by using a statistical approach to understand the hurdles faced the citizens who interested in getting a new passport. The questionnaire form contained four questions regarding the difficulties and possible solutions, and was distributed to 219 random citizens. Briefly, the results show that there are two kinds of challenges in the process of gain the passport; the first ones are the challenges faced by the employees of the passport directorate, and the second ones are faced by the citizens who have to go through long waiting periods. In conclusion, the researcher present a summary of all the challenges faced both of the general directorate for nationality employees as well as the citizen, and he have been able to identify a set of recommendations for the implementation of electronic government, and its effective role in the elimination of all these challenges, which contributes to improve and develop Iraqi passport services. Keywords: challenges of e-government, e-government, Iraqi passport, General Directorate of nationality, statistical analysis.

Samıd Hassan Hadı Al-baıaty
Çankaya University · Institute of Graduate Studies in Science
2017
00
Master'sOpen AccessEN

Automatic scoring approach for Arabic short answers essay questions

There are different types of questions produced by the students in their exams, such as multiple-choice questions, true/false questions, and essay questions which require free text answers. Evaluation and scoring these types of exams traditionally are an exhausting process that takes from the instructors a lot of efforts, time and activities. In this regard, applying automated approaches to evaluate and score exams are essentially required to reduce time and efforts. Although there are many commercial tools for scoring multiple-choice and true/false questions, yet there is lack of approaches and tools for evaluating and scoring essay questions, especially for the Arabic language. In this research, the aim is to propose an automated scoring approach for short answers to Arabic essay questions. The scoring process is based on the similarity between the student's answer and model answer which is provided by the instructor. Cosine similarity measures will be used for this purpose. Cosine similarity is a heuristic evolutionary measure that has succeeded to solve text to text similarity problems. In this research, we will use the word root for each keyword in the student's answer and the model answer in order to achieve accurate results. The proposed approach will be tested on a data set proposed and will be compared to other approaches.

Mohammed Abdulmunem Nsaıf Al-falahı
Çankaya University · Institute of Graduate Studies in Science
2017
00
Master'sOpen AccessEN

A systematic review of interactive information retrieval evaluation studies, 2007-2016

Since the last mid-century researchers start investigate in the performance of IIR (interactive information retrieval). Many methods and measures were innovated and used in this field studies. To ensure the replication of maturation researches there are several factors that effect on the maturation of researches such as introducing standards for measurement and analysis, and understand past endeavors. In this study, we analyzed a vast range of papers within the period (2007-2016), where 1110 papers were examined manually and only 78 articles were included. Based on the achieved results in our study, we found that the researchers increased their concentration on IIR evaluation. Due this expansion of researching in this topic, we noticed that researches used some new techniques and datasets for evaluation of IIR systems. Most of the included papers were conducted based on the help of participants and questionnaires.

Haıder Alı Hasan Alyaseen
Çankaya University · Institute of Graduate Studies in Science
2017
00
Master'sOpen AccessEN

A lexicon based method for subjectivity and sentiment analysis using an Arabic twitter corpus

Sentiment analysis for social media is an interesting area of data mining for decision making in various domains. Therefore, continuous research is carried out in this area to cover the huge amount of data being pushed by users. Arabic is one of the ten important languages used in social media; therefore, interest in decision making anywhere needs knowledge about this. Twitter provides a platform for the exchange of opinions and ideas among users, leading decision making to building a knowledge base towards the development and planning of future outcomes. We present and illustrate how to obtain models with a high accuracy of classification by using the Lexicon-based approach. Our approach is implemented in three phases, beginning with preprocessing steps for Arabic words. The second phase discusses the extraction of more features relating to statistical and semantic orientations. We demonstrate how the extracted features (weight, score and negation) depend on two types of Arabic lexicon being clearly useful. Finally, the third phase applies a feature selection method with the Information Gain attribute evaluation and Ranker search method to find the features that have greater impact on the performance measures. We keep the features that have high rankings and remove those that have low rankings from the dataset. In the last two phases, we carry out our evaluations for all tasks using two machine-learning algorithms, namely K-Nearest Neighbor and Naïve Bayes. The accuracy for classification was found to have reached 93.56 with the Naïve Bayes classifier with a score feature, and this task determined which one of the two selected machine-learning models is more suitable for classifying the sentiment of Arabic tweets. Keywords: Arabic sentiment analysis, lexicon-based, feature extraction, feature selection, KNN, Naïve Bayes, Ranker, information gain attribute.

Naseer Mohammed Jasım Al-buhruzı
Çankaya University · Institute of Graduate Studies in Science
2017
00
Master'sOpen AccessEN

An evaluation of three search engines (google, yahoo, bing) based on arabic user perception

With the rapid increase of Arabic users around the world, a need of understanding the preferences of this wide Stratum of society is important. To do so we designed a questionnaire to evaluate three search engines (Google, Yahoo, and Bing) to reveal the Arabic user's perception of these three search engines. Each engine has its own features and it is own search methods to achieve the main purpose of search engines which is retrieving relevant results. This study exposes the pros and cons of the three tested search engines which give the opportunities to search engine developers to enhance their products. The results showed that Google performs very well comparing with Yahoo and Bing in some kind of data (text and video search) whereas image and advertisements Yahoo seems to be more accurate, Bing captivated many participants with its attractive homepage and useful features except supporting Arabic language in Bing map. Our results will help engineers and designers to improve their search engines in term of the design of homepages and result, improving their search engines algorithms for Arabic queries.

Suad Shattı Azeez Azeez
Çankaya University · Institute of Graduate Studies in Science
2017
00
Master'sOpen AccessEN

Improving classification of damaged buildings post hurricane using satellite imagery

There is a growing need for efficient and accurate methods for assessing building damage, especially in post-disaster scenarios. Traditional manual inspection is time-consuming and prone to human error, highlighting the need for automated systems. Leveraging advanced deep learning models can improve the accuracy and speed of image classification, contributing to timely disaster response. This study focuses on developing an advanced deep learning model for building damage classification using image data. The proposed model leverages a hybrid architecture combining ResNet50 for transfer learning with a custom Convolutional Neural Network (CNN) to capture both global and local features effectively. The dataset used includes labeled images of buildings under different conditions, providing a diverse set for training and evaluation. The evaluation results showed an accuracy of 98.9% on the balanced dataset and 98.01% on the unbalanced dataset. The proposed model outperformed various models and demonstrated robustness across different data distributions. The study provides insights into the efficacy of hybrid models combining transfer learning and custom-designed CNNs for image-based classification tasks.

Sarah Muayad Ismael Al-sumaıdaee
Çankaya University · Institute of Graduate Studies in Science
2025
00
Master'sOpen AccessEN

Hepatitis C virus prediction in machine learning

The infection of the hepatitis C virus is a considerable medical field challenge globally that can require the development of effective as well as accurate diagnostic approaches. Traditional diagnostic techniques, while widely used, often have limits when it comes to accuracy, accessibility, and cost-effectiveness. This study proposes a predictive model utilizing machine learning to early diagnose the liver HCV, utilizing the Extra Trees Classifier in conjunction with the Synthetic Minority Over-Sampling Technique to address the challenge of class imbalance within the dataset. Three freely accessible datasets, HCV-EGY, ILPD, and HCV, have been used in both training and evaluation, thereby ensuring robustness and generalisability across diverse population groups. The model of this study achieves an accuracy of 98% of both the HCV and HCV-EGY datasets, while the ILPD achieved 95%. exceeding the performance of traditional diagnostic methods and demonstrating the effectiveness of machine learning in improving early HCV detection. An analysis of feature importance was performed to determine the key biomarkers that significantly influence the classification process. The interpretability component is essential, offering insights into the biological markers linked to HCV infection, which may assist in refining diagnostic criteria and treatment strategies. This study highlights the potential of non-invasive, data-driven diagnostic methods in clinical settings through the application of advanced machine learning techniques. The results indicate that machine learning models can function as dependable, efficient, and interpretable instruments to aid healthcare professionals in the early diagnosis of HCV. This research enhances the existing evidence for AI-driven methodologies in medical diagnostics, facilitating the development of more accurate and accessible disease detection frameworks.

Alhasan Salıh Ibrahım Ibrahım
Çankaya University · Institute of Graduate Studies in Science
2025
00
Master'sOpen AccessEN

Opinion mining with text operations and extracting data from user reviews for e-commerce applications

This thesis tries to make an understanding of using opinion mining and sentiment analysis and applying these methods for extracting data from user feedbacks. Before the data extraction step, opinion mining is examined in detail to understand better the goals to be achieved.To extract the data, www.kobiform.com e-commerce web site is used as pilot platform. www.kobiform.com is furniture selling e-commerce web site originated in Ankara which will be mentioned in details later. Beside user reviews, multiple choice and text based surveys are given to the users while they browse thorough products for the data gathering.The algorithm used to evaluate the data which is based on the classic view of generating word sets. The application used to gather data is developed with .net framework and Visual Studio 2008 is used as IDE. Microsoft SQL Server is used as database.

Data analysisData mining
Muhammed Burak Uytun
Çankaya University · Institute of Graduate Studies in Science
2013
00
Master'sOpen AccessEN

Analysis of Turkish art music; identification of makam signatures

This study aims to gather more scientific-based concrete data by using computer technology and to figure out what extent the makam structure of Traditional Turkish Art Music interpreted in the compositions. For this reason, 120 compositions from Muhayyer kürdi, Acem kürdi and Kürdi makams from Traditional Turkish Art Music were analyzed by computer program in terms of makam and gained data were compared with the makam structure of Traditional Turkish Art Music. Also, makam of the composition is being tried to determine with the help of computer software. Musical note?s frequency of usage, usage duration and effectiveness level were calculated and progression analysis of the scales in the compositions were done. It was seen that data gained by research show parallelism to a large extent with the makam structure of the Traditional Turkish Art music. In addition, thanks to computer software the makam of the composition can be determined but it is seen that the definition of the progression cannot be expressed by mathematical data with the gained data.

Mehmet Bilal Er
Çankaya University · Institute of Graduate Studies in Science
2013
00
Master'sOpen AccessEN

Matching resumes with job descriptions using latent semantic indexing

In this thesis, Vector Space Model of Information Retrieval is examined. First, the classical method of term frequency inverse document frequency is presented as an introduction to the problem. After introducing basics, the thesis explains the concept of Latent Semantic Indexing. Singular Value Decomposition, which is the fundamental of Latent Semantic Indexing, is explained without going too deep into Linear Algebra. Relationship between Singular Value Decomposition and Latent Semantic Indexing is also explored. Finally, thesis presents the results of its demonstration, which is matching a Resume with an appropriate Job Description by using Latent Semantic Indexing and comparing it with the classical Vector Space method.

Murat Pojon
Çankaya University · Institute of Graduate Studies in Science
2014
00
Master'sOpen AccessEN

Apache Nutch ve Lucene kullanarak web tarama

The availability of information in large quantities on the Web makes it difficult for user selects resources about their information needs. The good link between the internet users and this information is Search engine. Search engine is kind of Information Retrieval (IR). It works on data collection from the Web by software program is called crawler, bot or spider. Most of Search Engines users don't know the mechanism of action the Search Engine, like how Search Engine works and how it catch information in the Web and how it rank the results to users. For this reason in this thesis used the open-source Search Engine is researched in detail. In this study, we used each of (Apache Nutch and Lucene) to clarify work of Web crawling open source. They are released under the Apache Software Foundation. Nutch is a web Search Engine working to search and index Web Pages from the World Wide Web (WWW). Nutch is based or built on top of Lucene. It uses in the information retrieval technology. It has more software libraries to indexing of large-size data. Lucene doesn't care about information existing in the Web, like PDF, TEXT, and MS Word. It is working to indexing these documents and convert them to the data can be utilized. The benefit of using both Nutch and Lucene in this study, they are free and we can their development. The Nutch and Lucene are written by Java language, it is a computer programming language. Furthermore, we used Tag Cloud Technology to analysis and view the Lucene content or its index.

Computer networksDigital image processingDatabase+2
Nıbras Abdulwahıd
Çankaya University · Institute of Graduate Studies in Science
2014
00

Other supervisors