Content based image search engine
2015
0 views
0 downloads
Advisor: Prof. Dr. Abdullah Bal
Abstract (EN)
Nowadays, since the amount of the information is constantly increasing, reaching accurate and clear data is gaining a great importance. Information has an enormous strength and people who attain it first move one step further at all levels. People are searching everything on the internet, and in their computers, e-mails, as well as photograph archives almost every day. Hence, searching and reaching the information became an everyday routine. People are able to learn almost anything which they are curious about. For instance, 5.8 billion words were searched per day using Google in 2014. However scientists argue that not only Google damages the minds but also causes laziness. Since there is enormous data and information out there, it is very hard to process different types of information and create a value. In addition, a requested analysis may take long time periods such as weeks or even months, when storing the information is added to these difficulties. In order to read and process quickly big amount of data, Big Data technologies are developed. In this thesis, to manage content-based image search engine, we built a potential architectural design using Big Data technologies. Data set is prepared from news websites on the internet. Headline news of 20 news websites are scanned twice a day for a week. Images are stored in Apache HBASE database. By doing so, a dataset containing 200 thousand images is prepared. The features of images are extracted by using MapReduce framework with the help of image processing algorithms and then extracted features are indexed by using Apache Solr. Increasing the number of pictures will also increase the response time. To solve this problem we designed cascading search architecture by using different image processing algorithms. In cascading search architecture, images which are irrelevant to queried image inside the data set are filtered and then ranked with the help of more prosperous image processing algorithms. At this point we choose Fuzzy Color Texture (FCTH) algorithm which has less response time then other image processing algorithms compared (FREAK, SIFT, SURF, BRIEF, BRISK, ORB). By doing so the number of the pictures is decreased from 200 thousand to 500, approximately in 0.2 seconds. After filtering stage, 500 images are ranked according to similarity of the queried image with Fast Retina Keypoints (FREAK) and Scale Invariant Feature Transform (SIFT) algorithms respectively. Finally, the images having different types of texture but having similarity are ordered again among all the peers with the help of Color Mapping algorithms. The similary between query and result images decreases according to the descending order of result images. The performance of systems is measered by recall and precision values. Recall measured as 92 % without treshold usage. To measure precision, 6 different tresholds are used. We used treshold values ascending order and relaize that recall value decreases if precision is increasing.
Author
Dr. Mehmet Zahid Yüzügüldü
Institution
How to Cite
Mehmet Zahid Yüzügüldü (Master Thesis). Content based image search engine, 2015, Yıldız Technical University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Yıldız Technical University
- Examining ?Historical housing structures" within the confines of protecting ecological balance(2012)
- Approximate solutions of integral equations(2012)
- Stepper motor speed control with labVIEW(2014)
- Determining supply chain risk factors in food industry(2014)
- TiO2/Cu2O ince film fotovoltaik hücrelerin karakterizasyonu(2014)
- Study of the problem of evil from a philosophical perspective(2015)