Master'sOpen Access

Adult Content Filtering Using Text and Image Analysis

2012
0 views
0 downloads

Abstract (EN)

ABSTRACT: The working principle of the Internet is such that anyone who sets up a server computer and connects it to the local area network in their neighborhood becomes equipped to shear with the world any type of information they deem appropriate. Generally, some of this information dispatched is not appropriate for viewing of our children and some steps should be taken to help the society so that classification and controlled access become possible. Throughout this thesis, we designed and implemented a text and image based web-page filtering system that makes use of web page parsing, HTML tags removal and string in string search procedures along with various other criteria for processing images downloaded from a web site using a custom written JAVA program. For the text, there are some words and phrases that are common to pornographic sites and are rarely seen in regular sites. To find out such words and phrases, a survey was done on a number of sites. With the words and phrases determined, our expectation is that any site which may contain pornographic oriented text will have in it some of these words and phrases. Hence, once the tested web page was parsed and the pure text string was obtained from the downloaded HTML code the string would be searched for the type of words and phrases previously determined and final decision would be made based on the frequency of words detected. From literature survey, everyone seems to agree that pornographic images have too much skin exposure which is why detecting skin is generally the starting point. To find out the amount of skin in an image, improved YCbCr color segmentation was implemented. The improved YCbCr segmentation would satisfactorily segment out the skin from the other regions but some skin like objects would still be falsely detected. Therefore, texture property was used to differentiate bearing in mind that skin is generally smooth and most others textures aren’t (many are more coarse). In order to classify a web site from which images have been extracted through the help of a JAVA program, criteria such as face detection, lacunarity, edge sum, uniformity, entropy and percentage of skin region have been employed and when three or more of the criteria were met this was taken as an indication for containing adult nature material. Final decision was made by computing percentages for the results obtained for both the text and image analysis and comparing the average of the two to some previously selected threshold ranges. For the five randomly selected adult content containing web sites that were used for test purposes the text analysis always gave 95-100% accuracy and the image analysis resulted in 56.83, 54.83, 52.63, 57.14, 66.67 % accuracy respectively for sites 1-5 as detailed in chapter 5. In chapter five it was also shown how the two results (text and image analysis) can be combined to get an average percentage. For the five different web sites considered the lowest average percentage obtained was 73.82%. Keywords: HTML parsing, skin color segmentation, texture analysis, lacunarity. ………………………………………………………………………………………………………………………………………………………………………………

Author

Dr. Halidu Sule

How to Cite

Halidu Sule (Master Thesis). Adult Content Filtering Using Text and Image Analysis, 2012, Eastern Mediterranean University, Department of Electrical and Electronic Engineering.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Eastern Mediterranean University