Yüksek LisansAçık Erişim

Enhancing mobile application age ratings usingnatural language processing techniques

2023
0 görüntülenme
0 i̇ndirme
Danışman: Dr. Öğr. Üyesi Muhammed Fatih Adak

Özet (EN)

As the huge number of mobile applications in stores increases, it gets difficult to verify all the information about each application. Especially Age Rating. For instance, Play Store and App Store have an algorithm to determine the age rating of an application, the algorithm is applied after the developer is asked questions about the application such as if the application includes any violence scenes and if It is frequent or not or containing offensive language or crude humor and many other questions, also rules of each country are put into account with the developer's answers while giving the appropriate rating. The challenge here is that a developer may answer any of the questions wrongly and this would affect the process of the rating of the application negatively, either by making it accessible by users who must not use it because of age restriction or limiting users from reaching the application while the application is suitable for them. In this thesis, a novel method is presented to classify mobile apps by analyzing their app store descriptions. The study utilizes a dataset obtained from the Apple App Store, consisting of over 365,000 app descriptions. The research aims to accurately classify and age-rate apps, enabling users to find apps suitable for their age group and preferences, and helping developers identify direct competitors and market trends. The approach consists of three major steps: data pre-processing, vectorization (using word embeddings and Bag-of-Words), and classification. Text pre-processing techniques such as lowercasing, tokenization, removal of non-ASCII characters, numbers, URLs, and stop-word removal are applied. Two vectorization methods are used: Bag-of-Words with a maximum of 1000 most-frequent words and word embeddings (GloVe, Word2Vec, fastText). The generated word embeddings represent the entire app description using simple unweighted averaging. Classification is performed using various algorithms such as Support Vector Machines, Multi-layer Perceptron, Random Forests, Nearest Neighbor, AdaBoost, Decision Trees and Logistic Regression. The performance of these classifiers is evaluated using accuracy, recall, precision, and F-measure. The study compares word embeddings with topic modeling (LDA) and demonstrates that word embeddings outperform LDA in app classification tasks. The limitations of LDA with short text inputs are highlighted, emphasizing the advantage of word embeddings in capturing semantic meaning and achieving better classification performance. Furthermore, the study compares word embeddings with the bag-of-words approach (VSM) and shows that VSM outperform word embeddings in capturing the semantic relationships and similarities between words in app descriptions. Among the word embedding models (GloVe, Word2Vec, fastText), GloVe performed the best, potentially due to its more comprehensive vocabulary. The inclusion of character n-grams in fastText did not significantly contribute to classification accuracy. Overall, the proposed approach has practical implications for users and developers, enabling better app discovery and market insights. Future directions include conducting more extensive analysis with expert-generated categorizations and developing a working prototype to bring the research closer to real-world application. In summary, the thesis presents an effective approach for app classification based on app store descriptions, utilizing word embeddings, and achieving accurate age rating predictions.

Yazar

Dr. Muhammed Emin İdelbi

Bu Yayına Nasıl Atıf Yapılır

Muhammed Emin İdelbi (Master Thesis). Enhancing mobile application age ratings usingnatural language processing techniques, 2023, Sakarya University.

Anahtar Kelimeler

Lisans

Tüm Hakları Saklıdır

Bu eser belirtilen lisans koşulları altında paylaşılmaktadır.

Sakarya University tezlerinden daha fazlası