Social media analysis and user representation with deep learning methods
2021
0 views
0 downloads
Advisor: Doç. Dr. Galip Aydın
Abstract (EN)
In today's world, where a significant part of life such as education, shopping and entertainment has shifted to the online world, huge amounts of data are being produced at every moment. It is very likely that besides the representation of simple text, sound and images, studies on semantic representation of things such as people (social media profile, author, educator, student, etc.), products (clothes, books, movies, songs) and many other objects in the online world will find an important place in the near future of machine learning. In this thesis, we propose a method for obtaining profile representations of Twitter users that will find a wide range of use cases in social media analysis applications. Our method, based on state-of-the-art deep learning-based text representation methods, can learn from both structured and unstructured user data in an unsupervised manner. We created a one-of-a-kind social media dataset to implement and evaluate the models. The dataset consists of rich user information such as profile information (biography, location, etc.), shares (retweets), comments and likes from semantically grouped users. The proposed user representation models were compared in detail for the different scenarios of exploring different text representation techniques such as tf-idf, word2vec, doc2vec, ELMO, BERT. In particular, a comprehensive parameter review has been performed for distributed document representation models. In addition to obtaining user representations, another contribution has been made by proposing approaches to two important problems encountered in text classification and sentiment analysis applications using deep learning methods on the text content created by social media and other Internet users. In the sentiment analysis problem, a Turkish dataset was created from multiple sources and a comprehensive study was conducted using the oversampling technique to overcome the imbalanced class distribution problem that limits the success of deep learning models. In the case of the text classification problem, a domain adaptation approach is proposed that can be applied to deep learning models that classify with high accuracy under conditions with a small amount of labelled tweets. For use in the study, a Turkish tweet dataset was created and labelled for 5 classes. For both studies, a comprehensive investigation of the proposed approaches was conducted using classical deep learning methods such as CNN, LSTM and hybrid architectures where they are used together.
Author
İbrahim Rıza Hallaç
Institution
How to Cite
İbrahim Rıza Hallaç (Doctorate thesis). Social media analysis and user representation with deep learning methods, 2021, Fırat University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Fırat University
- Using social media as an integrated marketing communication tool(2018)
- Foundation of Dutch East İndia Company and her rising in İndonesia in the 17th century(2013)
- Examination of stress state between Doğanyol (Malatya) and Çelikhan (Adıyaman) on the east Anatolian fault zone(2020)
- Color usage at Turkish Divan of Fuzûlî(2013)
- Yavuzeli (Gaziantep) surrounding volcanic outcropping of rocks petrographic and geochemical features(2014)
- Hizbu?t-Tahrir and the religions and political thoughts of Ercumend Özkan(2008)
