Theses supervised by Doç. Dr. Arzucan Özgür Türkmen

15 theses · Boğaziçi University

Master'sOpen AccessEN

Targeted sentiment analysis on Turkish texts

Sentiment Analysis (SA) is one of the Natural Language Processing (NLP) tasks whose goal is to understand subjective information from a piece of text. The increased accessibility to the Internet and thus to social media leads people to create an enormous amount of textual data. This data can store valuable information that waits to be extracted. Targeted Sentiment Analysis (TSA) specifically aims to extract sentiment towards a particular target from a given text. Sentiment analysis in English texts is a well-studied area and mainly requires human-annotated data for training. For languages such as Turkish, there is a lack of such annotated data. In the context of this study, we introduce an annotated dataset that consists of Twitter data in Turkish. It contains almost 4K sentences that are labeled for both sentiment analysis and targeted sentiment analysis. The proposed dataset allows us to train a TSA model on Turkish texts. We propose BERT-based models with different architectures, one of which is to be used as our baseline for TSA and the others are to improve this baseline. We observe that the performance of conventional SA models degrades when used for TSA data. We investigate the performance of several BERT-based architectures for this task. Our best performing model with target markers and max-pooling layer outperforms the F1-score of conventional BERT-based SA models by 13%.

Natural language processingSentiment analysisCellular neural networks+2
Mustafa Melih Mutlu
Boğaziçi University · Institute of Graduate Studies in Science
2022
00
Master'sOpen AccessEN

Predicting intracellular functions of proteins from amino acid sequences using language processing methods

Rapidly increasing computational power and sequencing technologies, which are at the peak of their development, enable the use of advanced algorithms with high processing volume to predict the intracellular functions of proteins, which is one of the most important problems in computational biology. The functionalities of proteins emerge primarily through their three-dimensional folded structures. When these structures are interpreted as graphs, the application of graph neural networks leads to promising results. However, these approaches are limited as the three-dimensional folded structures are not yet known for most proteins. The fact that the amino acid sequences of proteins have properties similar to natural languages and the large amounts of sequence data suggest that these sequences can be processed using natural language processing (NLP) methods. In this thesis, two different NLP methods are adapted to the problem of protein function prediction, assuming that the protein sequence data contain necessary and sufficient information to predict both three-dimensional folded structure and intracellular function: (i) Bidirectional Transoformer BERT model (ii) Heterogeneous Graph Convolutional Network (GCN) model. The results show that it is more advantageous to treat the proteins as graphs. The GCN model performs better than the BERT model and achieves performance close to the state-of-the-art model that uses three-dimensional folding information. In addition, we find that tokenizing the sequences instead of using the individual amino acids as tokens increases the performance.

Amino acidsNatural language processingTransformers+1
Bedirhan Çaldır
Boğaziçi University · Institute of Graduate Studies in Science
2022
00
Master'sOpen AccessEN

Drug-target affinity prediction using a graph-based approach enriched with molecule words

Wet-lab experiments to predict the affinity of drugs for their targets are costly and time consuming. Computational methods can provide an alternative to early stage experiments and guide the research process. Recently, the use of natural language processing techniques to represent molecules has become popular and has led to successful results. In our work, we assume that proteins and ligands, like human languages, have their own languages and that these languages consist of meaningful smaller parts that we call words. We identify protein and ligand words based on their 1D sequences using a subword tokenization method and represent protein-ligand interactions with a heterogeneous graph consisting of four different node types corresponding to proteins, ligands, protein words, and ligand words. A graph-based approach is used to learn embeddings for the nodes in the graph. These embeddings are fed into a deep learning model for predicting protein-ligand binding affinity. We show that using their word embeddings to represent novel proteins and/or ligands not present in the training set improves the results compared to the case where no words are used. Using pre-trained word embeddings for previously unknown molecules is also efficient in terms of complexity, as we do not need to re-train the input graph to learn the embeddings for these new molecules.

Cansu Damla Yılmaz
Boğaziçi University · Institute of Graduate Studies in Science
2023
00
Master'sOpen AccessEN

A SEQ2SEQ transformer model for Turkish spelling correction

Natural language processing (NLP) is a fascinating area of artificial intelligence. It allows humans to interact with machines through natural language. There are two main concepts in NLP model architectures, namely input vectorization and contextual representation. The input vectorization process starts with tokenization, where there are three approaches: character-level, word-level, and subword-level. Word-level tok- enization results in a large vocabulary, and in agglutinative languages such as Turkish, words derived from the same stem are treated as different words. This makes it difficult for NLP models to understand their relationships and the meaning of the morphological affixes. Furthermore, all NLP models suffer from a common problem: spelling errors in the data. In case of spelling errors, the misspelled tokens become completely different and the models cannot understand them. In this thesis, a character-level seq2seq trans- former model is developed for spelling error correction. To train the model, a dataset for Turkish spelling correction is created by collecting correctly spelled Turkish sen- tences and systematically adding spelling errors to them. Seq2seq models suffer from multiple decoding iterations and have high prediction time. To address this problem, a novel model architecture, one-step seq2seq transformer model, is proposed in which the transformer model predicts the outputs in one iteration. The proposed models are tested with the exact match criteria. The standard seq2seq model and the one-step seq2seq model achieved 68.64% and 42.69% accuracy, respectively. Finally, the stan- dard seq2seq model makes predictions for 160 input characters in 8.47 seconds, while the one-step seq2seq model makes predictions for the same number of characters in 73 milliseconds on CPU and 28 milliseconds on GPU.

Deep learningNatural language processingArtificial intelligence+1
Şahin Batmaz
Boğaziçi University · Institute of Graduate Studies in Science
2022
00
Master'sOpen AccessEN

Prediction of druggability properties of chemicals using machine learning techniques

Drug discovery is the process of designing and developing new medicine. The research and clinical experiments for new drug proposals are costly and take a long time. Although there has been proposed a lot of drug-like molecules, the number of drugs that are confirmed by regulating bodies and released to the market is very low. That is because most of the drug candidate molecules have low pharmacokinetic properties. Therefore, early assessments of ADMET properties have gained extreme importance for pharmaceutical industry, to be able to avoid costly failures. Here, our aim is to come up with an approach that reliably predicts druggability features of drug candidate molecules as well as to point out the relations between ADMET properties and molecular descriptors. \par In this thesis study, we examine and compare 4 different molecule representations to predict druggability features of molecules, using 3 different machine learning algorithms; namely k-nearest neighbor, support vector machine classifier and random forest on 9 ADMET property datasets. We conclude that among all molecular representations, morgan fingerprint performs better in terms of accuracy and F-measure, however run time for parameter tuning and train is longer with fingerprint representations. As far as the machine learning algorithms are concerned, SVM classifier with morgan fingerprint performs better with higher accuracy and F-measure. With descriptor vector representation, we examine a set of molecular descriptors and using RF classifier, we evaluate most effective molecular descriptor for each ADMET property. For some datasets, we add the most contributive descriptor to morgan fingerprint representation and report an increase on evaluation metrics.

Gamze Ege Kahya
Boğaziçi University · Institute of Graduate Studies in Science
2020
00
Master'sOpen AccessEN

Analyzing the generalizability of deep contextualized language representations for text classification

This study evaluates the robustness of two state-of-the-art deep contextual language representations, ELMo and DistilBERT, on supervised learning of binary protest news classification and sentiment analysis of product reviews. A ``cross-context'' setting is enabled using test sets that are distinct from the training data. Specifically, in the news classification task, the models are developed on local news from India and tested on the local news from China. In the sentiment analysis task, the models are trained on movie reviews and tested on customer reviews. This comparison is aimed at exploring the limits of the representative power of today's Natural Language Processing systems on the path to the systems that are generalizable to real-life scenarios. The models are fine-tuned and fed into a Feed-Forward Neural Network and a Bidirectional Long Short Term Memory network. Multinomial Naive Bayes and Linear Support Vector Machine are used as traditional baselines. The results show that, in binary text classification, DistilBERT is significantly better than ELMo on generalizing to the cross-context setting. ELMo is observed to be significantly more robust to the cross-context test data than both baselines. On the other hand, the baselines performed comparably well to ELMo when the training and test data are subsets of the same corpus (no cross-context). DistilBERT is also found to be 30\% smaller and 83\% faster than ELMo. The results suggest that DistilBERT can transfer generic semantic knowledge to other domains better than ELMo. DistilBERT is also favorable in incorporating into real-life systems for it requires a smaller computational training budget. When generalization is not the utmost preference and test domain is similar to the training domain, the traditional ML algorithms can still be considered as more economic alternatives to deep language representations.

Berfu Büyüköz
Boğaziçi University · Institute of Graduate Studies in Science
2020
00
Master'sOpen AccessEN

Türkçe metinlerde nefret söylemi tespiti

It is well known that prejudiced and discriminatory language is being widely used and spread through several channels such as printed or social media. The discriminatory language, in particular hate speech as its more aggressive, degrading and openly targeting form, which poses a threat to the values of democracy and human rights is a global problem that needs an immediate solution. Since we find the detection of hate speech important in the fight against hate speech, we have developed a model to detect it. For this purpose, we created a dataset by retrieving printed media news that the Hrant Dink Foundation systematically annotated in the context of hate speech from the website of the PRNet media monitoring company. To the best of our knowledge, with this study, the first model developed for Turkish language that runs on a labeled dataset is produced. In particular, the fact that most of the hate speech in printed media is based on context and implications requires a system that can detect changing discursive cues and understand the context around these discourses. With different word representations, we have examined the Hierarchical Attention Network (HAN) model, which aims to capture the changing meanings of expressions by using the hierarchical structure of the text. We studied the compatibility of our model with the problem by comparing it with Convolution Neural Network (CNN), which provided important results in text processing, and with machine learning models. In order to improve our study, we developed linguistic features for the problem based on critical discourse analysis techniques. We enhanced the HAN model using these features. Our results show that performance increases with a set of features that point out the use of 'othering language'. We believe that these feature sets created for the Turkish language will encourage new studies in the quantitative analysis of hate speech.

Zehra Melce Hüsünbeyi
Boğaziçi University · Institute of Graduate Studies in Science
2020
00
Master'sOpen AccessEN

Mention extraction and normalization using ontologies in the biomedical domain

This thesis proposes a machine learning- and rule-based system for the identification of adverse drug reaction (ADR) entity mentions in the text of drug labels and their normalization through the MedDRA dictionary. The machine learning approach is based on a recently proposed deep learning model that works on the sentence level. The model makes use of the combination of the pre-trained word embeddings and Convolutional Neural Network (CNN) embeddings generated from the characters of a given token. These tokens are initially passed through bi-directional Long Short-Term Memory (Bi-LSTM) layers for feature extraction. Finally, a Conditional Random Fields (CRF) classifier is trained on those extracted features for the prediction of the target mentions. The rule-based approach, used for normalizing the identified ADR mentions to MedDRA terms, is based on an extension of the text-mining system called SciMiner. The proposed system is evaluated with the TAC-ADR 2017 challenge dataset. Since this dataset contains mentions that are disjoint and overlapping, the model also uses a recently proposed chunking scheme designed to handle those types. The model obtained 76.97 f-score performance on the TAC dataset. Some of the challenges for the worse performance compared to performance of the models trained on the generic newspaper text are the small size of the training dataset and the uneven distribution of the class instances.

Mert Tiftikci
Boğaziçi University · Institute of Graduate Studies in Science
2019
00
Master'sOpen AccessEN

Hit song prediction using feature-based machine learning

Music industry is making big investments every year to produce hit songs. The increasing number of songs available through digital platforms can enable the development of learning models for predicting hit songs and identifying their common features. This thesis investigates classifying a song as hit or non-hit by using various machine learning methods. Besides the basic musical features provided by Spotify, more complex features based on the chords and melody extracted from the music files by utilizing music theory information are designed. Chord based features are created using the important chord progressions based on tonal harmony, while the features based on melody are designed in an intuitive way. In addition, new benchmark datasets are created by using both hit and non-hit songs from dance and rock music genres. The results show that using chord and melody based features with the basic musical features may lead to an improvement in hit song prediction performance. For rock songs, the Random Forest classifier achieves a significant improvement on the results by using these features. It is also observed that using a specific feature combination with Support Vector Machine classifier increases the accuracy score of hit dance song prediction. Furthermore, all the features used in this study are analyzed in the last part of this study for each dataset.

Machine learningMachine learning methods
Anıl Orhan Çalışkol
Boğaziçi University · Institute of Graduate Studies in Science
2019
00
Master'sOpen AccessEN

Identifying event nuggets in turkish news texts using natural language processing and machine learning methods

Event nuggets are smallest textual instance that marks the existence of an event. Detecting event nuggets in a given text opens door to further research and many practical applications, therefore it has been studied extensively for some lan- guages including English, Spanish and Chinese. In this study, event nugget detection and event type classification for Turkish is studied for the first time. Due to lack of annotated data for event nugget detection in Turkish, we developed a new annotated dataset for this task. In this study we described how we manually annotated our dataset as well as our system to identify event nuggets in Turkish news texts. The dataset consists of words from Turkish news texts. Each word in the dataset is manually annotated in terms of sequence type, nugget type, realis value and whether the event nugget is the main event, thus enabling us to make analysis on this dataset for event nugget detection, event type classification, realis classification and main event detection. We made use of language specific features like morphological features and dependency parser features in Turkish as well as some other features. We aimed to see the effect of language specific features on this kind of analysis. We also experimented with different machine learning algorithms to find the best fitting model for our tasks. After having completed our experiments, we have shown that Turkish specific morphological features, dependency tree related features as well as word embeddings enabled us to achieve better results.

Mehmet Durna
Boğaziçi University · Institute of Graduate Studies in Science
2019
00
Master'sOpen AccessEN

Using transformer networks for detection andnormalization of named entities in biomedical texts

The increasing difficulty of retrieving relevant information from rapidly growingliterature has raised the interest for natural language processing (NLP) systems in thebiomedical domain. In many of these systems, detection of named entities such asdiseases, genes, and molecules (named entity recognition) and matching them to thecorresponding entries in ontologies (normalization) are important intermediate steps.As these two tasks are related and datasets in this domain are relatively small, multi-task learning has been frequently used in the literature for this problem. Meanwhile,in recent years, the success of transformer-based pre-trained language models suchas BERT in various NLP tasks has led them to be also applied in the biomedicaldomain. The different characteristics of biomedical text such as abbreviations andspecific terminology motivated the development of new language models, which weretrained specifically for this domain using a biomedical corpus. In this study, we proposea multi-task learning approach for named entity recognition and normalization byutilizing transformer-based pre-trained language models. To enable the optimal sharingof information, both tasks are formulated with text span embeddings obtained witha common encoder network. Promising results are obtained and compared with theresults of state-of-the-art systems from the literature for commonly used named entityrecognition datasets.

İlkay Ramazan Pala
Boğaziçi University · Institute of Graduate Studies in Science
2021
00
Master'sOpen AccessEN

Machine learning based language models on nucleotide sequences of human genes

The use of computers for different fields of science has provided tremendous benefits. This phenomenon is expected to be more common as the speed of computers and the amount of data available for different kinds of scientific problems increase. This study focuses on genomics, one of the most exciting areas of science. We have applied several techniques to obtain a model for nucleotide sequences of genes that are found in human beings so that the model can learn the general pattern in these nucleotide sequences and predict how likely it is that an unseen sequence is a gene that belongs to human beings. They can even generate new nucleotide sequences. All of the methods used are examples of machine learning, where the programs are designed to learn from data for a specific task, rather than explicitly programming what to do at each step. Traditional approaches such as N-grams and more recent deep learning-based techniques such as recurrent neural networks and transformer architecture language models are used. In addition to the classical metrics, the strength of the methods is measured using a real-world task from the field of genomics. Finally, the results show an interesting comparison of how all these models perform on a task that is inherently different from classical natural language processing tasks, and how sometimes simple models like N-grams can be as good as, if not better than, more sophisticated techniques such as transformer for solving certain types of problems. Furthermore, the significance of evaluating obtained models on real-life tasks is seen because the transformer model was superior to the N-gram model according to perplexity, although it performed worse on real-world task.

Musa Nuri İhtiyar
Boğaziçi University · Institute of Graduate Studies in Science
2023
00
Master'sOpen AccessEN

Prediction of pathogen-host interactions with protein sequence embeddings using deep learning

Infections caused by pathogens are a significant problem around the world. Determining protein interactions between pathogens and hosts is critical to understanding infection mechanisms and developing prevention and treatment strategies. Wet-lab experiments to identify protein interactions are expensive and time-consuming. Therefore, computational approaches have been proposed as a promising complementary solution. While 3D structures of proteins contain helpful information about protein functions, with advances in sequencing technology, 1D sequences of proteins are widely available and are often utilized because they are easier to process with less computational power. The main goal of this thesis is to develop a sequence-based approach for predicting pathogen-host protein interactions based on the hypothesis that protein sequences can be viewed as sentences, therefore, can be decomposed into chunks, which we refer to as protein words. We first adapt the Byte Pair Encoding (BPE) tokenization method from the field of natural language processing to protein sequences and then apply a graph-based approach using the Metapath2Vec algorithm to learn representations of sequences. The results show that incorporating a word-based representation of proteins improves the performance of the graph-based approach. In addition, two other methods for learning text representations, SeqVec and ProtBERT, are evaluated for predicting pathogen-host protein interactions. The results on three virus-host protein interaction datasets show that the sequence-based protein representation approaches are promising and achieve comparable performance to the state-of-the-art methods.

Büşra Oğuzoğlu
Boğaziçi University · Institute of Graduate Studies in Science
2023
00
Master'sOpen AccessEN

Hate speech detection in turkish news using a transformer-based model enhanced with linguistic features

Hate speech directed at ethnicities, nationalities, religious identities, and specific groups has increased not only in social media, but also in print media. This creates a need for automated hate speech detection systems that can quickly review and filter print media content before it is provided to readers if it contains hate speech. However, most of the existing automatic hate speech detection models are limited to detecting hate speech without considering the hate speech target group-specific discourse that is often used in news articles. Moreover, there are few datasets that include Turkish print media articles in the hate speech domain. In this study, a new BERT based model enriched with a set of target-oriented linguistic features for hate speech detection is proposed. The effects of weighting different BERT hidden vectors are also investigated, instead of using only the first hidden vector of the BERT encoder, which is the classical approach. New BERT based models that integrate different attention techniques are proposed for combining hidden vectors. A new preprocessed Turkish dataset for hate speech is also published, in which the target group for all hate speech articles is annotated. Experiments on a comprehensive Turkish dataset of news articles labeled for hate speech show that competitive performance in terms of accuracy and F1-score is achieved compared to previous approaches.

Deep learningNatural language processingHate speech+1
Atıf Emre Yüksel
Boğaziçi University · Institute of Graduate Studies in Science
2022
00
DoctorateOpen AccessEN

Utilizing weakly-supervised learning for hashtag segmentation and named entity disambiguation

Today's high-performing machine learning algorithms learn to predict by the supervision of large amounts of human-labeled data. However, the labeling process is costly in terms of time and effort. In this thesis, we design weakly-supervised approaches, which are based on automatically labeling raw data, for two different Natural Language Processing (NLP) tasks, namely hashtag segmentation and Named Entity Disambiguation (NED). Hashtag segmentation's aim is to identify the words in the hashtags, so as to process and understand them better. We propose a heuristic to obtain automatically segmented hashtags using a large tweet corpus and use these data to train a maximum entropy classifier. State-of-the-art accuracy is achieved for hashtag segmentation without using any manually labeled training data. The target of NED, which is the second task that we address, is to link the named entity (NE) mentions in text to their corresponding records in the Knowledge Base. We hypothesize that the types of the NE mentions may provide useful clues for their correct disambiguation. The standard approaches for identifying mention types require a type taxonomy and large amounts of mentions annotated with their types. We propose a cluster-based mention typing approach, which does not require a type taxonomy or labeled mentions. This weakly-supervised approach is based on clustering the NEs in Wikipedia by using different levels of contextual information and automatically generating data for training a mention typing model. The mention type predictions lead to significant F-score improvement when incorporated to a supervised NED model. This thesis shows that designing weakly-supervised approaches by considering the underlying characteristics of the addressed problem can be an effective strategy for NLP.

Arda Çelebi
Boğaziçi University · Institute of Graduate Studies in Science
2020
00

Other supervisors