Theses supervised by Prof. Dr. Lale Akarun Ersoy

10 theses · Boğaziçi University

Master'sOpen AccessEN

Unsupervised routing strategies for conditional deep neural networks

Deep convolutional neural networks are considered state-of-the-art solutions due to their high classification performance in image classification tasks. The apparent drawback is the amount of computing power required to process a single input. To deal with this, this thesis proposes a conditional computation method that learns to process an input using only a subset of the network's computation units. Learning to execute only a part of a deep neural network by routing individual samples has several advantages. Firstly, it is beneficial to lower the computational burden. Furthermore, if images with similar semantic features are routed to the same path, that part of the network learns to discriminate finer differences among this subset of classes, resulting in improved classification accuracy with fewer parameters and computational resources. Investigating the network's activation on a single sample can also help interpret the neural network's prediction. Several works have recently exploited this idea using tree-shaped networks or taking a particular child of a node and skipping parts of a network. In this thesis, we follow a trellis-based approach for generating specific execution paths in a deep neural network. We have also designed a routing mechanism that uses unsupervised differentiable information gain-based cost functions to determine which subset of units in a layer block will be executed for a sample. We call our method Conditional Unsupervised Information Gain Trellis (CUTE). We tested the clustering performance of our unsupervised information gain-based objective function under different scenarios. Finally, we tested the classification performance of our trellis-shaped CUTE network on the Fashion MNIST dataset. We show that our conditional execution mechanism achieves comparable or better model performance than unconditional baselines, using only a fraction of the computational resources.

Convolutional neural networksNerve netRouting
Tuna Han Salih Meral
Boğaziçi University · Institute of Graduate Studies in Science
2022
00
Master'sOpen AccessEN

Blur assessment in sign language videos

Blur impairs the sharpness of visual features and the clarity of details. It may sometimes be desired for artistic effect. However, in general, it is regarded as a defect. There are different problems studied about blur, such as blur detection, segmentation, estimation, and deblurring, but despite its abundance in visual media such as photographs and videos, there is limited annotated data about blur. This lack of data inhibits the usage of deep learning models because they require a lot of annotated data. Annotating that much data is expensive and cumbersome. In this thesis, we investigate blur-vs-sharp classification using deep learning, also we experiment with weak supervision as a remedy against the lack of data for blur assessment and localization. We compare our results with the classical approaches found in the literature. We use the data we annotated from four different datasets, three of which are sign language datasets and the other one is an action recognition dataset. We focus our research on sign language videos where motion blur is frequently encountered. Sign languages are the primary communication method of Deaf community and for that reason sign language recognition (SLR) is an important task. Determining the intensity of blur and its location may be beneficial for SLR research.

TurbidityImage processingSign language
Giray Naim Eryılmaz
Boğaziçi University · Institute of Graduate Studies in Science
2022
00
Master'sOpen AccessEN

Driver behavior modeling

Autonomous vehicles are set to be a part of everyday traffic. Their presence in traffic dominated by humans possesses some challenges. Any experience with driving in traffic shows us that each driver is unique in their driving style. So far, this richness in differences in human behavior has not been projected into the models used in traffic simulations. These models are an essential part of the development of autonomous vehicles; from the inference of other vehicle intentions to virtual testing. Therefore creating a more realistic traffic environment is a very important task. In this work, a deep dive into the state of the problem is given. Then, a framework that accounts for different driving styles, as well as different vehicle types, is introduced. Firstly, an in-depth analysis of distinct patterns of driving is carried out in the dataset. Then these distinct patterns are modeled with simulated agents using reinforcement learning. In inference time, a traffic scene is observed, each vehicle is assigned to the pre-trained driver model and a simulation is carried out. As a result, a traffic scene is reconstructed with data-validated models. This new approach that incorporates previous driver modeling work with a behavioral component, paves the way for a more realistic model of the traffic. This realistic traffic model can be used in AV testing and validation.

Driver behaviorsDriversUnmanned vehicles
Ferhat Melih Dal
Boğaziçi University · Institute of Graduate Studies in Science
2022
00
Master'sOpen AccessEN

Attention modeling with temporal shift in sign language recognition

Sign languages (SLs) are the main communication language of deaf people. They are visual languages that establish communication through multiple cues including hand gestures, upper-body movements and facial expressions. Sign language recognition (SLR) models have the potential to ease communication between hearing and deaf people. Advancements in deep learning and the increased availability of public datasets have led more researchers to study SLR. These advancements shifted solution methods for SLR from hand-crafted features to 2 Dimensional Convolutional Neural Network (2D CNN) models. Inadequacy of 2D CNNs on temporal modeling and 3D CNNs' ability of spatio-temporal modeling made 3D CNNs a popular choice. Despite its successful results, high computational costs and memory requirements of 3D CNNs created a need for alternative architectures. In this thesis, we propose an SLR model that uses 2D CNN as backbone and attention modeling with temporal shift. Usage of 2D CNN decreases the number of parameters and required memory size compared to its 3D CNN counterpart. In order to increase adaptability to other datasets and simplify the training process our model uses full frame RGB images instead of cropped images that focus on specific body parts of signers. Since communication in SL is established by using multiple visual cues at the same time or at different moments, the model must learn how these cues are collaborating with each other. While temporal shift modules give our 2D CNN backbone model the ability of temporal modeling, attention modules learn to focus on what, where and when in videos. We tested our model with BosphorusSign22k dataset which is a Turkish isolated SLR dataset. The proposed model achieves 92.97% classification accuracy. Our study shows that attention modeling with temporal shift on top of 2D CNN backbone gives competitive results in isolated SLR.

Open dataConvolutional neural networksSign language
Ahmet Faruk Çelimli
Boğaziçi University · Institute of Graduate Studies in Science
2022
00
Master'sOpen AccessEN

Spatially varying single image deblurring using cyclegans

In this thesis, we propose a novel method that learns to deblur hand images in the presence of spatially-varying object motion blur and unpaired blurry/sharp training pairs. While some success has been achieved in image deblurring by learning disentangled representations from synthetically blurred data, these methods do not perform well when objects in the frame are moving rapidly; consequently resulting in inferior pose estimation performances. This commonly occurs when the hands of a signer moves abruptly in a sign language setting. We propose to solve these problems by disentangling blur information from image content (hand texture, background). Lack of non-corresponding training pairs is dealt with cross-cycle consistency losses in blurring/deblurring branches based on disentangled representations and spatially-variant blur is extracted from blur-degraded regions using partial convolutions. We test our results both qualitatively and quantitatively on a novel hand blur dataset consisting of real blurry images and sharp frames as well as a reference synthetically blurred dataset.

Gizem Esra Ünlü
Boğaziçi University · Institute of Graduate Studies in Science
2020
00
DoctorateOpen AccessEN

Transfer learning for sign language recognition

Sign languages are visual languages that use hands, arms, and faces to communicate concepts. In the last decade, sign language recognition (SLR) research has made significant progress but still requires massive amounts of data to recognize signs. Despite efforts to create large annotated sign language datasets, applications that can translate for ordinary users in daily settings are yet to be produced. Most SLR research focuses on a few popular sign languages, leaving most sign languages, especially Turkish Sign Language (TID), under-resourced for sign language technology development. This dissertation addresses several open research questions about the development of SLR technology for TID from several perspectives. We generated BosphorusSign22k, an isolated SLR dataset for TID with 22k videos, and benchmarked state-of-the-art techniques on it. We proposed aligned temporal accumulative features (ATAF) to efficiently model sign language movements as dynamic and static subunits. Combined with methods using other modalities, the method achieves state-of-the-art performance on BosphorusSign22k. We then used regularized regression-based multi-task learning and presented task-aware canonical time warping for isolated SLR. The technique aligned and grouped signs to minimize discrepancies across different sources and emphasize class differences. Finally, we established a benchmark for cross-dataset transfer learning in isolated SLR. We evaluated supervised transfer learning algorithms using a temporal graph convolution-based SLR method. Experiments with closed and partial-set cross-dataset transfer learning reveal a substantial improvement over combined training and fine-tuning-based baseline techniques.

Convolutional neural networksImage processing-computer assisted
Ahmet Alp Kındıroğlu
Boğaziçi University · Institute of Graduate Studies in Science
2023
00
Master'sOpen AccessEN

Score level multi cue fusion for sign language recognition

In this thesis, we propose a Score-Level Multi Cue Fusion approach that improves the sign language recognition performance of the three dimensional convolutional neural networks. Sign Language is the communication language of the Deaf and Hearing-impaired individuals and performed using hand movements, facial gestures, and body alignment. Sign Language Recognition is the task that aims to understand sign language and gaining increasing popularity with the task becoming feasible due to the efficiency of the neural network. Previous work uses 3D CNN network variants to inspect SL properties in different settings. The vanilla 3D variant uses 3D kernels with high processing cost, the mixed convolution variant applies both 3D and 2D kernels respectively, and R(2+1)D variants exploit bottleneck connections to exploit the bottleneck dimension. Various studies use these networks to generate an end to end framework for tasks such as sign classification and translation. To achieve better performance, 3D CNN methods use the complicated neural network architectures that have a branch for every cue system. We evaluate the 3D network performances and propose a more straightforward approach which only adopts a single neural network that can process multiple cues at test time. We exploit the hand, body, and face cues by training single individual networks and fuse results by using a weighted score fusion. We test our method on the recently published Turkish Isolated SLR dataset. Despite the simple architecture, our method achieves \%94 percent classification rate on 744 different sign glosses. We hope that the multi cue approach can help with the other SLR tasks such as translation, which is stated as future work.

Image processing-computer assisted
Çağrı Gökçe
Boğaziçi University · Institute of Graduate Studies in Science
2020
00
Master'sOpen AccessEN

Statistical time series analysis methods with applications to portfolio management

In many domains of science and engineering, such as signal processing, bioinformatics and computational finance, sequential data modelling and analysis is essential for various tasks including clustering, anomaly detection and forecasting. In this work, we present the fundamentals of time series analysis methods with a focus on modelling dynamical systems. Our goal is to make statistical inferences that are able to account for our uncertainty about the system while also being able to incorporate domain specific knowledge into these system representations. Hidden Markov models have been extensively studied in the literature because of their relative simplicity and flexibility. We propose an extension to this model called the Gaussian-Gamma hidden Markov model which introduces an additional latent scale parameter, along with its state inference and parameter estimation algorithms. The model is inspired by our prior knowledge of financial markets in terms of displaying persistent regimes and having heavy-tailed distributions. The intuition behind the need for such a model in computational finance is discussed with respect to the shortcomings of the standard mathematical framework of constructing an optimal portfolio called the Mean-Variance Analysis. We illustrate the performance of our model in both synthetically generated and real financial data sets with regime identification and portfolio management problems. Results show that our model is able to discover meaningful insights about the dynamics of financial markets.

Yaman Kındap
Boğaziçi University · Institute of Graduate Studies in Science
2020
00
DoctorateOpen AccessEN

Conditional computation techniques in deep neural networks with conditional information gain

Recently, deep neural networks, particularly convolutional neural networks, have excelled in computer vision tasks such as image classification, object detection, and semantic segmentation. Their high performance stems from numerous layers and learnable parameters. However, this complexity poses challenges for efficient inference, especially on devices with limited computing power, like edge devices. Among numerous similar approaches in the literature to address this issue, conditional computing is an efficient inference method where parts of a deep neural network are used or skipped based on the properties of the input. In this thesis, we develop two main conditional computing approaches: "Conditional Information Gain Networks", where a neural network is designed in the shape of a tree, and the samples are routed based on network elements that are trained with information gain. The second one is the "Conditional Information Gain Trellis", which describes a trellis-shaped network that again allows the routing of the samples based on the decisions of routing units trained by information gain objectives. We develop loss functions, training methodologies, and regularizers for both models. For both models, we develop inference methodologies that allow the routing of samples over more than one route in these networks, where we try to achieve a balance between additional model performance and extra computational burden. These multiple-path routing approaches, which we call "Sparse Mixture of Experts" inference, are implemented using algorithms such as Bayesian Optimization, Cross-Entropy Entropy Search, and Reinforcement Learning. We show the results of these model designs with various experiments.

Ufuk Can Biçici
Boğaziçi University · Institute of Graduate Studies in Science
2024
00
Master'sOpen AccessEN

Occlusion-aware benchmarking in 3D human pose and shape estimation

3D human pose and shape reconstruction is a widely studied area in computer vision. In addition to non-rigid features and highly articulated joints, another challenge in this area is occlusion, which is common in nature. Although some methods explicitly try to handle occlusion cases, the benchmark against which they are evaluated is vague. The typical approach in the literature is to report the performance of the method on an occlusion-oriented subset. However, to form such a subset, it is necessary to quantify the occlusion in the samples. The existing approach uses the keypoints and bounding boxes to quantify and rank samples based on occlusion. However, it fails in several cases and tends to produce false positives. This study proposes the Occlusion Index, a novel index to quantify occlusion in images with high accuracy. The instance mask-based approach not only successfully quantifies occlusion, but also discriminates between occluders and occluders. It also reports the self-occlusion of a person, which is an unavoidable phenomenon in single-view reconstruction. The experiments show the superiority of the Occlusion Index by forming more challenging subsets, causing state-of-the-art occlusion-robust methods to fail more often. Also, some of the most occluded samples in the popular 3D human pose and shape estimation datasets are included.

Digital image analysisDigital image processing
Emre Girgin
Boğaziçi University · Institute of Graduate Studies in Science
2024
00

Other supervisors