Yıldız Technical University
Anabilim Dalı

Bilgisayar Bilimleri ve Mühendisliği Anabilim Dalı

Yıldız Technical University

189

Arşivlenen Tez

0

DOI Atanmış

0%

DOI Oranı

Anabilim Dalı

50 Tez
Yüksek LisansAçık ErişimTR

Konuşma dili işleme

İster Türkçe, ister Tagalog dilinde olsun, herhangi bir konuşma dilini kullanarak iletişim kurma kabiliyeti sadece insan ırkına mahsus bir üstünlüktür. Bilgisayarlar, bu dil kullanma yeteneğini insanlarla paylaşana kadar insanların günlük hayatlarında yapabildikleri pek çok işi yerine getiremeyeceklerdir. Sadece 3 yaşında olan bir çocuk, geçerli bir satranç oyunu oynayamamasına, en azından bir ustayı yenememesine rağmen kendi ana dilini rahatça konuşur ve anlar. Buna karşılık dünyada hala insanların bu dil yeteneğinin üstesinden gelebilmiş bir bilgisayar programı geliştirilebilmiş değildir. Konuşma dilinin yapısını anlayabilmek çok zor bir işlemdir. Bunun için kullanılan dilin gramer bilgisine ve tartışılmakta olan konuya bağlı olarak genel kültüre ihtiyaç vardır. Ancak, her şeyden önce bilgisayarın, işleyeceği konuşma dili hakkında yeterli bir alt yapıya sahip olması gerekir. Bunun için de o dilin kelimeleri, kelimelerin fonksiyonları, cümle kurarken kullanılan yapılar ve kurulmuş cümlelerin çeşitleri hakkında ayrıntılı bilgisi olmalıdır. Bu bilgilerin veri tabanına yerleştirilmesinden sonra dışarıdan girilen bilginin analiz edilmesi, cümlenin manalandırılması ve metin içinde ne anlama geldiğinin değerlendirilmesi aşamaları gelir. Bu aşamaların ortak adı "konuşma dili işleme"dir. Bu tezde, zekânın ve yapay zekâ'nın tanımları ayrıntılı olarak incelendikten sonra, bu kavramları kullanarak anlama için gerekli olan dilbilimsel kültürün düzgün bir konuşma dili için nasıl kullanıldığı incelenmiştir. Daha sonra PROLOG dili kullanılarak, bir Türkçe cümle parser'ı yazılmış ve Yunus Emre'nin 180 beyiti üzerinde kullanılarak, şairin kullandığı kelimelerin köken ve görev dağılımları çıkarılmıştır.

Doğal dil işleme
Mihriban Betül Yılmaz
Yıldız Technical University · Fen Bilimleri Enstitüsü
1992
00
Yüksek LisansAçık ErişimTR

World-Wide-Web üzerinde çokluortam veritabanı uygulaması

ÖZET Son yıllarda entegre devre ve iletişim teknolojilerindeki ilerileme ile bilgi işlemede daha önce görülmemiş işlem ve iletişim hızlarına erişilebilmiştir. Bunun doğal bir sonucu olarak da önceden işlemesi çok pahalı olan çokluortam veri de bilgi işlemde yaygın olarak kullanılmaya başlandı. Ses, video ve resim gibi çokluortam verisi daha çok gerçek hayatın aynen bilgi işleme hazır hale getirilmesi ile oluşturulur ve sekizli başına en düşük anlamı ifade eden veri tipidir. Sonraki kuşak bilgisayar sistemleri ise daha da büyük bilgiyi yine daha hassas ve yaygın işleme yeteneğine sahip olacaklardır. Bu gelişmeler doğrultusunda, çokluortam verisini daha etkin işlemek için yeni bilgisayar algoritmaları ve işlem teknikleri geliştirilmeye başlanmıştır. Yeni teknikler çokluortam verisinin hatalara karşı neredeyse duyarsız olmasından faydalanıp işlem ve/veya iletişim hızını arttırmaya yönelik olarak geliştirilmişlerdir. Öte yandan Internet dünya çapında ve her geçen gün yayılarak, özellikle uzak mesafeler arasında haberleşmenin en önemli araçlarından biri haline gelmiştir. Internet' in böylesi genişlemesindeki en önemli etmen ise basitleştirilmiş bazı iletişim kurallarına uyabilen her bilgisayara kullanım imkanı vermesidir. Bu tezde yukarıda sayılar gelişmelerin genel yapısı üzerinde durulup, çokluortam verisinin yönetimi, düzenlenmesi ve WWW üzerinde yayınlanmasını sağlayacak bir "Çokluortam Haberler Veritabanı" geliştirilmiştir. XII

MultimedyaNesnesel veri tabanıİnternet
F. Önder Yıldırım
Yıldız Technical University · Fen Bilimleri Enstitüsü
1998
00
DoktoraAçık ErişimTR

Hiyerarşik grupsal kurumlarda kullanılacak bir şifre sistemi

ÖZET Kurumlar, etkinlik alanlarına bağlı olarak askeri, devlet güvenliği ya da ticari kaygılardan dolayı bünyelerinde mevcut bilgilerin güvenliğini sağlamak isterler. Alınabilecek fiziksel güvenlik yöntemleri yetersiz kaldığında ise bilgi güvenliğini sağlamanın tek yolu, şifreleme yöntemlerinin kullanılmasıdır. Kurumlar, işlevsel farklılıklar içeren gruplardan oluşurlar. Bilgiler konularına göre farklı grupları ilgilendirebilir. Gruplar içinde de personelin görev farklılığım temel alan bir hiyerarşik seviyelendirme mevcuttur. Bir bilgi bir grubun yalnızca en üst seviyedeki yöneticisini ilgilendirirken, başka bir bilgi daha alt seviye personelini de ilgilendirebilir. Bu çalışma, hiyerarşik grupsal kurumlarda veri aktarma iletişim güvenliğini sağlamak için kullanılabilecek bir modern (açık anahtarlı) şifre sistemi tanımım içermektedir. Anahtarların üretim, korunma ve dağıtımı ile mesajların dağıtımından Mesaj Dağıtıcısı (MD) sorumludur. MD, sistemde mevcut her kişiye ait olduğu gruba ilişkin bir, ve hiyerarşik seviyeye ilişkin olarak bir olmak üzere iki anahtar dağıtacaktır. Mesajlar, sistemde dağıtılmadan önce hedef grup ile hiyerarşik anahtarların kullanımı ile iki kez şifrelenecektir. İlgili grup ile seviyede bulunan kişiler, ellerindeki anahtarları kullanarak şifrelenmiş mesajları çözebilirler; böylece mesaj larm ilgili seviye ve gruptaki personel tarafından okuyabilmesi garantilenmiş olur. Anahtar mevcut olmadığı durumda ise şifrelenmiş mesajlardan asıl mesajları üretmek mümkün değildir. Çalışma: kullandığı algoritma ile, tanımlanan grupsal hiyerarşik modelde şifrelemeyi sağlaması, alt seviye anahtarlarının basit bir işlem ile üretilebilmesi ve yalnızca iki anahtarın muhafazasına ihtiyaç duyurması nedenleri ile özgün bir algoritmayı içermekte olup, bu algoritmada RSA ile El Gamal şifre sistemleri tarafından kullanılmış ve güvenli oldukları ispatlanmış olan tek-yönlü NP-Complete fonksiyonlar temel olarak alınmıştır.

Bilgi erişimKriptografiKurum+1
Vedat Coşkun
Yıldız Technical University · Fen Bilimleri Enstitüsü
1998
00
Yüksek LisansAçık ErişimEN

Real-time intelligent strawberry harvesting and quality determination system using computer vision and deep learning

Strawberries have a comparatively extended harvesting period, which poses the need for intelligent harvesting systems that identify various stages of strawberry ripeness to alleviate fatigue and reduce the costs associated with this task. These systems offer potential solutions to enhance productivity while minimizing labor requirements. Research in real-time strawberry detection still has several gaps to address. These gaps include managing imbalance label distribution, exploring data augmentation techniques, optimizing preprocessing and training parameters, and investigating advanced topics such as fine-tuning. This research focuses on developing an accurate and efficient real-time strawberry detection model. The main objective is to locate and identify strawberries in agricultural orchards and assess their quality by detecting overripe and decaying fruits. For this purpose, a diverse High-Quality Annotated Strawberry Dataset(HQASD) was collected with 5000 RGB images belonging to ten strawberry maturity levels. Different types of data augmentation techniques were applied, and the most advanced and novel approach used in this context was the Cycle-Consistent Generative Adversarial Network (CycleGAN) for generating images of an underrepresented class from an overrepresented one. The latest real-time object detection model, You Look Only Once (YOLOv7), is trained on real and synthetically generated data. Deep experiments and various scenarios were conducted to investigate all the aspects that impact the model's overall performance. Hyper-tuning is performed on the model's parameters to optimize its performance. The best-obtained models for identifying ten classes achieved Mean Average Precision (mAP) 98.7% at Intersection over Union (IoU) threshold 0.5, and the best model for identifying three classes scored 96.51% mAP for all types at the same threshold, where the model trained on synthetic images performed 98.4% mAP for all classes which shade the effectiveness of employing CycleGAN along with fine-tuning techniques. In addition, the research succeeded in creating HQASD, which provides a comprehensive representation of the strawberry maturity spectrum, which makes it well-suited for computer vision tasks.

Real time systemsObject detectionStrawberry
Nagham Yassın Alhawas
Çukurova University · Fen Bilimleri Enstitüsü
2023
00
Yüksek LisansAçık ErişimEN

The detection and recognition of faces in the internet of things for security applications

In recent years, the security constitutes are the most important section of human life. Security of the house and the family is important for everybody. Automation of a home is an exciting field for security applications. This area has developed with new technologies such as the Internet of things (IoT). In IoT, each device behaves as a small part of an internet node and each node communicates and interacts. Currently, security cameras are used in order to construct safety in areas, cities and homes. The camera records the events, and when a problem occurs, it will detect by monitoring the old recording. At this time, the cost is the greatest factor. This system is very helpful to reduce the cost of monitoring the movement from outside. The purpose of this thesis is to describe a security alarm application by utilizing low preparing power chips and Internet. Also, a new online method is proposed to detect and recognize faces on Raspberry Pi in the IoT. Raspberry Pi operates and controls movement detectors. It will monitor and record the motions for future playback. This thesis proposes an analysis of images via computer vision to detect and recognize faces in the analyzed images. If these frames contain a face, the system will detect and recognize the face. This system is appropriate for small personal range surveillance, as, in personal office security, home, parking entrance and bank locker room. The face detection and recognition in the IoT is a very important problem for a security and surveillance system. Also, face detection and recognition is presently a very active research area. The proposed system is very helpful to reduce the cost of monitoring the movement from outside. On the other hand, in this research, an IoT-based system is combined with computer vision in order to detect human's body. A Raspberry Pi 3 cards with the size of a credit card is used for this purpose. A motion is detected by the PIR sensor mounted on the Raspberry Pi. PIR sensor helps to monitor and get alerts when a movement is detected. Afterward, the human's body detects in the captured image and is sent to a smartphone by using telegram application.

Image processingFace detectionFace recognition
Nashwan Adnan Othman
Fırat University · Fen Bilimleri Enstitüsü
2018
00
Yüksek LisansAçık ErişimEN

ComScribe: A communication monitoring tool for multi-GPU platforms

GPU communication plays a critical role in performance and scalability of multi-GPU accelerated applications. With the ever increasing methods and types of communication, it is often hard for the programmer to know the exact amount and type of communication taking place in an application. Though there are prior works that detect communication in distributed systems for MPI and multi-threaded applications on shared memory systems, to our knowledge, none of these works identify intra-node GPU communication. In this work we present ComScribe, a tool that identifi es and categorizes types of communication among all GPU-GPU and CPU-GPU pairs in a node. Our tool is built on top of NVIDIA's pro lfier nvprof for capturing intra-node point-to-point communication resulting from explicit communication primitives, Uni ed Memory operations, and Zero-copy Memory transfers. For monitoring collective GPU-GPU communication in a node, ComScribe intercepts NCCL's collective primitives at runtime and records data transfers among GPUs. It visualizes data movement as a communication matrices for both number of bytes transferred and the number of transfers. To validate our tool on 16 GPUs, we present communication patterns of 13 micro and 3 macro-benchmarks from NVIDIA, CommjScope, and MGBench benchmark suites. To demonstrate tool's capabilities in real-life applications, we also present insightful communication matrices of three deep neural network models. All in all, ComScribe can guide the programmer in identifying groups of communicating GPUs, the volume of communication, and types of primitives used. This o ers avenues to detect performance bottlenecks and more importantly communication bugs in an application.

Palwısha Akhtar
Koç University · Fen Bilimleri Enstitüsü
2021
00
Yüksek LisansAçık ErişimEN

Increasing efficiency of combinatorial optimization problems on quantum annealers using classical computers

Google's claim of quantum supremacy is a big milestone in the history of quantum computing. Despite the claim, the practical applicability of quantum computers remains questionable due to the low number of quantum bits and high noise rates. An alternative model of quantum computing is quantum annealing, which is capable of solving only an optimization problem in a specific format. Quantum annealing is being actively researched due to the fact that the scale of such devices has increased up to thousands of qubits. This relatively high number of qubits enables quantum annealers to solve problems of larger sizes, hence makes them usable in real-life scenarios. The specific format of the solved optimization problem is called an Ising formulation, which can also be represented as quadratic unconstrained binary optimization (QUBO). The QUBO consists of a set of qubits with corresponding bias weights and the quadratic weights between the qubits. This thesis presents two schemes for weight optimization in the QUBO formulation of two different combinatorial optimization problems. Both schemes involve a classical computer redefining the QUBO weights that is interfaced with an annealing device, which solves the QUBOs. The first combinatorial problem is the task assignment problem, in which the biases represent the computational costs of tasks and quadratic terms model communication between tasks. The second problem is the circuit mapping problem, where the biases represent the fidelity of quantum gates, while the quadratic terms model qubit movement. The first approach named weight optimization algorithm (WOA) searches for a desirable ratio between the qubit biases responsible for fidelity of mapping quantum gates to physical qubit topology and the quadratic terms responsible for qubit movement. The desirability of the ratio is defined by the total fidelity resulting from both qubit movement and mapping. The second presented model uses ant colony optimization (ACO) to update the quadratic terms of the QUBO that solves the task assignment problem. However, this model can be generalized to any combinatorial optimization problem solvable by the ACO. Efficient updates of the weights based on the answers from previous QUBOs are expected to guide the reformulated QUBOs towards the optimum of the objective function. At the same time, this algorithm would allow utilizing the stochasticity of quantum annealers for better exploration of the solution space and their speed for faster generation of candidate solutions. The introduction of the WOA into the quantum annealing workflow for quantum circuit mapping resulted in reduced qubit movement in 72.9% of all problem samples. Moreover, it allowed to increase the total fidelity of the mapped circuit by 39% on the IBM Vigo device and 107% on IBM QX2. The experiments have been performed on the tabu search QUBO solver from the D-Wave quantum annealing software stack. The results for the ant colony weight optimizer are limited due to the unavailability of a quantum annealing device.

Quantum computersMetaheuristicsHeuristic methods
Ilyas Turımbetov
Koç University · Fen Bilimleri Enstitüsü
2021
00
DoktoraAçık ErişimEN

Platform and data-aware execution of sparse triangular solve on CPU-GPU heterogeneous systems

Sparse triangular solve (SpTRSV) is an important computational kernel used in many scientific and numerical linear algebra applications such as direct methods, iterative solvers and least square problems. In comparison with other sparse kernels such as sparse matrix-vector multiplication (SpMV), SpTRSV is an inherently sequential operation owing to the presence of dependencies among computations of different unknowns. Often, it has also been observed to be one of the most time consuming operations in an application. The performance of a given algorithm highly depends upon the sparsity characteristics of the input matrix and the underlying hardware. Several SpTRSV algorithms and their implementations are available for the CPUs and the GPUs. Unfortunately, there is no single algorithm or hardware platform that has been shown to achieve the best performance for all input matrices. In this dissertation, we propose tools and techniques aimed at extracting higher SpTRSV performance for a given input matrix on modern CPU-GPU heterogeneous systems. Towards this end, we adopt a two-pronged approach: Depending upon the sparsity characteristics of the input matrix, we propose to (i) automatically select the best CPU or GPU SpTRSV algorithm for the matrix, (ii) split a single SpTRSV execution into parallel and sequential parts such that the parallel part is executed with a parallel-friendly algorithm and the sequential part is executed with a sequential-friendly algorithm. The aim is to achieve a higher SpTRSV performance than using a single algorithm. The two algorithms can also potentially execute on two different platforms (CPU and GPU). For the SpTRSV algorithm selection, we propose a supervised machine learning-based prediction framework to predict the SpTRSV implementation giving the fastest execution time for a given sparse matrix based on its structural features. The framework works by extracting matrix features, collecting algorithm performance data, and training a prediction model with around 1000 real square matrices from the SuiteSparse Matrix Collection. Once trained on a given machine, the model can predict the fastest SpTRSV implementation for a given matrix by paying a one-time matrix feature extraction cost. The framework is also capable of taking into account CPU-GPU communication overheads that might be incurred in scientific applications such as iterative solvers. We test the framework on a modern CPU-GPU machine. Experimental results on a modern CPU-GPU platform (Intel Xeon Gold + NVIDIA Tesla V100 GPU) show that the fastest algorithm is selected with a reasonable accuracy (87%) as well as the predicted SpTRSV implementation achieves significant speedups compared (1.4-2.7× harmonic mean) with a lazy choice of a single algorithm. Secondly, for dividing the SpTRSV execution between the CPU and GPU, we propose an SpTRSV split execution model. The model has been designed based on the empirical evidence from the existing research that; (i) a highly parallel algorithms perform well for SpTRSV requiring few sequential steps, with high number of unknowns computed per step, (ii) a sequential algorithms perform better for SpTRSV requiring more steps with few unknowns computed per step. For matrices having a mix of highly parallel and sequential steps, a parallel algorithm is expected to provide benefits for steps with large number of unknowns, its performance is expected to deteriorate for steps with few unknowns. The converse can be said about the performance of sequential algorithms for such matrices. With the aim of performing split-execution for such matrices using a highly parallel and a sequential algorithm, our split execution model can automatically determine the suitability of an SpTRSV for split-execution, find the appropriate split point, and execute SpTRSV in a split fashion using two SpTRSV algorithms while automatically managing any required inter-platform communication. The model is implemented as a C++/CUDA library supporting multiple CPU-GPU algorithms. Experimental evaluation of the model on two CPU-GPU with a matrix dataset of 327 matrices from the SuiteSparse Matrix Collection shows that our approach correctly selects the fastest SpTRSV method (split or unsplit) for 88% of the matrices on Intel Xeon Gold + NVIDIA Tesla V100 and 83% of the matrices on Intel Core I7 + NVIDIA G1080 Ti platform achieving speedups up to 10× and 6.36×, respectively. We expect that the tools and methods proposed in this dissertation will benefit both academia and industry while at the same time will pave the way for further research in the area of efficient exploitation of CPU-GPU systems for sparse linear algebra computations.

Najeeb Ahmad
Koç University · Fen Bilimleri Enstitüsü
2021
00
DoktoraAçık ErişimEN

Cognitively-inspired deep learning approaches for grounded language learning

Designing machines that can perceive the surrounding world and interacting with us using human language is one of the long-standing goals of artificial intelligence. Although tremendous progress has been made to model the linguistic meanings computationally, how to best integrate linguistic and perceptual processing in multi-modal tasks is a significant open problem. This thesis explores several cognitively-inspired neural architectures that consider the different aspects of the language's role in cognition, visual perception, and task execution. Proposed models incorporate design choices motivated by cognitive science studies and are based on the common patterns in vision-language tasks. We begin by presenting an encoder-decoder network with a novel channel-based perceptual attention mechanism and its application to the navigational instruction following task. The perceptual processing component of this architecture is designed to focus on individual objects and properties within the environment using the language priors while preserving the spatial relations. To benefit from the designed component, we also propose an improved agent-centric world representation to allow the model to reason over the perception spatially. Next, we explore the usage of the Neural Module Networks approach in a real robotic system for the first time. Since collecting large-scale real world data is a labor-intensive and expensive work, the system learns the language grounding on simulated data and the perceptual representation separately to overcome the scarce data problem. However, because of the separate learning processes, inconsistencies arise between the user's and robot's world models. To overcome this, we propose a Bayesian learning approach that uses the implicit information in the instruction to update the perceptual belief to align what the user sees and what the robot perceives. In both parts, we demonstrate systems that use the high-level effect of language on visual processing, which operates on high-level representations. In addition to this, in the last part, we investigate the effect of language on low-level visual processing. To this end, we condition one or both low-level and high-level visual processing branches of a backbone architecture on language using language filters and apply these models to the image segmentation from referring expression task. Experiments show that modulating both low-level and high-level visual processing with language significantly improves the language grounding performance.

Ozan Arkan Can
Koç University · Fen Bilimleri Enstitüsü
2021
00
Yüksek LisansAçık ErişimEN

Self-supervised representation learning from demonstration

Given the growing demand for the application of robotics in an increasingly wide range of tasks and environments, robot learning as opposed to programming is getting more relevant by the day, since programming robots is tedious and usually only feasible in controlled domains. However, gathering data with robots is an expensive and time-consuming task, so the learning methods must cope with the limited amount of data available. When working with high dimensional perceptual data, learning low-dimensional, useful representations is a key aspect in dealing with the low-data setting. In this thesis, we develop a neural network architecture to learn perceptual representations from few human skill demonstrations in a self-supervised manner. The developed models take the sequentiality of the data and the low-data nature of the problem into account. These representations are used to learn perceptual goal models of the demonstrated skills. These models can monitor the learned skill executions and be used in reinforcement learning to generate reward signals, without explicit reward engineering. Our simulated and real robot evaluations with object manipulation skills show that the learned representations result in better goal models in terms of monitoring and reinforcement learning performance compared to generic dimensionality reduction methods. We further introduce transfer learning approaches in the context of learning from demonstration and show positive transfer between different objects for the same skill, between the same object for different skills under certain conditions, and also between different perceptual domains. Transferring knowledge as enabled by our modular neural architecture allows us to leverage existing data for continual learning into the future. Overall, we show that our proposed self-supervised representation learning architecture has the potential to improve learning from demonstration approaches with a perceptual component.

Ercan Alp Serteli
Koç University · Fen Bilimleri Enstitüsü
2021
00
Yüksek LisansAçık ErişimEN

Stacked frequency-timeGRUs for continuous arousal recognition from musical audio

We address the problem of continuous arousal detection for emotion recognition in musical audio pieces where emotions are represented in the two-dimensional arousal-valence space. We propose a novel method which is a combination of two recurrent neural networks using mel-spectrogram features: A bidirectional GRU network along the frequency dimension as a feature extractor, stacked with a GRU network along the temporal dimension, which is unidirectional for real-time adaptability. The method is evaluated on the MediaEval2015- Emotion in Music Dataset, achieving an RMSE of 0.215 which is better than the results reported by real-time adaptable state-of-the-art models.

Aslıhan Çeliker
Koç University · Fen Bilimleri Enstitüsü
2021
00
DoktoraAçık ErişimEN

Precise event sampling: In-depth analysis and sampling-based profiling tools for data locality

Precise event sampling is a profiling feature in current commodity CPUs that allows sampling of hardware events and identifies the instructions that trigger the sampled events. It offers the ability to detect performance bottlenecks with low overhead as well as the locations of the bottlenecks in source code. There have been a number of profiling tools developed using this feature that detect various sources of performance bottlenecks. However, none of these tools detects inter-thread data movement nor measures data locality in multithreaded applications, which have become widely used due to the ubiquity of multicore architectures. Furthermore, though this hardware facility has been used in multiple profiling tools, there have been only few works that analyze it in terms of accuracy and overhead. All of these works target only the facility in Intel architecture, and none of these works evaluates other aspects of precise event sampling such as memory overhead, stability, and functionality of the facility. In this dissertation, we present threefold major contributions. First, we perform the most comprehensive and in-depth qualitative and quantitative analyses to date on PEBS and IBS, which are the precise event sampling facilities of two major vendors, Intel and AMD, respectively. Next, we show the potential for imaginative use of precise event sampling in developing low overhead yet accurate profiling tools for multicore and design two diagnostic tools with a particular focus on data movement as it constitutes the main source of inefficiencies. First of such tools is ComDetective that detects inter-thread communications, classifies them into true sharing or false sharing, and records them in the form of communication matrices. Second is ReuseTracker that measures data locality in private and shared caches of multithreaded applications. ComDetective and ReuseTracker leverage precise event sampling to profile multithreaded applications accurately and with low overheads compared to their state-of-the-art alternatives. To analyze key differences between Intel PEBS and AMD IBS, we firstly developed a series of carefully designed microbenchmarks. Through our qualitative analysis and quantitative study using the microbenchmarks, we found that Intel PEBS samples hardware events more accurately and with higher stability in terms of the number of samples that it captures, while AMD IBS records richer set of information at each sample. We also discovered that both PEBS and IBS are afflicted with bias when sampling the same event across multiple different instructions in a code. Moreover, we also show how our findings from the quantitative experiments using the microbenchmarks are relevant for a full-fledged profiling tool that runs on Intel and AMD machines. We develop ComDetective, a profiling tool that captures inter-thread communications accurately and with low runtime and memory overheads. ComDetective employs precise event sampling to sample memory accesses and utilizes hardware debug registers to detect inter-thread communications. In addition to detecting communications, ComDetective can also classify them into true or false sharing. Its time and memory overheads are only 1.30× and 1.27×, respectively, for the 18 applications studied under 500K sampling interval. Using ComDetective, we generate insightful communication matrices from several microbenchmarks, PARSEC benchmark suite, and some CORAL applications and compare the produced matrices against the matrices of their MPI counterparts. Using ComDetective, we identify communication bottlenecks in a few codes and achieve up to 13% speedup from code refactoring those codes. We also design ReuseTracker, which is a profiling technique that measures reuse distance - a widely used metric that measures data locality. Reuse distance is a measurement of data locality as it is the number of unique memory locations that are accessed between two consecutive accesses to a particular memory location (use and reuse). ReuseTracker leverages precise event sampling to capture uses and debug registers to detect reuse in measuring reuse distance. ReuseTracker can measure reuse distance in multithreaded applications by also considering cache-coherence effects with much lower overheads than existing tools. It introduces only 2.9x time and 2.8x memory overheads. It achieves 92% accuracy when verified against a carefully crafted configurable microbenchmark that can generate user-specified reuse distance patterns. We demonstrate in two use cases how ReuseTracker can be used to guide code refactoring by detecting spatial reuses in shared caches that are also false sharing and how it can also be used to predict whether certain applications can benefit from adjacent cache line prefetch optimization. We expect that the analysis, algorithms, and the tools presented in this dissertation will benefit hardware architects in designing new precise event sampling features and performance engineers in performance tuning of their software while also paving the way for a new generation of low-overhead profiling tools. Moreover, the outcomes of the dissertation can be used by the end-users (e.g., data analysts, engineers, compiler developer) to identify the performance issues and increase the data locality aspects of their software.

AnalysisLinuxSampling
Muhammad Adıtya Sasongko
Koç University · Fen Bilimleri Enstitüsü
2022
00
Yüksek LisansAçık ErişimEN

Utilizing coarse-grained data in low-data settings for event extraction

Annotating text data for event information extraction systems is hard, expensive, and error-prone. We investigate the feasibility of integrating coarse-grained data (document or sentence labels), which is far more feasible to obtain, instead of annotating more documents. We utilize a multi-task model with two auxiliary tasks, document and sentence binary classification, in addition to the main task of token classification. We perform a series of experiments with varying data regimes for the aforementioned integration. Results show that while introducing extra coarse-grained data offers greater improvement and robustness, a gain is still possible with only the addition of negative documents that have no information on any event.

Event tree methodData setsInference
Osman Mutlu
Koç University · Fen Bilimleri Enstitüsü
2022
00
Yüksek LisansAçık ErişimEN

Exploring mixed and multi-precision SpMV for GPUs

Sparse Matrix-Vector Multiplication (SpMV) is one of the key memory-bound kernels commonly used in industrial and scientific applications. To improve its data movement and benefit from higher compute rates, there are several efforts to utilize mixed precision for SpMV. Most of the prior-art focus on performing the SpMV in different precisions throughout the entire application, such as an iterative solver (e.g., CG, GMRES) where certain steps can be done with lower precisions. More recently, methods of using mixed-precision within a single SpMV has been of consideration. For instance, one can decide precision for each non-zero value in the matrix, and then split the input into multiple matrices with their respective precisions; or, decide the precision for each block in a given block-diagonal matrix format. In this work, we are interested in this more fine-grained approach of mixedprecision SpMV. To this extent, we extend an existing entry-wise precision based approach by deciding precision for each row, motivated by the granularity of parallelism on a GPU where groups of threads process rows in row-compressed sparse matrices. We propose mixed-precision CSR storage methods with row permutations and describe its greater load-balance compared to the existing method. We also consider a multi-precision case where single and double-precision copies of the matrix are stored priorly, and further extend our mixed-precision SpMV approach to comply with it. To evaluate our methods, we apply them in two real-life applications: a multiprecision Jacobi method and a multi-precision Cardiac modeling application. We further extend our mixed-precision methodology to be used with the ELLPACK-R format. We demonstrate the effectiveness of the proposed SpMV methods on an extensive dataset of real-valued large sparse matrices from the SuiteSparse Matrix Collection using an NVIDIA V100 GPU.

Jacobi formsSparse matrixesVector multiplication
Erhan Tezcan
Koç University · Fen Bilimleri Enstitüsü
2022
00
Yüksek LisansAçık ErişimEN

GAN inversion based image manipulation with text-guided encoders

Style-based Generative adversarial networks (StyleGAN) enable very high quality image synthesis while learning disentangled latent spaces. Hence, there is a lot of recent work focusing on semantic image editing by latent space manipulation. A particularly emerging field is editing images based on target textual descriptions. Existing approaches tackle this problem either by performing instance-level latent code optimization which is not very efficient or by mapping predefined text prompts to editing directions in the latent space. In contrast, in this thesis work, we present two novel approaches that enable image editing guided by textual descriptions. Our idea is to use either a text-conditioned encoder network or a text-conditioned adapter network that predicts a residual latent code in a feed forward manner. Both quantitative and qualitative results demonstrate that our methods outperform competing approaches in terms of manipulation accuracy, i.e., how well the synthesized images match the textual descriptions while ensuring highly realistic results and preserving features of the original image. We also demonstrate that our method can generalize to various domains including human faces, cats, and birds.

ManipulationAutoencodersProducer networks
Ahmet Canberk Baykal
Koç University · Fen Bilimleri Enstitüsü
2022
00
Yüksek LisansAçık ErişimEN

Domain-adaptive self-supervised pre-training for face & body detection in drawings

Drawings are powerful means of pictorial abstraction and communication. Understanding diverse forms of drawings, including digital arts, cartoons, and comics, has been a major problem of interest for the computer vision and computer graphics communities. Although there are large amounts of digitized drawings from comic books and cartoons, they contain vast stylistic variations, which necessitate expensive manual labeling for training domain-specific recognizers. In this work, I show how self-supervised learning, based on a teacher-student network with a modified student network update design, can be used to build face and body detectors. My setup allows exploiting large amounts of unlabeled data from the target domain when labels are provided for only a small subset of it. I further demonstrate that style transfer can be incorporated into my learning pipeline to bootstrap detectors using a vast amount of out-of-domain labeled images from natural images (i.e., images from the real world). My combined architecture yields detectors with state-of-the-art (SOTA) and near-SOTA performance using minimal annotation effort. Through the utilization of this detector architecture, I accomplish a set of additional tasks. First, I extract a large set of facial drawing images (~1.2 million instances) from unlabeled data and train SOTA generative adversarial network (GAN) models to generate and a SOTA GAN inversion model to reconstruct faces. When the detector-aided data is leveraged, these generative models successfully learn diverse stylistic features. Secondly, I implement an annotation tool to enlarge the existing set of annotated data. This tool offers users to annotate bounding boxes of panels, speech bubbles, narrations, faces, and bodies; to associate text boxes with faces and bodies; to transcript the text; to match the same characters in the image.

Barış Batuhan Topal
Koç University · Fen Bilimleri Enstitüsü
2022
00
Yüksek LisansAçık ErişimEN

Learning markerless robot-depth camera calibration and end-effector pose estimation

Robot arms are being used more and more in unstructured environments. As such, they are relying more on vision sensors compared to traditional factory robots which are placed in highly structured, controlled and caged work-cells. Some applications that rely on vision data include bin-picking, box picking and placing, assembly and part feeding in mixed human-robot work cells, inspection, quality control etc. Vision data is mainly required to localize the objects to be manipulated, perform measurements and detect near-by humans. Vision based robot systems require extrinsic calibration between the robot and camera in order to work properly. This is a time consuming and tedious procedure which can be expensive as well. Fast, flexible and precise robot-camera calibration is essential for not only industrial environments but also academic lab environments where the location of the camera and/or robot needs to be frequently or is accidentally changed. Extrinsic calibration between a robot arm and camera is a decades old challenge still prevalent to this day. Traditional techniques work by estimating pose of the camera relative to a fiducial marker from multiple points and matching these estimations with the robot's pose. Recent learning based approaches predict extrinsic calibration from images relying heavily on simulation data. In this thesis, we present a learning based markerless extrinsic calibration system that uses a depth camera. We learn models for end-effector (EE) segmentation, single-frame rotation prediction and keypoint detection, from automatically generated real-world data. Our models are based on MinkUNet and PointNet++ architectures. We use a transformation trick to get EE pose estimates from rotation predictions and a matching algorithm to get EE pose estimates from keypoint predictions. We further utilize the iterative closest point (ICP) algorithm, multiple-frames and outlier detection to increase calibration robustness. Our results on the test set with previously unseen camera locations give sub-centimeter (0.74 cm) and less than 0.05 radians (1.69 degrees) average calibration errors and 1.00 cm and 2.74 degrees average pose estimation errors. In addition, we released an open source easy to use tool for robot users to handle robot-camera calibration with a few mouse clicks by seamlessly integrating all the models and algorithms discussed in this thesis.

Computer visionDeep learningFlexible robot arm+3
Buğra Can Sefercik
Koç University · Fen Bilimleri Enstitüsü
2022
00
Yüksek LisansAçık ErişimEN

Addressing the static scene assumption and the scale ambiguity in self-supervised monocular depth estimation

Self-supervised monocular depth estimation is the task of estimating per-pixel depth from a single image without any supervision. Typically, there are two networks to estimate depth and camera pose between consecutive frames, which are then used to reconstruct one view from another for self-supervision. There are two problems with this approach. Firstly, the scene is assumed to be static and the only motion is due to the camera, namely the static scene assumption, however, this is frequently violated in real-world driving scenarios. Due to this assumption, monocular depth methods struggle to produce accurate predictions in the moving regions of the scene. Current methods either ignore moving regions or require an additional instance segmentation input to identify and separately process moving regions. In this thesis, we first propose MonoDepthSeg to jointly estimate the depth and decompose the scene into moving regions to model the motion of dynamic objects. We show that going beyond the static scene assumption improves the accuracy of depth prediction, especially in moving regions. The second problem of self-supervised monocular depth estimation methods is the scale ambiguity. The estimated depth values are in an unknown scale which is typically handled with normalization with respect to the ground truth scale value during inference. We revisit the traditional paradigm of plane and parallax to address this issue and propose DepthP+P to estimate depth in metric scale. Our method shows promising results that are metric scale without any additional normalization.

DepthStageEstimation techniques
Sadra Safadoust
Koç University · Fen Bilimleri Enstitüsü
2022
00
Yüksek LisansAçık ErişimEN

Two-level temporal relation model for video instance segmentation

In Video Instance Segmentation (VIS), current approaches either focus on the quality of the results, by taking the whole video as input and processing it offline; or on speed, by handling it frame by frame at the cost of competitive performance. In this work, we propose an online method that is on par with the performance of the offline counterparts. We introduce a message-passing graph neural network that encodes objects and relates them through time. We additionally propose a novel module to fuse features from the feature pyramid network with residual connections. Our model, trained end-to-end, achieves state-of-the-art performance on the YouTube-VIS dataset within the online methods. Further experiments on DAVIS demonstrate the generalization capability of our model to the video object segmentation task. We also evaluate our work on autonomous driving setting and show comparable results in KITTI MOTS dataset.

Object detectionVideo segmentationFace perception
Çağan Selim Çoban
Koç University · Fen Bilimleri Enstitüsü
2022
00
Yüksek LisansAçık ErişimEN

Answering spatial density queries under local differential privacy

Spatial density queries are fundamental in many geospatial data analysis tasks and have numerous applications in the real world, such as determining crowded areas, estimating traffic density, navigation, passenger demand analysis, and so forth. However, answering spatial density queries based on users' data may violate users' privacy by exposing their true locations to an untrusted third party (e.g., a server acting as the service provider or data collector). In this thesis, we propose a solution for answering spatial density queries while preserving Local Differential Privacy (LDP), a state-of-the-art privacy protection standard. Our solution consists of four main steps: partitioning, finding sensitivity, user-side noisy response computation, and server-side estimation. For the first step, we initially propose three basic partitioning strategies: Singleton Partitioning, Holistic Partitioning and Random Partitioning. Based on our qualitative and empirical analysis of the three basic strategies, we design and implement an improved strategy called Advanced Partitioning. For the second step, we adapt graph-based modeling of query sets from the centralized DP literature. Advanced Partitioning also leverages and extends this technique by formulating the partitioning problem as a vertex coloring problem on the graph representation of a query set. For the third and fourth steps, in addition to adapting two popular LDP protocols (GRR and RAPPOR) to our solution, we propose an extension for the Optimized Unary Encoding (OUE) protocol so that it can be employed in our solution. We call the extended protocol Optimized Bitvector Encoding (OBE). OBE is applicable to not only the problem of answering spatial density queries, but also in arbitrary LDP problems with bitvector encodings. We formally prove that the user-side perturbation step of our OBE protocol satisfies LDP and its server-side estimation step produces unbiased estimates. Combining the different partitioning strategies and LDP protocols, we obtain a total of 8 different approaches for answering spatial density queries under LDP, all of which can be parsed as instances of our four-step solution with different choices in the individual steps. We perform an extensive experimental evaluation of these approaches using 4 real-world datasets, varying number of queries, varying query sizes, varying degrees of privacy, and multiple error metrics. Results show that Advanced Partitioning and OBE protocol typically yield the lowest error, demonstrating the superiority of our proposed methods.

Density
Ekin Tire
Koç University · Fen Bilimleri Enstitüsü
2022
00
Yüksek LisansAçık ErişimEN

Keyframe demonstration seeded and Bayesian optimized policy search

Reinforcement learning (RL) is a promising approach to endow robots with skills. However, RL requires many trials to get satisfactory results. Learning from Demonstration (LfD) seeded RL alleviates this problem by learning an initial skill from human demonstrations. Nevertheless, this approach still requires robots to perform a non-trivial amount of trials. In this thesis, we develop an approach to further reduce these for manipulation skills with perceptual goals. Our main contributions are (1) an algorithm to focus the exploration by using the learned relationship between action and perception and a (2) Black-Box RL Policy Search (PS) method that improves upon the popular Policy Improvement with Path Integral (PI²) algorithm, called the Bayesian Optimized PI² (BO-PI²), that uses reward predictive UCB-type exploration. Our underlying LfD framework utilizes a Dynamic Bayesian Network (DBN) learned from keyframe demonstrations to jointly model the action (end-effector pose) and the goal (object-specific perceptual data) of the skill. The action part is used to generate robot trajectories, and the goal part is used to monitor the success of trajectory executions and to create a Partially Observable Markov Reward Model in order to learn rewards. BO-PI$^{2}$ is used to improve the action part of the DBN with trial-and-error using the learned returns. The novelty of BO-PI$^{2}$ comes from its exploration strategy. The coupling between the action and the goal is used to pick the part of the model to focus on to reduce the effort, in a sense to solve the credit attribution problem. After picking the part to focus on, BO-PI$^{2}$ samples trajectories from the action model to get rollouts, which is typical of PS approaches. In addition, BO-PI$^{2}$ uses a Gaussian Process (GP) to learn local returns from these rollouts, which is improved with each executed trajectory. The next samples are selected by utilizing an Upper Confidence Bound (UCB) approach, using the predicted return and uncertainty of the possible candidate points. This is in contrast to random sampling, used in most PS approaches. BO-PI$^{2}$ also utilizes a skill success based termination criteria, using the goal model to monitor success autonomously. We evaluate BO-PI$^{2}$ with expert and non-expert keyframe demonstrations for three skills. In the expert case, the models are perturbed so that the initial skill execution starts from a failure condition. In the non-expert case, we pick skill models that fail to begin with. We test our approach against the current state-of-the-art PI$^{2}$-ES-Cov algorithm using three metrics: (1) skill success rate, (2) total accumulated reward, and (3) number of trials. In both the expert case and the non-expert case, on average, our approach performed better than the baseline on all three metrics. Our results show that utilization of keyframes allows us to focus on failed sub-goals rather than the entire trajectory, and combined with reward predictive exploration strategies, are beneficial to improve RL performance and reduce the number of trials to endow robot arms with real-life manipulation skills.

KeypointsDeep learningMachine learning+1
Onur Berk Töre
Koç University · Fen Bilimleri Enstitüsü
2022
00
Yüksek LisansAçık ErişimEN

Hierarchical clustering attention for unsupervised object-centric representation learning

Extracting object-centric representations from a complex multi-object scene is indeed a crucial milestone for modern neural network architectures to achieve near human level cognition capabilities. Nevertheless, most of the contemporary neural networks that address object-centric representation learning problem require apriori initialization of a fixed set of object describing vectors or cannot manage to handle images of higher resolution. Contrary to long-standing paradigms in the literature, this work proposes Query Breaking Visual Attention (QBVA) module, an efficient and effective building block that introduces a divide and conquer strategy to object-centric representation learning while solving the unsupervised scene segmentation task. QBVA is essentially a stand-alone attention based clustering module that is capable of extracting object-centric representations from a multi-object scene when cascaded into a hierarchical network architecture. QBVA leverages a novel, fully differentiable and non-parametric clustering scheme named Query-Breaking Clustering (QBC) which eliminates the need for initializing a fixed set of clusters and holds the promise to provide dynamic representation for a variable number of objects. We demonstrate that QBVA-Net is indeed a competitive approach to address object-centric representation learning paradigm and prove to be advantageous compared to the state-of-the-art in the sense that it can provide better segmentation performance at the end of the encoder network and theoretically scale up to images of higher resolution.

Graph neural networksHierarchical clusteringArtificial neural networks
Can Küçüksözen
Koç University · Fen Bilimleri Enstitüsü
2022
00
DoktoraAçık ErişimEN

Incentive compatible and provably secure blockchain applications

With the introduction and advent of blockchain technologies, money transfers have become easier with lower costs, no matter how far the recipient is from the sender. Moreover, the blockchain technology has not only affected the contemporary economical model, but also inspired many novel applications that have been thought impossible without the help of a trusted third party (such as governments or well-recognized companies). However, being a newly developed technology, blockchain still lacks enough research to ensure its security (in terms of mining, transactions, or its applications) and incentive compatibility (i.e., the participants are motivated for following the expected protocol steps). Cryptography and game theory are two well-known tools with traditional proving methods to help us for this task. In this thesis, we utilize game theory and cryptography based security models for building secure applications on blockchain. First, we develop a provably secure blockchain application that we name as e-donation, for secure, fair and decentralized online donations. Second, we develop a practical attribute based digital signature scheme. Third, we analyze and improve the security of blockchain by providing an incentive compatible defense solution against the selfish mining attack of Eyal and Sirer (CACM '18) on proof-of-work blockchain. Forth, we develop a network-based simulation tool for both attacks and defenses. Fifth, we propose a framework for better game-theoretical analysis of multi-player mechanisms when the sizes of the coalitions can be bounded by a threshold, via successfully combining threshold security ideas of cryptography and transferable utility of game theory. Sixth, we apply our threshold coalition notions to show the incentive compatibility of our solution to the outsourced incentivized computation problem in blockchain with multiple outsourced parties.

Blockchain systemGame theoryDigital signature
Osman Biçer
Koç University · Fen Bilimleri Enstitüsü
2022
00
Yüksek LisansAçık ErişimEN

Elastic pipeline load balancing for dynamic DNNS

Training of dynamic models is gaining traction in DNNs as it reduces computational and memory requirements of large-scale training. Gradual pruning, one of the prominent approaches for dynamic training, prunes (or sparsifies) the parameters of a model during training. However, one of the side effects of gradual pruning is that sparsification introduces an imbalanced workload across accelerators, which in turn affects the pipeline parallelism efficiency. This work introduces DynPipe which dynamically load balances the stages of the pipeline to offset the negative performance effects of pruning. On top of load balancing dynamic models, DynPipe can dynamically pack work into fewer GPUs, while sustaining performance. DynPipe works on single nodes with multi-GPUs and also on systems with multinodes. Experimental results show that DynPipe can speed up the training up to 5.64% in a single node, and 8.43% in a multi-node setting, over state-of-the art solutions used in training production large language models. DynPipe is available at https://anonymous.4open.science/r/DynPipe-1EC5

Muhammet Abdullah Soytürk
Koç University · Fen Bilimleri Enstitüsü
2023
00
DoktoraAçık ErişimEN

Building, analyzing and interpreting classroom engagement: Apps and machine-learning models for an affordable programming education

Ensuring students remain engaged in the classroom is crucial for their success in any given topic. In programming education, promoting hands-on interactions and demonstrating real-world use cases are effective methods to foster engagement. However, schools located in socio-economically disadvantaged areas often lack adequate digital infrastructure, such as computer laboratories, to support such engagement-building tasks. Nevertheless, utilizing mobile devices can support programming education due to their availability and affordability. Additionally, mobile devices can help augment the tangible materials to use in collaborative programming experiences. In this respect, the first goal of my research is to develop an affordable tangible programming education environment using mobile devices that supports the collaborative work of students. On top of building tools to foster engagement, analyzing and interpreting student engagement is also a critical component of the teaching process. Analyzing classroom engagement requires a multi-component evaluation of affective, behavioral, and cognitive states. Yet, limited research has been conducted to create a multimodal classroom engagement dataset and analysis model. To fill this gap, my second goal is to build a multimodal engagement evaluation tool using a single camera to ease teachers' workload in group activities. Overall my thesis contributes to Human-Computer Interaction (HCI), Artificial Intelligence (AI), and educational technology research areas with six research outputs: 1. An open-source, user-centered, AI-powered, affordable tangible programming environment that was informed by iterative user studies. 2. A set of design considerations for developing paper-based intelligent tangible programming environments which are informed by iterative and classroom-wide user experience studies. 3. A set of curricular activities to help teachers adapt our programming environment into Turkish national curricula. 4. An open-source audio-visual dataset comprising eight-hour-long video recordings of thirty-three students to predict classroom engagement levels using their self-evaluation scores. 5. Multimodal machine-learning models that can address the multi-component definition of engagement. The image models achieved up to 84% test accuracy on person-based engagement level prediction. The real-time video model that can run on streaming videos achieved 71% test accuracy. 6. A student-centric interactive dashboard to help students view their engagement over time and interpret the results of engagement model prediction.

Alpay Sabuncuoğlu
Koç University · Fen Bilimleri Enstitüsü
2023
00
Yüksek LisansAçık ErişimEN

HyperGAN-CLIP: A versatile framework for CLIP-guided imagesynthesis and editing using hypernetworks

Generative Adversarial Networks, particularly StyleGAN and its variants, have shown exceptional capability in generating highly realistic images. However, training these models remains challenging in domains where data is scarce, as it typically requires large datasets. In this thesis work, we introduce a versatile framework that enhances the capabilities of a pre-trained StyleGAN for various tasks, including domain adaptation, reference-guided image synthesis, and text-guided image manipulation even when only a small number of training sample are available. We achieve this by integrating the CLIP space into the generator of StyleGAN using hypernetworks. These hypernetworks introduce dynamic adaptability, enabling the pre-trained StyleGAN to be effectively applied to specific domains described by either a reference image or a textual description. To further improve the alignment between the synthesized images and the target domain, we introduce a CLIP-guided discriminator, ensuring the generation of high-quality images. Notably, our approach shows remarkable flexibility and scalability, enabling text-guided image manipulation with text-free training and seamless style transfer between two images. Through extensive qualitative and quantitative experiments, we validate the robustness and effectiveness of our approach, surpassing existing methods in terms of performance.

Abdul Basıt Anees
Koç University · Fen Bilimleri Enstitüsü
2023
00
Yüksek LisansAçık ErişimEN

Comicverse: Expanding the frontiers of ai in comic books with holistic understanding

Comics are a unique and multimodal medium that conveys stories and ideas through sequential imagery often accompanied by text for dialogue and narration. Comics' elaborate visual language exhibits variations from different authors, cultures, periods, technologies, and artistic styles. Consequently, the computational analysis of comic books requires addressing fundamental challenges in computer vision and natural language processing. In this thesis, I aim to enhance neural comic book understanding by making use of comics' unique multimodal nature and processing comics in a character-centric approach. The primary data source for this thesis is the Golden Age of American Comics due to its public accessibility and abundance of comic series. However, the availability of annotated data is limited. Thus, to achieve my goal, I have adopted a holistic approach composed of four main steps ranging from curating datasets to proposing novel tasks and architectures for comics. The first three steps aim to create a machine-readable comics database by locating comic book panels, identifying their components, and transforming character identities into a dialogue-like structure and the final step uses this database to train a transformer- based model. The first step involves extracting high-quality text data from speech bubbles and narrative box images using OCR models. I decompose comic pages into their constituent components in the second step through detection, segmentation, and association tasks with a refined Multi-Task Learning (MTL) model. Detection involves identifying panels, speech bubbles, narrative boxes, character faces, and bodies. Segmentation focuses on isolating speech bubbles and panels, while the association task involves linking speech bubbles with character faces and bodies. In the third step, I utilize the paired character faces and bodies obtained from the previous stage to create character instances and, subsequently, reidentify and track these instances across sequential panels. In the final step of my thesis, I propose a multimodal framework by introducing the ComicBERT model, which exploits the abovementioned structure. Cloze-style tasks were used to evaluate ComicBERT's contextual understanding capabilities. Furthermore, I propose a new task called Scene-Cloze, which predicts the next panel given n previous panels as context. As a result, my approach achieves a new state-of-the-art performance in Text-Cloze and Visual-Cloze tasks with accuracies of 69.5% and 77.1%, respectively, thus getting closer to the human baseline. Overall, the highlights of my contributions are as follows: 1. I curated and shared COMICS Text+ Dataset with over two million transcrip- tions of textboxes from the golden age of comics. In addition, I open-sourced the text detection and recognition models that are fine-tuned for the task and datasets used in their training. 2. I refined a MTL framework for detection, segmentation, and association tasks and achieved SOTA results in comic character face and body-to-speech bubble association tasks. 3. I proposed a novel Identity-Aware Semi-Supervised Learning for Comic Character Re-Identification framework to generate unified and identity-aligned comic character embeddings and identity representations. Furthermore, I generated two new datasets: the Comic Character Instances Dataset, encompassing over a million character instances used in the self-supervision phase, and the Comic Sequence Identity Dataset, containing annotations of identities within sets of four consecutive comic panels used in semi-supervision phase. 4. I introduced the multimodal Comicsformer, a transformer-encoder architecture capable of processing sequential panels and their constituents. It serves as the backbone for the Masked Comic Modeling (MCM) task, a novel self- supervised pre-training strategy for comics, resulting in ComicBERT, a potential foundation model for golden age comics. ComicBERT achieves SOTA performance in cloze-style tasks, particularly in text-cloze and visual-cloze tasks, approaching human-level comprehension.

Computer visionNatural language processingMachine vision+3
Gürkan Soykan
Koç University · Fen Bilimleri Enstitüsü
2023
00
Yüksek LisansAçık ErişimEN

Investigating the missing pieces of sensorimotor reinforcement learning agents for autonomous driving

Reinforcement Learning (RL) has the potential to surpass human capabilities in self-driving without needing any expert supervision. Despite its promise, the state-of-the-art in sensorimotor self-driving is dominated by imitation learning methods due to the inherent challenges of RL algorithms. Nonetheless, RL agents are able to discover highly successful policies when provided with privileged ground truth representations of the environment. In this work, we investigate what separates privileged RL agents from sensorimotor agents for urban driving in order to bridge the gap between the two. We propose vision-based deep learning models to approximate the privileged representations from sensor data. In particular, we identify aspects of state representation that are crucial for the success of the RL agent such as desired route generation and traffic light prediction, and propose solutions to gradually remove privileged information for each with the existing computer vision approaches. Through rigorous evaluation on the CARLA simulation, we shed light on the significance of the state representation in RL for autonomous driving and outline unresolved challenges for future research.

Ege Onat Özsüer
Koç University · Fen Bilimleri Enstitüsü
2023
00
DoktoraAçık ErişimEN

Advancing toward temporal and commonsense reasoning in vision-language learning

Humans learn to ground language to the world through experience, primarily visual observations. Devising natural language processing (NLP) approaches that can reason in a similar sense to humans is a long-standing objective of the artificial intelligence community. Recently, transformer models exhibited remarkable performance on numerous NLP tasks. This is followed by breakthroughs in vision-language (V&L) tasks, like image captioning and visual question answering, which require connecting language to the visual world. These successes of transformer models encouraged the V&L community to pursue more challenging directions, most notably temporal and commonsense reasoning. This thesis focuses on V&L problems that require either temporal reasoning, commonsense reasoning, or both simultaneously. Temporal reasoning is the ability to reason over time. In the context of V&L, this means going beyond static images, i.e., processing videos. Commonsense reasoning requires capturing the implicit general knowledge about the world surrounding us and making an accurate judgment using this knowledge within a particular context. This thesis comprises four distinct studies that connect language and vision by exploring various aspects of temporal and commonsense reasoning. Before advancing to these challenging directions, (i) we first focus on the localization stage: We experiment with a model that enables systematic evaluation of how language-conditioning should affect the bottom-up and the top-down visual processing branches. We show that conditioning the bottom-up branch on language is crucial to ground visual concepts like colors and object categories. (ii) Next, we investigate whether the existing video-language models thrive in answering questions about complex dynamic scenes. We choose the CRAFT benchmark as our test bed and show that the state-of-the-art video language models fall behind human performance by a large margin, failing to process dynamic scenes proficiently. (iii) In the third study, we develop a zero-shot video-language evaluation benchmark to evaluate the language understanding abilities of pretrained video-language models. Our experiments reveal that the current video-language models are no better than the vision-language models, processing static images as input in processing daily dynamic actions. (iv) In the last study, we work on a figurative language understanding problem called euphemism detection. Euphemisms tone down expressions about sensitive or unpleasant issues. The ambiguous nature of euphemistic terms makes it challenging to detect their actual meaning within a context where commonsense knowledge and reasoning are necessities. We show that incorporating additional textual and visual knowledge in low-resource settings is beneficial to detect euphemistic terms. Nonetheless, our findings on these four studies still demonstrate a substantial gap between current V&L models' abilities and human cognition.

İlker Kesen
Koç University · Fen Bilimleri Enstitüsü
2023
00
Yüksek LisansAçık ErişimEN

RbA: Segmenting unknown regions rejected by all using mask classifiers

Fine-grained visual semantic understanding is essential for autonomous driving and many other computer vision tasks. To this end, segmentation tasks have been witnessing rapid advancements as a result of the growing number of benchmarks proposed. However, existing segmentation benchmarks generally assume a fixed set of semantic categories. Consequently, the development of segmentation methods has been centered around this assumption, while little attention has been poured into handling novel or out-of-distribution (OoD) samples that can potentially be encountered in real-life scenarios. This poses an issue for an autonomous vehicle, as it is crucial to identify unknown objects so that a safety warning can be issued to avoid disastrous consequences in case of failure. As a result, the task of OoD segmentation has been addressed separately, leading to the emergence of methods that adapt to the existing segmentation methods and disregard the performance of the main segmentation tasks. Additionally, these methods also generally suffer from a lack of smoothness and objectness in their predicted anomaly maps due to the reliance on models that follow the per-pixel classification paradigm. In this thesis, we explore the potential of region-level classification models for unknown segmentation as a unified architecture with an inherent ability to express uncertainty. We show that the object queries in mask classification models tend to behave like one \vs all classifiers. Based on this finding, we propose a novel outlier scoring function called Rejected by All (RbA) by defining the event of being an outlier as being rejected by all known classes. We also propose an objective that optimizes this proposed score for boosting the unknown segmentation performance using pseudo-outlier data without hurting the closed-set performance. RbA performs well under high domain shifts and is capable of separating sources of uncertainty, such as at the boundaries, due to known class ambiguity. We evaluate RbA on several unknown segmentation benchmarks and show that it achieves state-of-the-art performance with significant margins compared to previous pixel-level unknown segmentation methods. We report extensive ablation experiments that validate the effectiveness of RbA.

Nazır Nayal
Koç University · Fen Bilimleri Enstitüsü
2023
00
Yüksek LisansAçık ErişimEN

A multilayer perceptron framework for sparse multiple kernel learning

Advancing cancer research and improving patient care heavily rely on understanding how biological mechanisms change as cancer progresses. Identifying the active biological processes within tumors is crucial for developing targeted treatments and potential biomarkers for early detection and prognosis. In this thesis, we focused on distinguishing between early-stage and late-stage cancers using gene expression profiles, aiming to develop a powerful and interpretable model. Machine learning methods have shown promise in cancer research by enabling the analysis of vast genomic and clinical data to uncover hidden patterns and predictive features. Our work centered around harnessing the capabilities of a multilayer perceptron (MLP) model to create a sparse solution within the context of multiple kernel learning (MKL). This enabled our model to proficiently differentiate between early-stage and late-stage cancers based on gene expression profiles. Our model not only offered high predictive performance but also provided valuable insights into the crucial genes and pathways driving cancer progression. To evaluate our MLP model, we benchmarked it against three well-established machine learning algorithms: random forest, support vector machine, and MKL. Remarkably, our model consistently achieved better or comparable predictive performance, as measured by the area under the receiver operating characteristic curve, across 15 cancer cohorts. The findings demonstrate that our proposed MLP model effectively identifies critical genes and pathways driving cancer progression, offering valuable insights into early-stage and late-stage cancer classification.

Binnur Şahin
Koç University · Fen Bilimleri Enstitüsü
2023
00
Yüksek LisansAçık ErişimEN

HAİSTA-NET: Dikkat yoluyla insan destekli obje segmentasyonu

Instance segmentation is a fundamental computer vision task with a wide range of applications. Some instance segmentation tasks such as medical image analysis, and image/video editing require high levels of precision. However, this precision is often beyond the reach of what even state-of-the-art, fully automated instance segmentation algorithms can deliver. The performance gap becomes particularly prohibitive for small and complex objects. Practitioners typically resort to fully manual annotation, which can be a laborious process. In order to overcome this problem, we propose a novel approach to enable more precise predictions and generate higher-quality segmentation masks for high-curvature, complex and small-scale objects. Our human-assisted segmentation model, HAISTA-NET, augments the existing Strong Mask R-CNN network to incorporate human-specified partial boundaries. We also present a dataset of hand-drawn partial object boundaries, which we refer to as "human attention maps." In addition, the Partial Sketch Object Boundaries (PSOB) dataset contains hand-drawn partial object boundaries which represent curvatures of an object's ground truth mask with several pixels. Through extensive evaluations, we show that HAISTA-NET outperforms state-of-the art methods such as Mask R-CNN, Strong Mask R-CNN, and Mask2Former, achieving respective increases of +36.7, +29.6, and +26.5 points in AP-Mask metrics for these three models. Our novel approach sets a baseline for future human-aided deep learning models by combining fully automated and interactive instance segmentation architectures.

Muhammed Korkmaz
Koç University · Fen Bilimleri Enstitüsü
2023
00
DoktoraAçık ErişimEN

FLAGS framework and decentralized federated learning under device volatility

Federated Learning (FL) has become a key choice for distributed machine learning. Initially focused on centralized aggregation, recent works in FL have emphasized greater decentralization supported by standardization of serverless interaction in the next-generation communication networks. However, the diversity of devices, data distributions, and communication settings, compounded by dynamic operating conditions, result in multiple challenges for Decentralized FL (DFL). There have been various approaches to DFL, from utilizing intermediate edge servers to fully device-to-device approaches. In decentralized settings, communication cost and learning performance are usually assessed together and certain trade-offs are made based on scenarios. However, there is a lack of existing work on comparing DFL approaches in an apples-to-apples manner in a multitude of scenarios and operating conditions. To bridge this gap between methods and their comparative analysis, we design and develop the Federated Learning Algorithms Simulation (FLAGS) Framework. One important challenge we noticed that most DFL methods struggle with is the extreme fluctuations in device availability, especially for purely decentralized approaches. This \textbf{device volatility} leads to poor learning performance. To address this issue, we investigate the effects of neighborhood selection, memory and multi-hop information passing on DFL performance. We introduce a fully decentralized FL approach that can operate under realistic and highly volatile device participation settings. The key contributions of this thesis are as follows: (i) development of a lightweight FL framework for benchmarking a large plethora of methods, (ii) analysis and comparison of multiple FL methods with this framework under multiple operating conditions, (iii) empirical analysis of various node selection strategies under heavy device volatility, and (iv) utilizing memory and relayed communication to enhance device-to-device FL by developing a novel algorithm that can operate under realistic operating conditions and heavy device volatility. Federated Learning supports a wide variety of node interactions and autonomous operations across the network edge. With the aim to encompass this multi-faceted heterogeneity, the FLAGS framework was proposed and developed as a lightweight FL implementation and testing platform. FLAGS framework allows for a wide range of device behaviors and cooperation mechanisms, enabling rapid testing of multiple FL algorithms. FLAGS's built-in features allow it to subject existing and novel FL algorithms to a wide range of data distributions, simulating the nodes with multiple neural networks as well as participation conditions ranging from homogeneous to highly volatile. Different network tiers and communication mechanisms enable various FL algorithms to be configured by employing various combinations of the aforementioned factors. In order to consolidate this very extensive FL landscape and offer an objective analysis of the major FL algorithms, comprehensive cross-evaluations for a wide range of operating conditions have also been conducted. Starting with the three foundational FL algorithms, including Hierarchical FL (HFL), Decentralized FL (DFL), and Gossip FL (GFL), this work evaluates six derived algorithms ranging from fully centralized to fully decentralized. The experiments indicate that fully decentralized FL algorithms achieve comparable accuracy under multiple operating conditions, including asynchronous aggregation and the presence of stragglers. Furthermore, DFL can also operate in noisy environments and with a comparably higher local update rate. However, the impact of extremely skewed data distributions on DFL is much more adverse than on centralized variants. The analysis of the cross-evaluation indicates that DFL performance is considerably impacted by node participation. This part of the thesis focuses on improving DFL performance under realistic and volatile device behavior. Node selection in various forms has been experimented with to improve both communication efficiency and convergence rate. We experimented with multiple node selection mechanisms and also proposed and evaluated a time-varying parameterized node selection method for DFL employing validation accuracy and its per-round change. The mentioned criteria are evaluated using both hard and stochastic/soft selection on sparse networks. The results indicate that the bias associated with node selection adversely impacts performance as training progresses, and a uniform random selection is preferable under extremely limited participation conditions. Continuing with volatile conditions, we investigate and propose mechanisms to improve DFL operating on sparse graphs in the presence of stragglers and non-participating nodes. We first propose two algorithms: Memory-Assisted DFL (MA_DFL) and Augmented-Graph Assisted DFL (AG_DFL). These algorithms employ memory and selective relaying to improve DFL performance. Both algorithms outperform the baseline DFL and gossip interaction for volatile node participation. Then, we propose a hybrid of these two algorithms, Memory and Augmented-Graph Assisted DFL (MAG_DFL), that employs memory and graph augmentation to improve the performance of DFL under highly volatile devices and extreme data conditions. The research conducted in this thesis evaluates the multi-faceted challenges to the DFL operation in volatile conditions and proposes mechanisms to improve its performance. Our work indicates that DFL holds the potential to assist learning operations distributed across the edge network. It may be used to augment the FL in the presence of costly upstream communication or limited connectivity. However, node density has a major impact on DFL, and sparse networks, along with volatile device behavior and non-IID distributions, tend to reduce its convergence rate. The enhanced neighborhood interaction and intelligent use of local information has the potential to improve DFL performance under such adverse conditions based on the presented results. The analysis, algorithms and results presented in this thesis pave the way for additional developments and more practical applications of DFL in the next-generation communication networks.

Deep learningMachine learningSwarm intelligence+1
Ahnaf Hannan Lodhı
Koç University · Fen Bilimleri Enstitüsü
2023
00
Yüksek LisansAçık ErişimEN

Autonomous execution for multi-GPU systems: CPU-free blueprint and compiler support

As multi-GPU systems become more prolific in the field of supercomputing, scientific applications are adapted and scaled up to take advantage of the highly parallel accelerators for increased performance. However, the traditional model of GPU programming leaves much to be desired in multi-GPU settings, wherein communication among devices - one of the largest points of contention and bottlenecks in scientific applications - is controlled by the CPU. This kind of one-sided control leads to undue latencies incurred by the constant back-and-forth of synchronization and API calls between the host and devices, and harms application scaling as the number of GPUs grows. This work first proposes the fully autonomous CPU-Free execution model for multi-GPU applications that completely excludes the involvement of the CPU beyond the initial kernel launch. We systematically combine several techniques such as persistent kernels, thread block specialization, and GPU-initiated communication and synchronization to significantly reduce host-incurred latencies and facilitate further optimizations. We benchmark our proposed model on a broadly used iterative solver, 2D/3D Jacobi Stencil and improve 3D stencil communication latency by 58.8% compared to CPU-controlled baselines on 8 NVIDIA A100 GPUs. The second part of this work adds compiler support to easily write performant CPU Free code in high-level Python by extending the DaCe framework with GPU-centric communication intrinsics. We compare automatically generated CPU-Free code to existing distributed facilities in DaCe and observe over 96\% performance improvement in Stencil benchmarks.

Javıd Baydamırlı
Koç University · Fen Bilimleri Enstitüsü
2023
00
Yüksek LisansAçık ErişimEN

Localizing knowledge in large language model representations

Large language models (LLMs) are very proficient in NLP tasks. In the first part of this work, we evaluate the performance of LLMs on the task of finding the locations of characters inside a long narrative. The objective of the task is to generate the correct answer when the input is a piece of a narrative followed by a question asking the location of a character. For the evaluation of the task, we generate two new datasets by annotating the characters and their locations in the narratives: Andersen and Persuasion. We show that the LLM performance is not satisfactory on these datasets when compared to the simple baseline we designed that does not use machine learning. We also experiment with in-context learning to improve the performance and report results. Moreover, we address the problem that the LLMs are limited by the bounded context length. We hypothesize that if we localize the character-location relation information among the activations inside an LLM, we can store those activations and inject them into other models that are run with a different prompt so that the LLM can answer the questions about the information that was carried from another prompt, even though the character and location relation is not mentioned explicitly in the current prompt. We develop five different techniques to localize the character-location relation information occurring in the LLMs: Moving and adding LLM activations to other prompts, adding noise to LLM activations, checking cosine similarity between LLM activations, editing LLM activations, and visualizing attention scores during answer generation. We report the observations we made using these techniques.

Batuhan Özyurt
Koç University · Fen Bilimleri Enstitüsü
2023
00
Yüksek LisansAçık ErişimEN

Investigating the potential of incorporating protein language models (pLMs) into ML/DL approaches for enhanced prediction of allosteric sites in proteins

Allosteri, proteinin bir bölgesindeki değişikliğin, mesela başka bir moleküle bağlanmanın, proteinin uzak bir bölgesini etkilediği süreç olarak tanımlanabilir. Allosteri protein fonksiyonu üzerindeki önemli etkisi sebebiyle ilaç geliştirme alanında önemli bir odak noktasıdır. Allosterik ilaçlar proteinleri aktive veya inhibe edebilir, allosterik olmayan ilaçlara göre avantajlar sunar. Bununla birlikte, allosterik bölgelerin tanımlanması zorlu bir iştir. Geçmişte allosterik bölgeleri tahmin etmek için Normal Mod Analizi (NMA), Moleküler Dinamik (MD) ve Makine Öğrenimi (MÖ) gibi hem statik cep özelliklerini hem de proteinlerin dinamiklerini kullanan çeşitli hesaplama teknikleri geliştirilmi olmakla birlikte bu yöntemlerin performansının daha da geliştirilmesi gerekmektedir. Bu araştırmada, pDM'lerin (örneğin, ProtTrans pDM ailesinden BERT mimarisine dayalı ProtBERT'in) allosterik kalıntıların tahminini iyilrştirmek için Protein Dil Modellerini (pDM'ler), MÖ ve/veya DÖ yaklaşımlarıyla birlikte kullanılma potansiyelini araştırılıyor. Tezde, amino asitler arasındaki mekansal ilişkiyi etkili bir şekilde öğrenerek, sonuçta allosterik alanların/ceplerin tanımlanmasını hedeflenmektedir. ProtBERT-BFD (ProtTrans), test veri kümesinde %61,54'lük bir F1 puanıyla allosterik kalıntıları tahmin eden protein dizilerinin Allosterik Veri Kümesine (AVK) göre ince ayar yapılmıştır. XGBoost, SVM, AutoML ve GNN'ler dahil olmak üzere çeşitli MÖ ve DÖ yaklaşımlarından yararlanılmıştır, İnce ayarlı pDM özelliklerinin dahil edilmesiyle, yukarıda belirtilen yaklaşımların tümü, allosterik bölgelerin tahmin performansını önceki çalışmalara göre önemli bir farkla artırdığı bulunmuştur. Bu çalışmada en yüksek performansa sahip model olan XGBoost, ince ayarlı ProtBERT'ten çıkarılan özellikleri FPocket tarafından çıkarılan cep özellikleriyle birleştirerek sonuçları iyileştiriyor ve allosterik cepler/bölgeler için %75,76'lık bir F1 puanı erişmektedir. Bilinen allosterik bölgelere sahip proteinler üzerinde örnek çalışmaların yanı sıra, farklı proteinler üzerindeki yeni allosterik bölgeleri tahmin etmek için de çalışmalar yapılmıştır.

Moaaz Ur Rehman Azhar Khokhar
Koç University · Fen Bilimleri Enstitüsü
2023
00
Yüksek LisansAçık ErişimEN

Unsupervised multi-object discovery and tracking using memory-augmented slot attention

Learning object-centric representations from static images is a promising research direction in the field of deep learning. However, adapting this approach to videos poses certain challenges due to the necessity of capturing the temporal dynamics of video content. Recent works have made significant progress in object discovery within synthetic video datasets. Nevertheless, these works do not fully exploit the motion of objects in videos and temporal cues. In this thesis, we aim to enhance the performance of object-centric representation learning on video frames by using the temporal information more carefully. To achieve this goal, we propose a new unsupervised learning method that utilizes a memory-augmented slot attention model for multi-object discovery and tracking. The key component of our approach is the integration of memory slots, which store information from past video frames, alongside object slots into the learning architecture. Object slots simultaneously attend to both memory slots for information from past frames and the current image input. Training memory slots requires longer video sequences. However, the most suitable learning structures for this, recurrent neural networks (RNNs), are not very effective in learning long-term temporal data due to the issues of exploding and/or vanishing gradients. To train our memory-augmented model more effectively on long video sequences, we employ truncated back-propagation through time. Experiments conducted on synthetic yet realistic video datasets have yielded promising results, indicating that memory slots significantly improve multi-object tracking and object segmentation performance. Our fully unsupervised learning method contributes to the problem of object-centric representation learning in videos and opens up new possibilities in this field.

Ahmed Imam Shah
Koç University · Fen Bilimleri Enstitüsü
2023
00
Yüksek LisansAçık ErişimEN

Self-collision aware reaching and pose control in large workspaces using deep reinforcement learning

Reaching, pose control, and inverse kinematics are fundamental robotic manipulator tasks that underpin other tasks and as such, there is a vast body of related literature from various fields. Control, planning, and more recently learning fields are among the main ones. Traditional control algorithms are prone to failure around singularities and joint limits and do not naturally handle self-collisions. Planning methods are not fast enough for reactive behaviours and require additional infrastructure, including controllers, to work. Learning-based methods have emerged to tackle these issues. However, most of them do not handle arbitrary initial and target poses, ignore self-collisions, do not include orientation information in their targets, work in small workspaces and evaluate themselves with coarse success metrics. In this thesis, we introduce a novel hybrid approach that combines Pseudo-inverse control (PinvC) and model-free reinforcement learning (RL), including state space and reward function design, to fill these gaps in the context of reaching, pose control and inverse kinematics. PinvC already calculates joint velocities given desired task-space (e.g. the end-effector pose) velocities and only requires the kinematic structure of the robot. PinvC is mostly reliable away from joint limits, singularities and when individual links are not prone to collisions. The main idea behind our approach is to use RL to handle these situations. Towards this end, we design a novel state space and reward functions. Our reward function aims to minimize position (reaching task) or pose (inverse kinematics and pose control tasks) errors, reduce self-collisions and reduce joint velocities near the target. Furthermore, we develop a curriculum learning methodology to aid learning. Lastly, we introduce a simple modification, which we call "switching" to further improve task performance. We evaluate our approach with four simulated robots for various problem settings and compare it against traditional and learning-based approaches. Our results show that our approach decidedly outperforms the baselines in terms of mean error, success rates at various thresholds and terminal speed for reaching tasks. In addition, we reduced the number of self-collisions across all the scenarios. Our approach achieved better results when orientation was included, but none of the methods performed very well, especially the learning baselines. We note that the learning-based methods in the literature almost always ignore orientation. As a result, we comprehensively discuss the reasons for orientation failure and potential remedies.

Industrial robots
Tumuçin Bal
Koç University · Fen Bilimleri Enstitüsü
2023
00
Yüksek LisansAçık ErişimEN

Audio-driven image generation and editing with pretrained diffusion models

We are witnessing a revolution in conditional image synthesis with the recent success of large scale text-to-image generation methods. This success also opens up new opportunities in controlling the generation and editing process using multi-modal input. While spatial control using cues such as depth, sketch, and other images has attracted a lot of research, we argue that another equally effective modality is audio since sound and sight are two main components of human perception. Hence, in this thesis we propose SonicDiffusion to enable audio-conditioning in large scale image diffusion models. Our method first maps features obtained from audio clips to tokens that can be injected into the diffusion model in a fashion similar to text tokens. We introduce additional audio-image cross attention layers which we finetune while freezing the weights of the original layers of the diffusion model. In addition to audio conditioned image generation, our method can also be utilized in conjuction with diffusion based editing methods to enable audio conditioned image editing. We demonstrate our method on a wide range of audio and image datasets. We perform extensive comparisons with recent methods and show favorable performance.

Burak Can Biner
Koç University · Fen Bilimleri Enstitüsü
2024
00
Yüksek LisansAçık ErişimEN

Interpretable cancer stage classification using sparse bayesian neural networks

Cancer requires an in-depth exploration of its molecular characteristics. For targeted treatment strategies, distinguishing cancer stages is essential. This thesis has focused on the differentiation of early- and late-stage cancers using gene expression profiles. With the integration of computational techniques into medical research, machine learning models excel in this task, offering insights into biological mechanisms. To harness these insights, we proposed a novel approach, which is Bayesian Neural Networks (BNNs) with sparsity-inducing priors. The proposed sparse BNNs are designed to deliver high predictive performance in identifying cancer stages while maintaining a high level of interpretability. To evaluate our sparse BNN models, we benchmarked them against three machine learning algorithms across 15 different cancer cohorts. The results of our study revealed that our sparse BNN models achieve predictive performances comparable to traditional benchmark models. Additionally, we addressed the black-box issue of neural networks in medicine, which obscures which input features are crucial for predictions, a serious issue in decision-making with significant implications. To address this issue, our primary contribution has been the development of a novel BNN architecture that considerably enhances data interpretability. In our approach, we have integrated three types of sparsity inducing priors, namely, Laplace, Student's t, and Spike-and-Slab. Each prior has a mean of zero and low variance, promoting a reduction in connections and thus enabling a focused feature selection process. This methodology allows us to identify and concentrate on the most influential gene expressions. Our analysis revealed that sparse BNNs show a distinct preference for specific gene sets. In conclusion, the development of sparse BNNs offers a biologically informative and interpretative tool, enhancing the field of cancer research by shedding light on key gene pathways and significantly improving the process of cancer staging.

Hazal Hasret Yurdakul
Koç University · Fen Bilimleri Enstitüsü
2024
00
Yüksek LisansAçık ErişimEN

Cha and core discovery on intel chips and generating optimized thread binding

In modern multi-core architectures with distributed directory-based cache coherence, each memory address is overseen by a distributed directory unit, known as a Caching/Home Agent (CHA), that monitors cache line state and location. Neither the CHA nor core locations in a processor are directly exposed to the programmer. In this work, we firstly analyze and compare the methodologies for uncovering both the CHA and core topology of Intel Xeon Scalable processors, as well as the methods to reveal the mapping of memory addresses to CHAs. Leveraging the topology and the address mapping information, we investigate the impact of spatial proximity between communicating cores and CHAs on application performance, and propose a thread mapping heuristic that assigns threads to cores by considering cache coherence traffic. We expect our heuristic to achieve significant performance gains on applications with high amount of on-chip cache coherence traffic due to high percentage of shared written data. We evaluated our heuristic on applications that exhibit high amount of on-chip communication traffic. The heuristic achieves up to 5.6% speedup over compact placement on merge-based SpMV application, up to 8% with an average of around 4.4% on Barnes application, around 25% for Fluidanimate application to simulate 60 frame per second, and lastly approximately 6% for LU across different matrices. We also prove the improved performance is in fact related to reduced on-chip traffic on the mesh.

Aydın Özcan
Koç University · Fen Bilimleri Enstitüsü
2024
00
Yüksek LisansAçık ErişimEN

An automated framework for concurrent graph processing on GPU

GPUs have become the bleeding edge in high-performance systems used for artificial intelligence and scientific computations. Running parallel jobs and algorithms on GPUs yields results that are orders of magnitude faster than running the same jobs on CPUs that employ fewer cores. A significant body of computational tasks incorporates graph traversals and computations associated with traversal of large graphs. These tasks are either graph traversal tasks in nature or are simplified or transformed into graph traversal tasks. Kernel fusion has been proposed in several bodies of research as a means of fusing computation cores (kernels) in parallel jobs so that the performance can be optimized. Specifically, kernel fusion is used to fuse accesses to same graph nodes and corresponding regions of memory by multiple parallel jobs so that memory caches are more efficiently utilized, resulting in performance improvements. As kernel fusion's fundamental idea is sharing resource access to increase utilization by multiple jobs, it is even better suited for GPUs that run hundreds of jobs in parallel than CPUs that run a few jobs in parallel. However, kernel fusion is a computationally expensive operation that incorporates control flow changes and manipulation of the order of execution for computational jobs. As such, kernel fusion is harder to utilize on GPUs compared to CPUs. This work introduces a streamlined kernel fusion framework for concurrent graph processing on GPUs. The framework enables definition and implementation of graph jobs that can be executed in parallel in GPUs, and are automatically fused by the kernel fusion algorithm implemented in the framework, to achieve better performance. The framework also introduces novel data handling structures, as well as a meta-compiler that enables static polymorphism for running multiple different jobs as a job queue on the GPU, without performance hits that are associated with typical polymorphism. The framework is then extensively evaluated by defining four common graph jobs (BFS, SSSP, PageRank, Label Propagation) and running up to 200 parallel instances of these jobs in homogeneous and heterogeneous manners, with and without kernel fusion. The evaluations show that kernel fusion provides about 5% to 10% performance improvement when job parallelism is reasonably high (i.e., more than 10 parallel jobs).

Mandana Bagherımarzıjaranı
Koç University · Fen Bilimleri Enstitüsü
2024
00
Yüksek LisansAçık ErişimEN

Distributed multi-party fair exchange

Over the years, researchers focused on either two-party cases or multi-party computation securely, fairly and optimistically. Most of the studies in the literature require a trusted third party (TTP) to be present for fairness either at every step of the protocol or only optimistically intervene in case of malicious actions. However, having a single TTP creates a single point of failure (in terms of security) in the protocol. To distribute this responsibility, literature has different solutions. Blockchain is one of the most used solutions for decentralizing exchange protocols and multi-party computations. Secret sharing is another procedure used for distributing responsibility, however, it is mostly used for 2-party settings. As far as we know, there is no optimistic multi-party fair exchange solution using multiple TTPs in the literature. In this work, we present a Distributed Multi-Party Fair Exchange Protocol, building on top of a previous research [Alper and Küpçü, 2021], by distributing the responsibility of a single TTP to multiple (m) TTPs. We are using secret sharing to distribute the decryption share used for fairness to TTPs. Even with a malicious subset of TTPs, the protocol provides a fair result, as long as there are threshold-many honest TTPs. Even though the performance of the protocol is worse (slower, as expected) than the previous study, it carries the importance of being the first optimistic multi-party fair exchange protocol with multiple trusted third parties in the literature.

Aybüke Buket Akgül
Koç University · Fen Bilimleri Enstitüsü
2024
00
Yüksek LisansAçık ErişimEN

O1O: Grouping of known classes to identify odd-one-out

Object detection methods trained on a fixed set of known classes struggle to detect objects belonging to unknown classes in real-world scenarios. Open-world methodologies have emerged in recent years as a solution for the limitations of closed-set approaches. The main goal of open-world object detection is to detect and identify novelties while maintaining closed-set abilities. One common approach involves incorporating approximate supervision with pseudo-labels corresponding to candidate locations of objects, typically obtained in a class-agnostic manner. While previous attempts mainly rely on the appearance of objects, we propose that geometric cues provide a better solution as the source of pseudo-labels. By considering not just how objects look but also their shapes and relative locations, we aim to improve the system's ability to detect unfamiliar objects. Although additional supervision from pseudo-labels improves unknown object detection, it also introduces confusion for known classes. We observed a notable decline in the model's performance for detecting known objects in the presence of noisy pseudo-labels. To address this problem, we drew inspiration from human cognitive science. Studies about how humans mentally represent objects found that humans group objects based on their common attributes, which then helps to compare and identify the different ones given a group of objects. We applied a similar concept by organizing known object classes into a smaller set of superclasses by learning discriminative superclass representations. By doing so, our model can identify similarities between classes within a superclass, thereby facilitating the detection of unknown classes through an odd-one-out scoring mechanism. Our experiments on open-world detection benchmarks demonstrate significant improvements in unknown recall consistently across all tasks. Crucially, we achieve this without compromising known performance, thanks to better partitioning of the feature space with superclasses.

Mısra Yavuz
Koç University · Fen Bilimleri Enstitüsü
2024
00
DoktoraAçık ErişimEN

Multiview contrastive autoencoder-transformer approach for protein-protein interface representation: Unveiling biological and functional insights

Protein-protein interactions (PPIs) play pivotal roles in various biological processes, orchestrating cellular functions essential for life. The interfaces where these interactions occur serve as focal points for understanding the mechanisms underlying disease pathways. Accurate representation of these interfaces is crucial for deciphering their biological significance and designing therapeutic interventions. This thesis introduces a novel approach for representing protein-protein interfaces using a graph-based multiview contrastive autoencoder combined with a transformer, which learns representations from a large dataset. Comprehensive evaluations demonstrate the method's effectiveness in capturing the structural and functional characteristics of protein-protein interfaces. The learned representations are applied to tasks such as biological relevance prediction, biological vs. crystal classification, and Gene Ontology term prediction, showcasing their versatility and utility in understanding PPIs. By integrating explainable AI techniques, key features contributing to model predictions are identified, enhancing the interpretability of the results. A detailed case study illustrates the practical application of these methods, highlighting their potential to provide actionable insights for biological research and drug discovery. Overall, this thesis advances the understanding of protein-protein interactions by providing interpretable representations that capture the complex structural and functional characteristics of interfaces, thereby facilitating biomedical studies and therapeutic developments.

Damla Övek
Koç University · Fen Bilimleri Enstitüsü
2024
00
Yüksek LisansAçık ErişimEN

Grounding language in motor space: Exploring robot action learning and control from proprioception

Language development, particularly in its early stages, is deeply correlated with sensory-motor experiences. For instance, babies develop progressively via unsupervised exploration and incremental learning, such as labeling the action of "walking" by first discovering to move their legs via trial and error. Drawing inspiration from this developmental process, our study explores robot action learning by trying to map linguistic meaning onto non-linguistic experiences in autonomous agents, specifically for a 7-DoF robot arm. While current grounded language learning (GLL) in robotics emphasizes visual grounding, our focus is on grounding language in a robot's internal motor space. We investigate this through two key aspects: Robot Action Classification and Language-Guided Robot Control, both within a "Blind Robot" scenario by relying solely on proprioceptive information without any visual input in pixel space. In Robot Action Classification, we enable robots to understand and categorize their actions using internal sensory data by leveraging Self-Supervised Learning (SSL) through pretraining an Action Decoder for better state representation. Our SSL-based approach significantly surpasses other baselines, particularly in scenarios with limited data. Conversely, Language-Guided Robot Control poses a greater challenge by requiring robots to follow natural language instructions, interpret linguistic commands, generate a sequence of actions, and continuously interact with the environment. To achieve that, we utilize another Action Decoder pre-trained on sensory state data and then fine-tune it alongside a Large Language Model (LLM) for better linguistic reasoning abilities. This integration enables the robot arm to execute language-guided manipulation tasks in real time. We validated our approach using the popular CALVIN Benchmark, where our methodology based on SSL significantly outperformed traditional architectures, particularly in low-data scenarios on action classification. Moreover, in the instruction following tasks, our Action Decoder-based framework achieved on-par results with large Vision-Language Models (VLMs) in the CALVIN table-top environment. Our results underscore the importance of robust state representations and the potential of the robot's internal motor space for learning embodied tasks.

Emre Can Acikgoz
Koç University · Fen Bilimleri Enstitüsü
2024
00
Yüksek LisansAçık ErişimEN

OPTKIT: Unleash the Green Potential of Your CPU Code

The analysis of energy consumption in software has traditionally taken a back seat to performance analysis. However, increasing environmental concerns, power and cooling limitations in expanding HPC systems and data centers, and the need for energy efficiency in power-constrained devices such as mobile phones, tablets, laptops, and IoT devices have brought significant focus to this area. While performance analysis has a wide range of tools, those specifically designed for software energy efficiency are scarce. This study introduces the Optimizer Toolkit (optkit), a dedicated library and toolset designed for the energy analysis and optimization of software runtime. optkit utilizes perf_event_open kernel call for PMU and RAPL, and msr-safe for cpu frequency operations. Inside the tool, there are Query* classes where the user can query anything regarding cpu, energy, and performance before taking any action. It includes useful utility tools that help users streamline the optimization process. In this work, we leverage optkit to evaluate the potential for reducing a system's energy consumption by strong-scaling core numbers in a given application. We then use these data to further reduce energy consumption through co-scheduling applications. In another use-case, we investigate the impact of CPU frequency on energy consumption and develop a methodology to dynamically detect the optimal frequency at runtime and adjusts frequencies based on execution phases to maximize energy savings. We find that not all applications scale linearly, and allocating appropriate resources yields better energy savings. Co-scheduling applications by using unassigned resources often saves more energy than serial execution counterparts. Although manually finding the best frequency offers the highest energy savings, our automated approach achieves up to 31.83% energy savings compared to normal execution of applications. The optkit tool is publicly available on GitHub

Osman Yasal
Koç University · Fen Bilimleri Enstitüsü
2024
00
Yüksek LisansAçık ErişimEN

Feature generation for SME behavioral credit scoring model using graph embeddings

Banks lend credit regarding customers' credit scores from in-house credit risk models. Corporate customers hold greater risk and are evaluated with already high-performing models. In this thesis, corporate customers' financial interactions (money transfers) with each other are used to develop an additional feature set to improve the behavioral credit risk scoring model of small and medium-sized enterprises. Graph representation learning does a good job in terms of extracting the essence of such relationships and projecting them into a multidimensional space. For this purpose, a graph of enterprises is constructed where they are nodes and their sum of money transfer amount over time are edges. An inductive graph representation learning algorithm, GraphSAGE, is employed. Therefore, the graph is created with directed and weighted edges and reduced to a strongly connected graph. Then, the algorithm is run in various settings to achieve the best performance. A credit scoring pipeline using different machine learning algorithms is added to improve the current model with these new embedding features. QNB Finansbank's real banking data for 2022 is used to perform this study. After looking at the results of the computational experiments, the most important contribution came from the concatenation of multiple embeddings generated with different aggregator functions. Starting from here, a potential economic savings is calculated. Additionally, it is seen that the new architecture improves the credit scoring of nodes with either low transfer amount or number of edges the most.

Bank creditsGraph neural networksMachine learning+2
Kerem Kaşıkcı
Koç University · Fen Bilimleri Enstitüsü
2024
00
Yüksek LisansAçık ErişimEN

Enhancing scene sketch understanding through a dual-network: Visio-temporal segmentation and context-aware sketch recognition

Understanding scene sketches involves segmenting and categorizing individual objects within the sketch. Semantic segmentation in scene sketches is crucial for distinguishing distinct sketches, but current methods often treat sketches as bitmap images, which can result in a loss of stroke order information. However, people tend to draw objects sequentially, so leveraging this temporal order could enhance segmentation performance. Moreover, traditional methods typically focus on class-level segmentation, failing to differentiate between instances within the same category. Another important aspect of scene sketch understanding is classifying individual sketch objects within a scene. These objects often lack the detail needed for standalone recognition, making their identification challenging without contextual information. For instance, a sketch that is fluffy and circular could be interpreted as a "bush" or a "cloud," depending on its position and size within the scene. Bushes are typically drawn on the ground, while clouds are usually sketched in the sky. Despite their similar appearances, their interpretation can vary based on their context. To enhance recognition accuracy, information about relative position and size is often used in computer vision tasks like object recognition and image classification. However, many current sketch recognition methods treat sketches in isolation, overlooking the contextual information present in the scene. To address these issues, I propose a dual-network approach comprising two novel networks for separate tasks: scene sketch segmentation and scene sketch recognition. The first network, the Class-Agnostic Visio-Temporal Network (CAVT), detects individual objects in a scene sketch using a class-agnostic object detector and groups strokes with its post-processing module. This network can distinguish object instances at the stroke level, independent of their categories. The second network, Context-Aware Graph Attention Transformer Network (CGAT-Net), processes individual sketch objects within the scene and leverages inter-object relationships to find their appropriate categories. This work is the first to apply a context-based sketch recognition approach by leveraging a novel Transformer-based Graph Attention Network within scene sketches. Additionally, the literature lacks free-hand scene sketch datasets with both instance and stroke-level class annotations. To fill this gap, I collected the largest Free-hand Instance- and Stroke-level Scene Sketch dataset (FrISS) that contains 1,000 scene sketches and covers 403 different object classes with dense annotations. Extensive experiments on FrISS and other scene sketch datasets demonstrate that the dual-network approach, combining CAVT and CGAT-Net, as well as each network individually, outperforms existing methods in their respective domains.

Aleyna Kütük
Koç University · Fen Bilimleri Enstitüsü
2024
00
Yüksek LisansAçık ErişimEN

Hierarchical spatial decompositions under local differential privacy

The popularity of smartphones, GPS-equipped devices, social networks, and connected vehicles continues to increase the volume of spatial data available for collection and analysis. Spatial decompositions assist in handling big spatial data, and they have been commonly used in the centralized differential privacy (DP) literature for range query answering, spatial indexing, count-of-counts histograms, data summarization, and visualization. However, their applications under the emerging local differential privacy (LDP) notion are relatively scarce. In this thesis, we study the problem of building hierarchical spatial decompositions under LDP, focusing on two methods: quadtrees and kd-trees. We develop two solutions for quadtrees: a baseline solution that is inspired by the centralized DP literature, and a proposed solution that utilizes a single data collection step from users, propagates density estimates to remaining nodes, and performs structural corrections to the quadtree. Since kd-trees rely on node medians which are data-dependent, we observe that it is not feasible to build kd-trees using a single data collection step. We therefore propose an iterative solution that constructs kd-trees in top-down fashion by utilizing a novel algorithm for estimating node medians at each tree depth. We experimentally evaluate our quadtree and kd-tree algorithms using four real-world spatial datasets, multiple utility metrics, varying privacy budgets, and tree parameters. Results demonstrate that our algorithms enable the building of accurate spatial decompositions that provide high utility in practice. Notably, our quadtrees and kd-trees achieve substantially lower errors in answering spatial density queries (up to 10-fold improvement) when compared with a state-of-the-art method.

Data privacy
Ece Alptekin
Koç University · Fen Bilimleri Enstitüsü
2024
00