DoctorateOpen Access

Derin sinir ağlarının yenı ve hibrit mimariler aracılığı ile transformlara genellemesi

2023
0 views
0 downloads
Advisor: Yrd. Doç. Dr. Mustafa Furkan Kıraç

Abstract (EN)

Object recognition is a foundational pillar for many computer vision tasks such as searching, tracking, navigating, scene understanding or information retrieval that require some kind of category knowledge at various levels. Even with the major advances in these tasks with the data-driven deep learning methods such as convolutional neural networks (CNNs), generalization to geometric variations and embedding part-whole relationships are still yet to be achieved when compared to human-level recognition. CNNs particularly fail to generalize to unseen viewpoints of a learned object even with substantial samples and are easily confused as the pooling operations lose the relation between existing entities in the input. Recently emerged capsule networks outperform CNNs in novel viewpoint generalization tasks even with significantly fewer parameters. Capsule networks group the neuron activations for representing higher-level attributes and their interactions for achieving equivariance to visual transformations. Capsules are designed to represent the pose of an existing visual entity and learned transformations are essentially pose transformations which are matrices. However, capsule networks have a high computational cost for learning the interactions of capsules in consecutive layers via the, so-called, routing algorithm in addition to the training stability problems. In this thesis, we propose to represent the pose information and transformations with quaternions in Quaternion Capsule Networks (QCNs). Quaternions are immune to the gimbal lock, have straightforward regularization of the rotation representation for capsules, and require a smaller number of parameters than matrices. QCNs directly inherit the existing EM-Routing for a fair comparison of the benefits of using quaternions instead of matrices. Experimental results show that QCNs generalize better to novel viewpoints with fewer parameters, and achieve on-par or better performances with the state-of-the-art Capsule architectures on well-known benchmarking datasets. Building on this proposal, we aimed to reduce the computational burden and embed feature vectors to the capsules in addition to pose information. In this context we propose, Alleviated Pose Attentive Capsule Agreement (ALPACA) which is tailored for capsules that contain pose, feature and existence probability information together to enhance novel viewpoint generalization of capsules on 2D images. For this purpose, we have created a Novel ViewPoint Dataset (NVPD) a viewpoint-controlled texture-free dataset that has 8 different setups where training and test samples are formed by different viewpoints. In addition to NVPD, we have conducted experiments on the iLab2M dataset where the dataset is split in terms of the object instances. Experimental results show that ALPACA outperforms its capsule network counterparts and state-of-the-art CNNs on iLab2M and NVPD datasets. Moreover, ALPACA is 10 times faster when compared to routing-based capsule networks. It also outperforms attention-based routing algorithms of the domain while keeping the inference and training times comparable.

Author

Dr. Barış Özcan

How to Cite

Barış Özcan (Doctorate thesis). Derin sinir ağlarının yenı ve hibrit mimariler aracılığı ile transformlara genellemesi, 2023, Özyegin University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Özyegin University