Master'sOpen Access

Performance analysis of contextual vectors in visual question answering problem

Is this your thesis?

This record came from a bulk archive import. If it’s yours, link it to your profile.

2022
0 views
0 downloads

Abstract (EN)

Visual question answering (VQA) studies aim to provide consistency as well as to make sense of visual images. The VQA problem deals with the connection between a visual image and the question asked to that image. The interpretation and analysis of the discussed link ensures that the expected answer to the question asked is obtained from within the picture. In order to perform the analysis process, it is necessary to represent the visual images on the mathematical plane. These representations are called vectors. In the acquisition phase of visual vectors, Xception and Inception-Resnet-V2 models which are trained with ImageNet data were used. The models obtain vector representation from visual data with high accuracy due to deep convolutional networks and residual layer structure. Visual vector representation is not sufficient for the VQA problem. The mathematical representation of the question asked to the image is required. Representation of textual data, also known as word embeddings, can be obtained independently of the semantic context with the pre-trained models Word2Vec, Global Vectors for Word Representation (GloVe) and FastText, Bi-directional Encoder Representations from Transformers (BERT)learns and represents the sub-context between words with the multi-headed attention structure it is built on. BERT contextual vectors were adapted to strengthen the semantic integrity of the question asked in this study. When the results of the study were evaluated, it was seen that the BERT method achieved higher accuracy rates than the Word2Vec, GloVe and FastText methods. Thus, the success of the BERT contextual vectors method, which has just entered the literature, in the GSC problem has been demonstrated.

Author

Özlem Hakdağlı

How to Cite

Özlem Hakdağlı (Master Thesis). Performance analysis of contextual vectors in visual question answering problem, 2022, Bursa Uludağ Üni̇versi̇ty.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Bursa Uludağ Üni̇versi̇ty