Master'sOpen Access

Classification of document images using visual, textual, and layout features with deep learning methods

2024
0 views
0 downloads
Advisor: Prof. Dr. Rifat Edizkan

Abstract (EN)

The rapid increase in digitalization has made document management and classification processes more significant. In this context, substantial advancements have been made in the field of document classification. This thesis compares the performance of different models under three main approaches: convolutional methods based on visual features, Transformer-based methods relying on visual features, and methods incorporating visual, textual, and layout information. A comprehensive comparison was conducted using convolutional neural networks, Transformer-based ViT and ViC models, and the LayoutLMv3 architecture. These architectures were tested on the Tobacco-3482 and RVL-CDIP Small-200 datasets, and the impact of each method on document classification processes, document types, and success rates was evaluated. The thesis offers a unique analysis in the field of document classification by integrating multi-faceted classification approaches based on visual, textual, and layout information. Among deep learning methods, the highest accuracy in document classification was achieved using the LayoutLMv3 model, with a success rate of 95.78% on the Tobacco-3482 dataset. Keywords: Document Classification, Deep Learning, Convolutional Neural Networks, ViT, ViC, NLP, LayoutLMv3, Transformer

Author

Melike Burcu Ayhan

How to Cite

Melike Burcu Ayhan (Master Thesis). Classification of document images using visual, textual, and layout features with deep learning methods, 2024, Eskişehir Osmangazi University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Eskişehir Osmangazi University