Master'sOpen Access

Rag supported knowledge-based question-answer system with multimodal (text + visual) data

2025
0 views
0 downloads
Advisor: Dr. Öğr. Üyesi Volkan Altıntaş

Abstract (EN)

To address the inadequacy of traditional systems and Large Language Models (LLMs) in handling multimodal data (text, tables, figures) and complex PDF documents requiring privacy, this thesis presents a comprehensive and energy-efficient offline RAG (Retrieval-Augmented Generation) framework that overcomes existing online dependencies. The technical distinguishing feature of the work lies in its offline dual-channel multimodal architecture, which processes visual (Image) data through separate structures. This approach provides a closed and reliable solution by hosting Large Language Models (LLMs) locally on the device (Edge AI) without requiring an external API; this developed Hybrid System significantly outperforms the accuracy of basic VLM models (LLaVA and VILA's approximately 0.72-0.73 accuracy) in real-world tests, achieving an average success rate (CheckAvail score) of 0.8827. This proves it is a competitive alternative to the online, high-performance XSum architecture (0.97 score). The XSum score was found to be consistent with text-focused metrics. The Hybrid Score Model, which shows similar success to other hybrid models, particularly in evidence capture success (true0), unfortunately has the lowest value (0.69) in the metrics for excluding irrelevant evidence (true1); this indicates that the model lacks specificity and is prone to false positives (FP). Detailed analyses have revealed that this low performance in the true1 metric is primarily due to the evaluation of evidence from the visual (Image) channel, while the text channel produced satisfactory results. Additional analyses showed that these false positives were eliminated when verified using an offline visual LLM (Vision LLM). Work is ongoing to find a better solution in the future. Therefore, future work will focus on developing new-generation filtering mechanisms that will increase the model's specificity to overcome this problem, especially in the visual channel, architectural optimizations for resource-constrained LLMs, and robust image/visual extraction functions that can adapt to different PDF structures.

Author

Dr. Onur Yazıcı

How to Cite

Onur Yazıcı (Master Thesis). Rag supported knowledge-based question-answer system with multimodal (text + visual) data, 2025, Manisa Celal Bayar University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Manisa Celal Bayar University