Rag supported knowledge-based question-answer system with multimodal (text + visual) data
2025
0 views
0 downloads
Advisor: Dr. Öğr. Üyesi Volkan Altıntaş
Abstract (EN)
To address the inadequacy of traditional systems and Large Language Models (LLMs) in handling multimodal data (text, tables, figures) and complex PDF documents requiring privacy, this thesis presents a comprehensive and energy-efficient offline RAG (Retrieval-Augmented Generation) framework that overcomes existing online dependencies. The technical distinguishing feature of the work lies in its offline dual-channel multimodal architecture, which processes visual (Image) data through separate structures. This approach provides a closed and reliable solution by hosting Large Language Models (LLMs) locally on the device (Edge AI) without requiring an external API; this developed Hybrid System significantly outperforms the accuracy of basic VLM models (LLaVA and VILA's approximately 0.72-0.73 accuracy) in real-world tests, achieving an average success rate (CheckAvail score) of 0.8827. This proves it is a competitive alternative to the online, high-performance XSum architecture (0.97 score). The XSum score was found to be consistent with text-focused metrics. The Hybrid Score Model, which shows similar success to other hybrid models, particularly in evidence capture success (true0), unfortunately has the lowest value (0.69) in the metrics for excluding irrelevant evidence (true1); this indicates that the model lacks specificity and is prone to false positives (FP). Detailed analyses have revealed that this low performance in the true1 metric is primarily due to the evaluation of evidence from the visual (Image) channel, while the text channel produced satisfactory results. Additional analyses showed that these false positives were eliminated when verified using an offline visual LLM (Vision LLM). Work is ongoing to find a better solution in the future. Therefore, future work will focus on developing new-generation filtering mechanisms that will increase the model's specificity to overcome this problem, especially in the visual channel, architectural optimizations for resource-constrained LLMs, and robust image/visual extraction functions that can adapt to different PDF structures.
Author
Dr. Onur Yazıcı
Institution
How to Cite
Onur Yazıcı (Master Thesis). Rag supported knowledge-based question-answer system with multimodal (text + visual) data, 2025, Manisa Celal Bayar University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Manisa Celal Bayar University
- TÜRK FİKİR HAYATINDA MİLLİYETÇİ MUHAFAZAKAR KADIN FİGÜRÜ: AYŞE DERGİSİ, EMİNE IŞINSU VE SAMİHA AYVERDİ(2025)
- Dissolution kinetics of celestite ore with acid and base solutions and the production of SrCrO4 in the Sivas region(2018)
- The Mawlid of Behiştî (Examination -text)(2019)
- The influence of activity based teaching on historical thinking skills, based upon active learning and academic achievement in history course subjects of fourth grade social studies(2019)
- Identity perceptions of Izmir Jews(2019)
- Systematic examination of mites of Tydeidae family in Foça district (İzmir)(2019)
