DoktoraAçık Erişim

Deep learning based multi modal data summarization approaches using large language models

2025
0 görüntülenme
0 i̇ndirme
Danışman: Prof. Dr. Mehmet Karaköse

Özet (EN)

In data summarization problems, natural language processing, computer vision, artificial intelligence and data mining techniques can be used depending on the data format to be summarized. The need to use different types of data together in the same problem or system or to summarize a data with a different type of data mode has led to the development of multimodal data summarization applications. In this context, different multimodal and single-modal data summarization approaches have been developed in this study. Within the scope of the text summarization applications of the thesis study, a method focusing on abstractive and extractive summarization of news texts, a method focusing on abstractive summarization of dialogue texts and a method for abstractive summarization of texts containing patient questions have been developed. While an original encoder-decoder neural architecture is presented in the first text summarization approach developed, fine-tuning of pre-trained large language models has been performed in other text summarization approaches. The developed text summarization approaches are compared with existing methods in the literature and their contributions are clearly presented. Within the scope of video and audio summarization applications, the focus is on surveillance videos used in smart cities, query-driven multi-modal summarization of activity-based videos presented on online platforms, and multi-modal summarization of online meeting videos. In the approach focused on surveillance videos, video summarization process is performed with different video summaries obtained from two summary modules, object-centric and event-centric. While a genetic algorithm-based method that can preserve over 90% statistical features is developed for the object-centric summary module, a unique transformer architecture with over 90% classification performance is developed for the event-centric video summarization module. A multi-modal video summarization approach that uses a cross-attention mechanism focused on user query is developed for summarizing online activity videos. The video summarization performance of the presented approach is 87.1%. After obtaining the query-focused summaries of the videos, a new vision language model is also developed within the scope of the proposed method for the text descriptions of the summary videos. For the meeting video summarization, after automatic speech recognition, the extractive text, video and audio summaries of the meeting were obtained, and abstractive highlights of different sections of the meeting were produced. In the proposed method, pre-trained large language models were used for summarization operations and over 75% of the meeting videos were shortened. The performance of all developed video summarization approaches was compared with other studies in the literature and the performance superiority of the proposed methods was clearly demonstrated.

Yazar

Turan Göktuğ Altundoğan

Bu Yayına Nasıl Atıf Yapılır

Turan Göktuğ Altundoğan (Doctorate thesis). Deep learning based multi modal data summarization approaches using large language models, 2025, Fırat University.

Anahtar Kelimeler

Lisans

Tüm Hakları Saklıdır

Bu eser belirtilen lisans koşulları altında paylaşılmaktadır.

Fırat University tezlerinden daha fazlası