Master'sOpen Access

Developing an application for converting the conversations in videos into text and time-based indexing

2024
0 views
0 downloads
Advisor: Prof. Dr. Mustafa Servet Kıran

Abstract (EN)

Advances in artificial neural networks have led to significant improvements in many data transformation and recognition processes such as text recognition and text-to-speech conversion. In particular, the Speech to Text method has gained popularity for transcribing audio conversations. Nowadays, as the popularity of video and audio content increases, people are reaching large audiences and generating revenue by producing special content on video platforms. In addition to analyzing the audio, the study also enables the querying of the data obtained so that users can access the section of the video related to the word they are looking for and watch the videos from this section. In this way, users can reach the part they want to view without having to watch an entire video to reach the seconds in which the word they are looking for occurs. Thanks to the developed methods, the texts obtained from the audio in the videos can give more accurate results compared to the STT method. With the segmentation method developed as a result of the identified deficiencies, frequency values were calculated and separated by Fourier transform method using the MFCC algorithm until the speaker stops in the data sets. With these processes applied to the STT method, data was recorded in a more accurate duration/text relationship than the STT method used by YouTube. With this method, second-sequence accuracy rate errors are prevented in data sets recorded without noise. All these datasets are saved to the database with duration/text matching so that users can see the seconds in which the words they are looking for occur and can play that part whenever they want. The working process of the application consists of two stages: the first stage is the downloading of the video and the second stage is the transcription of the audio from the video. Depending on the internet speed, the download time of the videos and the transcription time varies. Using a standard internet network, a one-minute video takes about 30 seconds to download and a five-minute video takes about 75 seconds to download. The second stage, the transcription of the audio in the videos, takes place after the video is downloaded. At this stage, it takes about 30 seconds to transcribe a one-minute video and 75 seconds to transcribe a 5-minute video. Although the developed program uses the same method as the YouTube automatic translation system, more successful results were obtained in the experiments. In the experimental process consisting of one hundred videos, it was determined that the developed program performed both proportionally more translations and numerically more accurate word translations according to the data obtained from the videos. Although both systems use the same methods, this difference is due to the partitioning method, MFCC and Fourier transform method used in the developed program. The main purpose of the application is to save time for the users. In regions where fiber internet speeds are provided, a fast solution can be offered to the user between three to five minutes, especially for videos over one hour. On a standard Wifi network, all steps can be completed in four to seven minutes for videos longer than one hour.

Author

Dr. Oğuzhan Mert Kiraz

How to Cite

Oğuzhan Mert Kiraz (Master Thesis). Developing an application for converting the conversations in videos into text and time-based indexing, 2024, Konya Technical University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Konya Technical University