Master'sOpen Access

Graph based text summarization

2014
0 views
0 downloads
Advisor: Yrd. Doç. Dr. Nilgün Güler Bayazıt

Abstract (EN)

Nowadays amount of data has increased very much as the techology improves. So that number of documents has also gained incredible acceleration. The huge amount of documents make harder for users to reach information or cause users to miss some information during search. These kind of problems can be solved by using text summarization systems. Text summarization is the process of extracting a shorter version of the given text by maintaining the main idea. Generally two kind of methods are examined for text summarization systems which are called extractive or abstractive. Since abstractive summarization needs deep knowledge of natural language processing, most of the studies cover extractive methods. In extractive based summarization, sentences are selected as it appears in the given text. The key point is to choose sentences which involve important information for being in the summary. There are a lot of methods proposed for selecting sentences such that methods that use word frequencies, sentence clustering, graph based ranking methods, machine learning methods and so on are some of the techniques that are worked on. Graph methods are widely used for text summarization systems. Because graph representation enables different interpretation of data which helps to expose some other properties that cannot be easily observed by conventional methods. In this study, research is done on graph based text summarization. In the scope of research TextRank technique whose performance is proved has been utilized. This technique is inspired from "PageRank" technique which is used for ranking web pages. Method uses links between pages for ranking web pages with respect to importance. In order to use this technique in text summarization system, a relation has to be defined between sentences. In the scope of study effect of four different relation methods on TextRank is analyzed. Experimental studies is done by using DUC 2002 and CAST corpus. According to experimental studies the best result is found with content overlap method by using DUC but for the CAST set NGD gives the best result. In addition to this study, a novel hybrid system which is not tried out before has been developed. The hybrid system has been achieved by combining hierarchical agglomerative clustering and TextRank methods. In the propsed study sentences are clustered according to a criteria, and then TextRank is applied for choosing sentences from clusters. Experimental studies for the novel system is done by using DUC 2002 and CAST corpus like in the previous study so that previous TextRank results can be compared with results of the novel system. According to research; it is observed that hybrid system works with better performance when DUC set is used. CAST set shows that two methods surpass the four different relation methods that are used in TextRank and for the other two method results are close to each other. Therefore proposed novel hybrid method is able to show its advance for different kind of texts.

Author

Dr. Can Yalkın

How to Cite

Can Yalkın (Master Thesis). Graph based text summarization, 2014, Yıldız Technical University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Yıldız Technical University