DoctorateOpen Access

Extractive text summarization based on independent sets in text graphs

2020
0 views
0 downloads
Advisor: Prof. Dr. Ali Karcı

Abstract (EN)

Within the scope of this thesis, two new graph-based approaches have been contributed to the general, unsupervised and extractive text summarization problem. The KUSH (Karcı, Uçkan, Seyyarer, Hark) tool used in the data processing stage of both approaches has been proposed and tried. The first proposed method is the CatSumm (Cengiz, Ali, Taner Summarization) model, which consists of three main steps. In the first step, normalization was performed with the KUSH tool. In the second step of the model, the graphs were clustered with spectral graph partitioning, so that the summaries were produced in accordance with the number of sentence ratios in the subgraphs. In the last stage, using the node weighting methods, sentences with high centrality values are included in the summary. In the second study, based on the prediction that the sentences corresponding to the nodes in the independent clusters should not be included in the summary, a limitation was made on the documents to be summarized before the effect of the nodes on the general graph was determined numerically. Both approaches were tested on the DUC (Document Understanding Conference, DUC-2002 and DUC-2004) data set and using ROUGE (Recall-Oriented Understudy for Gisting Evaluation) evaluation metrics. Experimental processes were repeated for summaries of 100, 200 and 400 words. The values reported with the proposed models reveal the contributions of innovative methods.

Author

Dr. Taner Uçkan

How to Cite

Taner Uçkan (Doctorate thesis). Extractive text summarization based on independent sets in text graphs, 2020, İnönü University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from İnönü University