DoktoraAçık Erişim

Mining Turkish documents by meaning based techniques

2007
0 görüntülenme
0 i̇ndirme
Danışman: Prof. Dr. Oya Kalıpsız

Özet (EN)

One of the first application areas of computer systems had been data collection and reporting. As data storage capacity and the computing power increased, it became possible to store and query larger amounts of data (Fayyad, 1996a). Based on the developments, it is now possible to find out relations and patterns in the data that were not easy to discover before. These techniques are known to be data mining techniques and are different than conventional techniques.Data mining is the analysis of (potentially large) data sets aimed at finding unsuspected relationships, patterns, rules, uncertainties and statistically important structures which are of interest or value to the database owners (Hand vd., 2001). Data mining techniques are not directly suitable for analyzing document typed data like querying documents, searching for the relationships between documents. For specific document analyzing needs, techniques different than data mining techniques have been developed and this new discipline is known as text mining, document mining, semi-structured data mining. The objective of document mining is to discover the content of documents by computers as if it is read by human beings. When this is the case, the language of the document becomes important. From this point of view there had been many studies under natural language processing field for decades. For Turkish language, natural language studies are quite new and very few of them has resulted with useful outcomes. Moreover not all them are shared among researchers. Based on this fact, a document mining study for mining Turkish documents using NLP techniques was out of question, especially by 2004 when this study has started. The study we are to present in this theses aims to mine Turkish documents by using Latent Semantics Indexing (LSI) technique and to develop LSI, we are proposing to enhance LSI by combining it with n-gram approach. In order to compare and evaluate different document mining techniques on the international scale, a standart document set had been developed for a specific group of techniques like clustering, questing answering etc. This point is also a missing point for Turkish language. We had to develop our own document set for the study and collected articles from the business magazines that were published in Turkey from year 2000 to 2006. This document set was used to test our technique by doing document querying and clustering. For both document querying and clustering, the test results has shown that, n-gram based LSI technique outperformed conventional LSI. To overcome the assertions that the document set we had collected is not a standard one as a result of which the test results may not show the reality, we tested our technique on the internationally accepted English document set known as Reuters21578. Parallel to Turkish tests, English document test also showed the same results both for document querying and clustering. Keywords : Data Mining, Tezt Categorization, Text Retrieval, Information Retrieval, Querying

Yazar

Dr. Ahmet Güven

Bu Yayına Nasıl Atıf Yapılır

Ahmet Güven (Doctorate thesis). Mining Turkish documents by meaning based techniques, 2007, Yıldız Technical University.

Anahtar Kelimeler

Lisans

Tüm Hakları Saklıdır

Bu eser belirtilen lisans koşulları altında paylaşılmaktadır.

Yıldız Technical University tezlerinden daha fazlası