Master'sOpen Access

The effect of parallel corpus quality vs size in English to-Turkish statistical machine translation

Is this your thesis?

This record came from a bulk archive import. If it’s yours, link it to your profile.

Abstract (EN)

Machine Translation (MT) is the process of translating an expression to another language automatically with the aid of computers. MT has been studied since the early 1950s. MT, which is thought to increase in importance after World War 2, has been invested due to political, social and economic facts. Although, many important studies have been conducted in the following years, the results couldn't meet expectations. The investments and studies in this field began to decline from the middle of 1960. The Automatic Language Processing Advisory Committee (ALPAC) which studies about costs, projections, expectations and requirements about MT, has issued a negative report about MT and caused loss of motivation and investment in MT field. During this first period of MT studies, MT was primarily performed using rule based transfers of some representation levels like morphological, syntactical or semantic representations. The statistical approaches which are developep under the fluence of internet and big data technologies have started to be utilized in signal processing and natural language processing. The hesitancy in MT has eliminated by Statistical Machine Translation (SMT) studies pioneered by IBM and many researchers has started to work in developing this new field. Another MT approach that based on training data is example based machine translation (EBMT). Nowadays, MT systems have reached a certain success and its applications in various fields have steadily increased because of the convenience of data acquisition. But, the research and development activities on the systems that are able to combine all of the features expected, is proceeding rapidly. The featetures that expected from a successful MT system are as follws: ability to process understandable and literal translations, ability to process automatic translations without any human intervention and ability to process general-purpose texts without any domain restriction. The most important training data for example based MT models and statistical MT models are parallel corpus. Parallel Corpus are consist of texts that translation of each other and aligned at sentence level. In addition to MT, parallel corpus are widely utilized in word disambiguation, information retrieval and some of other natural language processing fields. In this study, general information about history of MT and methods are presented, the point reached by SMT is investigated. Furthermore, publicly avaible parallel corpus between Turkish and English languages are studied and severalTurkish - English parallel corpus are constructed from various sources. The aim of this study is to figure out the effects of parallel corpus size and quality in statistical machine translation between Turkish and English languages. In this study, a machine learning based classifier is developed to classify parallel sentence pairs in a parallel corpus as high -quality or poor quality. This calassifier has been applied to a parallel corpus contains 1 million parallel English – Turkish sentence pairs and 600K high-quality parallel sentence pairs were obtained. The multiple SMT systems with various sizes of entire raw parallel corpus and filtered high quality corpus, their performances are evaluated in our experiments. As expected, the experiments show that the size of parallel corpus is a major factor in translation performance. However, instead of extended corpus with all available "so -called" parallel data, a better translation performance and reduced time-complexity can be achieved with a smaller high-quality corpus using a quality filter. Keywords: Machine Learning, Artificial Intelligence, Natural Language Processing, Machine Translation, Statistical Machine Translation, Parallel Corpus, Parallel Corpus Filtering, Data Selection

Author

Eray Yıldız

How to Cite

Eray Yıldız (Master Thesis). The effect of parallel corpus quality vs size in English to-Turkish statistical machine translation, 2014, Yıldız Technical University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Yıldız Technical University