Master'sOpen Access

Authors native language identification in web mediums

Is this your thesis?

This record came from a bulk archive import. If it’s yours, link it to your profile.

2012
0 views
0 downloads

Abstract (EN)

In the domain of Text Mining and Document Classification, an introduction into the field of Authorship Attribution is presented. On the other hand, with the rapid growth of Internet technologies and applications, text is still the most common Internet medium. Examples of this include social networking applications such as Twitter, Facebook, etc. and web applications such as newsgroups, email, blog, etc. are also mostly text based. We developed a framework to determine an anonymous author?s native language for short length and multi-genre writing in English such as the ones found in many Internet applications.This thesis describes the development of such a tool using techniques from the fields of stylometry and traditional machine learning techniques. An author?s style can be reduced to a pattern by making measurements of various stylometric features from the text. In this framework, four types of stylistic text features (Lexical, Syntactic, Structural, and Content-Specific Features) are extracted and two machine learning algorithms (Decision Tree, Support Vector Machine and Naïve Bayesian) are designed for author?s native language identification based on the proposed features. For this research, we used four different collections of writings online news messages by speakers of four different nationalities: native English as well as speakers of Turkish, German, and Persian.

Author

Parham Mohammadalipour Tofighi

How to Cite

Parham Mohammadalipour Tofighi (Master Thesis). Authors native language identification in web mediums, 2012, Karadeniz Technical University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Karadeniz Technical University