Master'sOpen Access

Büyük veri setlerinde destek vektör regresyonu için sütun ve satır seçme yöntemi

2010
0 views
0 downloads
Advisor: Yrd. Doç. Dr. Özden Gür Ali

Abstract (EN)

This study introduces an algorithm, which selects important observations and variables to estimate SVR models for very large data sets. In this two-stage methodology, namely the Row and Column Selection Algorithm, ?-SVR models with L1-norm regularization are used both for selecting rows and columns. The first stage penalizes support vector weights to identify few support vectors as important points to include in the training data set. These support vectors are then used in the second stage to select the variable subset to be kept in the training data by penalizing the variable weights. The accuracy of holdout test set of the RBF-SVR models trained on this set including selected rows with all variables is significantly better than the accuracy of the same model trained on the benchmark which is the randomly sampled data set of the same size with all variables and SVMTorch.The contribution of this thesis is the development of an algorithm which facilitates estimating SVR models with very large data sets which are accurate and low complexity. By using the proposed algorithm, it is possible to select the important observations and variables and use them for estimation. The experimental results validate that the resulting training data set works effectively and reduces the number of variables dramatically while improving the generalization error of the RBF-SVR models in the presence of redundant variables. Furthermore, we investigate how the selected points differ from others by analyzing their distribution with respect to their distance from the prediction line, target values and the input variables of data set. This analysis demonstrates that L1-norm ?-SVR provides much more sparse solution than standard ?-SVR. Further the observations with extreme target values are more likely to be selected than average observations. Interestingly, in contrast to standard?-SVR, the L1-norm ?-SVR support vectors can be located both inside and outside the ?-tube. Moreover, low multi-collinearity between selected columns gives face validity variable selection procedure of our algorithm, namely second part of the proposed algorithm. Lastly, we identify which points are selected with respect to variables' values. The result of this analysis indicates that the row and column selection algorithm select observations based on background knowledge.

Author

Dr. Kübra Yaman

How to Cite

Kübra Yaman (Master Thesis). Büyük veri setlerinde destek vektör regresyonu için sütun ve satır seçme yöntemi, 2010, Koç University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Koç University