Theses supervised by Prof. Dr. Serdar Kurt
10 theses · Dokuz Eylül University
Model selection methods for multivariate linear partial least squares regression
Having large numbers of predictor variables or having more predictor variables than the number of observations is a serious problem in regression analysis. When a data set contains many predictor variables, multicollinearity can become an issue. Multicollinearity arises when predictor variables measure the same concept or when there is a linear relationship among them. These problems can cause high degrees of correlation and violate the assumption of Ordinary Least Square Analysis. As a result, it causes poor estimates of parameter estimation in regression analysis. A possible solution to this problem is a statistical method called `Partial Least Squares Regression?. PLSR allows for the study of regression in many situations that Multiple Linear Regression does not.In this thesis, PLSR has been studied in the analysis of obtaining the number of new predictor variables called `latent variables?. After obtaining the latent variables, this thesis is concerned with analyzing how many of these latent variables are the most relevant for describing the variability of predictor and response variables. Some model selection methods, such as two of the Multivariate Akaike Information Criterion which are studied by Bozdogan and Bedrick respectively, use PRESS values obtained from k-fold cross validation and Wold?s R criterion to obtain the optimum number of latent variables. The simulation study presented in this thesis has been performed to compare the performance of these criteria. The simulation results of MAIC, PRESS and Wold?s R were obtained from different number of observations and different numbers of predictor variables. These results show that for small-sized design matrices, all criteria achieved the true number of latent variables. However, the results for the other-sized design matrices varied greatly and they consistently showed different numbers of latent variables. The whole analysis, including all simulations and calculations, were done using MATLAB statistical program.
Genetic algorithm based outlier detection using information criterion
Outlier, abnormal or unusual observation can be defined as an observation that lies outside the overall pattern of a distribution. Diagnostic methods for identifying a single outlier or influential observation in a linear regression model are relatively simple from both analytical and computational points of view. However, if the data set contains more than one outlier, which is likely to be the case in most data sets, the problem of identifying such observations becomes more difficult because of the masking and swamping effects.In this thesis, Genetic Algorithm (GA) based outlier detection using information criteria in multiple regression models has been studied. A GA was allowed simultaneous detection of outliers in data sets. Thus, this method is to overcome the problems of masking and swamping effects. It is derived additional penalized value of information criteria for Akaike Information Criterion (AIC) and Information Complexity Criterion (ICOMP) and named as AIC' and ICOMP' respectively in this study. They have been used as the fitness function of genetic algorithms to detect outliers in multiple regression. The simulation study has been performed to compare consistency and robustness properties of AIC' and ICOMP' against corrected Bayesian Information Criterion (BIC'). Simulation results of AIC', BIC' and ICOMP' obtained from different number of sample sizes, different penalized Kappa values of information criterion and different number of explanatory variables for different percentage of outlier in dependent variables. The numerical example and simulation results clearly show a much improved performance of the proposed approach in comparison to existing method especially followed by applying the ICOMP' approach in order to accurately (robustly) detect the outliers.
A statistical information system for poison control centers
A poison control center (PCC) is a modern health service unit that is able to provide immediate, free, and expert treatment advice and assistance over the telephone in case of exposure to poisonous or hazardous substances. The aims of PCC are to provide guidance for treatment strategies by giving right, current and comprehensive information rapidly in case of poisoning and to promote the safe, effective and proper use of medicines. Another major task of PCC is to disseminate and develop knowledge in these areas through teaching and research.In this study, after giving general information about information systems and poison control centers, the statistical information system being developed for poison control centers (SISPCC) has been presented. Development stages and structure of the developed system have been explained. The modules of the information system have been presented. Consequent to the entry into the developed information system of the collected data by Dokuz Eylül University Drug and Poison Information Center (DPIC) since 1993, results of the 2007 annual report have been given. This report analyzed the etiological, demographical and clinical characteristics of exposures reported to the DPIC in 2007. And finally, conclusion and some suggestions for further research were given.
Problem of omitted variable in regression model specification
In many non-experimental studies, the analyst may not have access to all relevant variables, and does not include these variables into the model and omits them. To omit some variables that affect the dependent variable from the model may cause omitted variables bias. In this thesis, it is aimed to investigate the omitted variable bias, its importance, reasons, and consequences and to research the methods for dealing with omitted variable bias and RESET test which is a method for detecting omitted variable(s).In this study, a simulation was performed by using the programs written in Minitab which is a statistical software package. Three types of populations with 1000 observations which varied depending on the correlations between the variables were generated and random samples were drawn from these populations. Though the true model had three independent variables, the models were estimated by omitting one and then two independent variables for each sample. 10,000 repetitions were generated for each of sample sizes. Therefore when correlations were changed and the number of omitted variables was increased, the effects of omitted variable bias were investigated. The amount of bias, the estimated coefficients, coefficients of determination and the adjusted coefficients of determination, standard deviations of the estimated coefficients were computed for every model and F statistics were also computed for applying RESET test and they were all compared for each population. Moreover, by increasing the sample size, it was investigated whether the effects of omitted variable bias were changed depending on sample size.Keywords: Regression analysis, model specification error, omitted variable bias, RESET test
Testing non-additivity in statistical models
In this thesis, testing non-additivity (interaction) in two-way ANOVA tables, and contingency tables are studied. For two-way ANOVA tables, methods designed especially for testing interaction when there is only one observation (no replication) per cell are the focus, whereas log-linear models are considered for contingency tables. Cressie & Read (1984) developed the family of power-divergence measures. The family of minimum power-divergence estimators is obtained by minimizing these measures for unknown parameter. Cressie & Pardo (2000, 2002) also developed unified approach to model selection problem in nested log-linear models with test statistics based on power-divergence measures. However, the weight put on empty cells is the problem with power-divergence measures which affects the performances of minimum power-divergence estimators and power-divergence test statistics. Basu & Basu (1998) developed the family of penalized power-divergence measures as a solution for this problem. In this thesis, simulation study has been performed to compare efficiency and robustness properties of ordinary and penalized minimum power-divergence estimators for log-linear independence model in 2 x 2 contingency tables. The new families of penalized power-divergence test statistics have also been proposed to over come with the problem of weight that power-divergence test statistics put on empty cells. Ordinary power-divergence test statistics developed by Cressie & Pardo (2000, 2002) and proposed penalized power-divergence test statistics have been compared by simulation study in terms of exact size and power properties for testing nested log-linear models in 2 x 2 and 2 x 2 x 2 contingency tables.
ANOVA methods for the group means with unknown variances
Analysis of variance (ANOVA) is one of the most powerful tools while investigating the sources of variability in many disciplines like medicine, engineering, agriculture, education, psychology, sociology and biology. In ANOVA, variance of the distributions in which the samples are drawn should be homogeneous to validate the underlying probability distribution of the method and to confine the errors within the desired limits. Violation of this equality of variances assumption is called as heteroscedasticity in literature. In this study, after describing a general appearance of one-way ANOVA and effects of it?s inevitable assumptions, the results of heteroscedasticity in one-way fixed effects ANOVA have been examined with a close concern on large sample approximations of treatment and error mean sum of squares and distortion of the distribution of the F ratio. Then, two new and simple approximation procedures which intend to create an easy and applicable alternative under heteroscedasticity and nonnormality have been presented. The purpose of these new approximation procedures is to preserve the actual Type I error rate at a level determined by the researcher and to increase the power as well. Performance of these two new approximation procedures under different experimental patterns have been observed with two separate simulation studies and finally, some recommendations about the preference of these tests and further research topic were given.
Comparison of the methods for constructing point estimates for variance components
The purpose of this investigation is to estimate the variance components parameters according to the analysis of variance (ANOVA), maximum likelihood (ML) and restricted maximum likelihood (REML) procedures in the one-way random effects model for balanced data, and to compare these estimation methods. In this study, the simulation studies were made by using the programs written in statistical software Minitab. The variance components estimators for ANOVA, ML and REML were calculated 1000 times by simulation made for the different number of observations and of levels; and the results were comparied and the most appropriate method for this study were investigated. The means and the standard deviations of the estimates were considered as the criteria of this comparison. After the evaluation of the results, it was observed that the values of ANOVA and REML were more and more close to each other. Furthermore, it was seen that ANOVA sometimes can give the negative estimates and REML always gives the nonnegative estimates. Though, the means of estimates of these two methods are close the real value. If we keep in mind the negative estimate situation for the ANOVA method, REML can be found appropriate. But although ANOVA gives the negative estimate, it has good results. Although ML estimation method gives the nonnegative estimates, the results of the treatment variance estimate are away from the real value for the balanced data. Likewise, ANOVA and REML estimates of the treatment variances are more influenced than ML about a increases.
Estimation of the parameters in the balanced incomplete block design
ABSTRACT In this study, the methods which are used for estimating the treatment effects in the balanced incomplete block design and the estimated values of the treatment effects obtained by using these methods have been compared. When it is impossible to make the required number of treatments, which are needed for each block in the randomized complete block design, the experiment is designed in the balanced incomplete block design. In the balanced incomplete block design, it is suggested that other methods should be used rather than the least squares estimators to estimate the treatment effects. Therefore, in order to estimate the treatment effects in the balanced incomplete block design, the intrablock, the interblock and the combined estimates methods are introduced in literature. In this study, the simulation studies were made by using the programs written in statistical software Minitab. The intrablock, the interblock and the combined estimates for the balanced incomplete block design and the least squares estimates for the randomized complete block design of the treatment effects were calculated 2500 times by simulation. The results were compared and the most appropriate estimator of the treatment effects for the balanced incomplete block design was investigated. The means and the standard deviations of the estimates were considered as the criteria of this comparison. After the evaluation of the results, it was observed that each of the three methods gave unbiased results when the block effects were insignificant for the balanced incomplete block design. When the block effects were significant, it was seen that the results of the means of the intrablock and the combined estimates were unbiased. The interblock estimates results were observed as inappropriate. In both of the situations in which the block effects were significant and insignificant, the standardVI deviation of the intrablock estimates is lower than the standard deviations of the interblock and the combined estimates.
Principal components in the problem of multicollineartity
ABSTRACT In this study, principal components regression and ridge regression are examined among the methods used to remedy multicollinearity problem in multiple linear regression model. One of the assumptions in multiple linear regression is that there must be no perfect linear relations among the regressors. The relationship among the regressors is called multicollinearity. In case of multicollinearity, parameter estimations by least square method have large variances and hypothesis tests result in contradictory. There are various methods for dealing with multicollinearity problem. Biased regression methods (BRM) are the ones that can explain the structure of multicollinearity and provide small standard errors among the methods used. In this study two of biased regression methods; principal components regression and ridge regression are examined as theoretically and researched which methods give the best consequence by simulation. In the application, 50 repetitions have been generated for each of the sample sizes of 40, 80 and 120. Least squares, ridge and principal components regression are used for each sample. Regression coefficients for each estimator were computed and the mean and the standard deviation of the estimates were used as statistical comparison criteria. According to comparisons among the estimators the principal components regression has been found to provide better estimates.
Poisson regression modelinde otokorelasyon
ABSTRACT In this study, time series of Poisson count model is concerned. In real situations, mean-variance equality, which is the basic property of Poisson data, cannot be provided. Generally, in such data variance exceeds mean, this is called overdispersion. When the overdisperison is detected, then there may be autocorrelation in latent process for Poisson regression model. Correlation is assumed to result from a latent process which is added to the linear predictor in a Poisson regression model. A quasi-likelihood approach is used as a parameter estimation technique. Tests for the presence of the latent process and autocorrelation of the latent process are examined. Asymptotic properties of the regression coefficients are investigated by using a simulation study. As an illustration, monthly number of deathes who were infected by pulmonary tuberculosis for the years 1996 to 2002 in Izmir are investigated as a parameter- driven model and the asymptotic properties of the regression coefficients are investigated, then a suitable model is constructed for forecasting. Keywords: Quasi-Likelihood Method, Latent Process, Poisson Regression, Overdispersion, Autocorrelation.