Outliers - reasons for screening data, Advanced Statistics

Assignment Help:

Outliers - Reasons for Screening Data

Outliers are due to data entry errors, subject is not a member of the population that the sample is trying to represent, or the subject is really different. Statistical tests are quite sensitive to outliers so this problem should be addressed.

Univariate outliers are easy to detect (z-scores, box plots, histograms, etc.) standard scores larger than +/-3 are outliers (consider 4 is n>100 or 2.5 if n<10)

Multivariate outliers are difficult to detect. Mahalanobis distance is one powerful technique to use in this case (discussed later). This is evaluated as a chi-square statistic with degrees of freedom equal to number of variables in the analysis. A chi-sqaure statistic value that is significant beyond p<0.001 level determines outliers.

In most cases, it is ok to drop the value from the sample. One can also take steps to reduce the relative influence of outliers if the researcher decides to include the values in the analysis.


Related Discussions:- Outliers - reasons for screening data

Bartlett decomposition, Bartlett decomposition : The expression for the ra...

Bartlett decomposition : The expression for the random matrix A which has a Wishart distribution as the product of the triangular matrix and the transpose of it. Letting each of x

LASPEYERES QUANTITY INDEX, HOW TO OBTAIN THE LASPEYRES QUANTITY INDEX AND T...

HOW TO OBTAIN THE LASPEYRES QUANTITY INDEX AND THE FORMULA

Define probability judgements, Probability judgements : Human beings often ...

Probability judgements : Human beings often require assessing the probability which some event will occur and accuracy of these probability judgements often determines success of o

Coefficient of concordance, Coefficient of concordance : The coef?cient is ...

Coefficient of concordance : The coef?cient is taken in use to assess the agreement among m raters ranking n individuals according to some of the speci?c characteristic. Which can

Fan-spread model, This term sometimes is applied to the model for explainin...

This term sometimes is applied to the model for explaining the differences found between naturally happening groups which are greater than those observed on some previous occasion;

Queuing theory, 1) Let N1(t) and N2(t) be independent Poisson processes wit...

1) Let N1(t) and N2(t) be independent Poisson processes with rates, ?1 and ?2, respectively. Let N (t) = N1(t) + N2(t). a) What is the distribution of the time till the next epoch

Quota sample, Quota sample is the sample in which the units are not select...

Quota sample is the sample in which the units are not selected at the random, but in terms of a particular number of units in each of a number of categories; for instance, 10 men

Graphics., how to calculate the semi average method when 8 observations are...

how to calculate the semi average method when 8 observations are given?

Log-linear models, Log-linear models is the models for count data in which...

Log-linear models is the models for count data in which the logarithm of expected value of a count variable is modelled as the linear function of parameters; the latter represent

Define kalman filter, Kalman filter : A recursive procedure which gives an ...

Kalman filter : A recursive procedure which gives an estimate of the signal when only the 'noisy signal' can be observed. The estimate is efficiently constructed by putting the exp

Write Your Message!

Captcha
Free Assignment Quote

Assured A++ Grade

Get guaranteed satisfaction & time on delivery in every assignment order you paid with us! We ensure premium quality solution document along with free turntin report!

All rights reserved! Copyrights ©2019-2020 ExpertsMind IT Educational Pvt Ltd