Outliers - reasons for screening data, Advanced Statistics

Assignment Help:

Outliers - Reasons for Screening Data

Outliers are due to data entry errors, subject is not a member of the population that the sample is trying to represent, or the subject is really different. Statistical tests are quite sensitive to outliers so this problem should be addressed.

Univariate outliers are easy to detect (z-scores, box plots, histograms, etc.) standard scores larger than +/-3 are outliers (consider 4 is n>100 or 2.5 if n<10)

Multivariate outliers are difficult to detect. Mahalanobis distance is one powerful technique to use in this case (discussed later). This is evaluated as a chi-square statistic with degrees of freedom equal to number of variables in the analysis. A chi-sqaure statistic value that is significant beyond p<0.001 level determines outliers.

In most cases, it is ok to drop the value from the sample. One can also take steps to reduce the relative influence of outliers if the researcher decides to include the values in the analysis.


Related Discussions:- Outliers - reasons for screening data

Collective risk models, Collective risk models : The models applied to insu...

Collective risk models : The models applied to insurance portfolios which do not create direct reference to the risk characteristics of individual members of the portfolio when des

Geographical analysis machine, Geographical analysis machine is the proced...

Geographical analysis machine is the procedure designed to detect the clusters of rare diseases in a particular area. Circles of fixed radii are created at each point of the squar

Sequencing of 4 machines, how to resolve sequencing problem if jobs 6 given...

how to resolve sequencing problem if jobs 6 given and 4 machines given. how to apply johnson rule for making to machines under this conditions. please give solution as soon as poss

Z-tests, Hello! I am currently in graduate school earning a masters in ment...

Hello! I am currently in graduate school earning a masters in mental health counseling. I am in a stats course at current and we are reviewing z-scores. I am a little lost because

Copulas, Invariant transformations to combine marginal probability function...

Invariant transformations to combine marginal probability functions to form multivariate distributions motivated by the need to enlarge the class of multivariate distributions beyo

Factorial moment generating function, The function of a variable t which, w...

The function of a variable t which, when extended formally as a power series in t, yields factorial moments as the coefficients of the respective powers. If the P(t) is probability

Explain literature controls, Literature controls : The patients with the di...

Literature controls : The patients with the disease of interest who have received, in the past, one of two treatments under the investigation, and for whom the results have been pu

Pascal''s triangle, Pascal's triangle  is an arrangement of numbers describ...

Pascal's triangle  is an arrangement of numbers described by Pascal in his Traité du Triangle Arithmétique published in the year 1665 as 'The number in each cell is equal to in the

Fan-spread model, This term sometimes is applied to the model for explainin...

This term sometimes is applied to the model for explaining the differences found between naturally happening groups which are greater than those observed on some previous occasion;

Conjoint analysis, Conjoint analysis : The method used basically in market ...

Conjoint analysis : The method used basically in market research which is similar in many respects to the various dimensional scaling. The method attempts to assign values to the l

Write Your Message!

Captcha
Free Assignment Quote

Assured A++ Grade

Get guaranteed satisfaction & time on delivery in every assignment order you paid with us! We ensure premium quality solution document along with free turntin report!

All rights reserved! Copyrights ©2019-2020 ExpertsMind IT Educational Pvt Ltd