Data reduction, Applied Statistics

Assignment Help:

The PCA is amongst the oldest of the multivariate statistical methods of data reduction. It is a technique for simplifying a dataset, by reducing multidimensional datasets to lower dimensions for analysis. It produces a small number of derived variables that are uncorrelated and that account for most of the variation in the original data set.'By reducing the number of variables'in this way, we can understand the underlying structure of the data. 'The derived variables are combinations of the original variables. For example, it might be that students take I0 examinations and some students do well in one examination while other students do better in another. It is difficult to compare one student with another when we have 10 marks to consider. One obvious way of comparing students is to calculate the mean score.

This is a constructed combination of the existing variables. However, one might get a more useful comparison of overall performances by considering other constructed cwbinations of the 10 exam marks. The PCA is one way of constructing such combinations, doing so in such a way as to account fer the maximum possible variation in the original data. We can then compare students' performance by considering this much smaller number of variables.

PCA states and then solves a well-defined statistical problem, and except for special cases always gives a unique solution wi.th some very nice mathematical properties. We can even describe some very artificial practical problems for which PCA provides the exact solution. The difficulty comes in trying to relate PCA to real-life scientific problems; the match is simply not very good. Actually PCA often provides a good approximation to common factor analysis, but that feature is now unimportant since both methods are now easy enough.


Related Discussions:- Data reduction

Null and alternative hypothesis, 1) Suppose you want to test a hypothesis t...

1) Suppose you want to test a hypothesis that two treatments, A and B, are equivalent against the alternative that the response for A tend to be larger than those of B. You plan to

Range, Range Official Exports Target 2000-2001 ...

Range Official Exports Target 2000-2001 Product ($ million) Plantation 500 Agriculture and Alli

Stratified random sampling, Stratified Random Sampling: This method of ...

Stratified Random Sampling: This method of sampling is used when the population is comprised of natural subdivision of units, The method consist in classifying the population u

Utility index , If the economy does well, the investor's wealth is 2 and if...

If the economy does well, the investor's wealth is 2 and if the economy does poorly the investor's wealth is 1. Both outcomes are equally likely. The investor is offered to invest

HLT 362, What is an interaction? Describe an example and identify the varia...

What is an interaction? Describe an example and identify the variables within your population (work, social, academic, etc.) for which you might expect interactions?

Determine the regression equation, The file Midterm  Data.xls has a tab lab...

The file Midterm  Data.xls has a tab labeled "National Grid vs. Alcoa" which presents historical price data for two stocks.  Using the National Grid price as the X-value and the Al

Regression analysis and experimental design, For many decades, there has be...

For many decades, there has been considerable attention paid to identifying various factors that help to reduce the number of fatalities on Australian roads. In 1964 Victoria and S

Write Your Message!

Captcha
Free Assignment Quote

Assured A++ Grade

Get guaranteed satisfaction & time on delivery in every assignment order you paid with us! We ensure premium quality solution document along with free turntin report!

All rights reserved! Copyrights ©2019-2020 ExpertsMind IT Educational Pvt Ltd