data project, Applied Statistics

Assignment Help:
Dr. Jim Mirabella
UNIT EIGHT: DATA ANALYSIS PROJECT
All Excel output should be copied into a single Word document where you must enter all of your
responses to the questions below. Format the document professionally so it flows well. Include
a table of contents.
? Choose any published database from the internet or Bethel library (such as those from the
Census Bureau or any financial sites). You may opt to use one of the data files provided
by the instructor if applicable.
? Get advanced approval from the instructor on your chosen database.
? If the file is large, randomly choose 200 of the observations from the data.
? Explain each variable in the file that you are analyzing. Be sure your file includes at least
3 scale variables and at least 2 nominal variables.
? Conduct a descriptive analysis on any 2 interval / ratio variables you wish using
Descriptive_Statistics.xls and Frequency_Distribution.xls. Explain the output.
? Conduct 3 different hypothesis tests of your choice using appropriate variables from the
file (note: you must use 3 different tests and not run one test on 3 different variables). In
each case, state the variables being tested as well as the hypothesis, decision and
conclusion. Use 3 of the following (1-Sample Test for Means, 1-Sample Test for
Proportions, 2-Sample Test for Means – Independent Samples, 2-Sample Test for Means
– Paired Samples, 2-Sample Test for Proportions, Analysis of Variance, Chi Square
Goodness of Fit Test, Chi Square Test of Independence, Correlation Test).
? Develop a model to predict an interval / ratio variable using at least 2 other variables.
Use Multiple_Regression.xls and state the regression model and which variables are or
are not significant. Also, use the model to make a prediction by making up values for
each of the independent variables.
? Write a one to two page summary of your findings. Include the data file in the appendix.
The project is due at the close of Week 8. You may work alone or in a team of 2 (you choose
your own partners and both of you must let your instructor know of your intent to work
together).
You may use a data set from the internet or from your workplace, or you may use one of the files
provided on www.drjimmirabella.com/bethel The files are described here, and the variables are
described within the files. If there is anything confusing about these data files, please ask your
instructor.
BASEBALL: This file includes actual team by team data for the 1997 MLB season. The key
variable to predict in Multiple Regression Analysis is the number of wins (or possibly the
attendance). Lots of interesting analysis possibilities here, including how team salary relates to a
team making the playoffs, or whether money buys wins, or how wins relate to attendance, or
how performance on the field relates to the field surface, etc. If you know something about
baseball, this file should make sense to you.
Dr. Jim Mirabella
CARS: This file is self-explanatory after you open it. Several variables describe the car (sports
car, SUV, engine size, horsepower, etc.), and several describe the car’s performance (CityMPG
and Highway MPG). It also includes the Dealer Cost and Suggested Retail Price. The key
variable to predict in Multiple Regression Analysis is the Suggested Retail Price. Lots of
crosstabulation options for Chi Square Analysis, lots of ANOVA and t-test options in which you
analyze miles per gallon or price as a function of any of the many variables included.
LOW BIRTH WEIGHT: This file looks at factors that might predict a baby being born with low
birth weight. Birth weights of 5.5 pounds or less are considered low in this file. Use the actual
birth weight as the key predicted variable in Multiple Regression Analysis. Lots of variables
about the mother regarding her weight, race, medical problems, and doctor’s visits can be used
for Chi Square analysis or as factors in ANOVA’s or t-tests.
MUTUAL FUNDS: This file looks at Large Cap, Mid Cap and Small Cap funds with either
Growth or Value objectives. Some funds have fees. Funds are either high, average or low risk.
Assets range from 50.7 million dollars to 66.5 billion dollars. For Multiple Regression Analysis,
you can choose to predict any of the three Return rates (measured in percents). Lots of
categorical variables to choose from in a Chi Square Analysis or as factors to analyze differences
in mean return rates.
TIPS: This file includes data on 75 patrons at the Spaghetti Warehouse on a given day. The key
variable here is the Tip Rate or the Tip Total. If you wait tables there, under what circumstances
are you most likely to get a better tip? You can compute mean Bills or Tips or Tip Rates as a
function of the meal time, the party size or the size of the party at the table.
Note that you should not use a nominal variable with 3 or more values in the Multiple
Regression Analysis (unless you convert to dummy variables, but that is unnecessary here).

Related Discussions:- data project

Lorenz curve , Lorenz Curve   It is a graphic method of measur...

Lorenz Curve   It is a graphic method of measuring dispersion. This curve was devised by Dr. Max o Lorenz a famous statistician.  He used this technique for wealth it i

Purposive or judgement sampling, Purposive or Judgement Sampling Under ...

Purposive or Judgement Sampling Under this method of sampling, the choice  of selection of sample  items from the universe  depends exclusively on the judgement  of the investi

Measures of dispersion, calculate variance and standard deviation of the f...

calculate variance and standard deviation of the following sample 12,22,32,13,12,23,34,52,56,23,44,32,11,11

Explain survey development process, Problem: A survey usually originate...

Problem: A survey usually originates when an individual or an institution is confronted with an information need and the existing data are  insufficient. Planning the questionn

Option price binomial tree, Modify your formulas from (1) to compute the pr...

Modify your formulas from (1) to compute the price at time 0 of an American put option with the same contract speci cations in the binomial model. Report the price of the American

Find the backward induction equilibrium, A rightist incumbent (player I) an...

A rightist incumbent (player I) and a leftist challenger (player C) run for senate. Each candidate chooses among two possible political platforms: Left or Right. The rules of the g

Confidence interval, a) List down several measures of central tendency and ...

a) List down several measures of central tendency and define the difference among them? b) What do you mean by confidence interval, and why it is useful? What is a confidence lev

PERCENTAGES, CALCULATE THE PERCENTAGE OF REFUNDS EXPECTED TO EXCEED $1000 U...

CALCULATE THE PERCENTAGE OF REFUNDS EXPECTED TO EXCEED $1000 UNDER THE CURRENT WITHHOLDING GUIDELINES

Sample, You want to know the thoughts of air travelers in fields such as ti...

You want to know the thoughts of air travelers in fields such as tickets, comffort, safety, securuty, services and economic growth. You are given a database and 20 questions to ask

Penman-monteith method, (a) Average rainfall during the month of January...

(a) Average rainfall during the month of January is found to be 58 mm. A Class A pan evaporation recorded an average of 8.12 mm/day near an irrigation reservoir. The average

Write Your Message!

Captcha
Free Assignment Quote

Assured A++ Grade

Get guaranteed satisfaction & time on delivery in every assignment order you paid with us! We ensure premium quality solution document along with free turntin report!

All rights reserved! Copyrights ©2019-2020 ExpertsMind IT Educational Pvt Ltd