Examine the model performance on the validation set

Assignment Help Computer Engineering
Reference no: EM131926086

Problem

Personal Loan Acceptance. Universal Bank is a relatively young bank growing rapidly in terms of overall customer acquisition. The majority of these customers are liability customers with varying sizes of relationship with the bank. The customer base of asset customers is quite small, and the bank is interested in expanding this base rapidly to bring in more loan business. In particular, it wants to explore ways of converting its liability customers to personal loan customers. A campaign the bank ran for liability customers last year showed a healthy conversion rate of over 9% successes. This has encouraged the retail marketing department to devise smarter campaigns with better target marketing. The goal of our analysis is to model the previous campaign's customer behavior to analyze what combination of factors make a customer more likely to accept a personal loan. This will serve as the basis for the design of a new campaign.

The file UniversalBank.csv contains data on 5000 customers. The data include customer demographic information (e.g., age, income), the customer's relationship with the bank (e.g., mortgage, securities account), and the customer response to the last personal loan campaign (Personal Loan). Among these 5000 customers, only 480 (= 9.6%) accepted the personal loan that was offered to them in the previous campaign. Partition the data (60% training and 40% validation) and then perform a discriminant analysis that models Personal Loan as a function of the remaining predictors (excluding zip code). Remember to turn categorical predictors with more than two categories into dummy variables first. Specify the success class as 1 (personal loan acceptance), and use the default cutoff value of 0.5.

a. Compute summary statistics for the predictors separately for loan acceptors and non-acceptors. For continuous predictors, compute the mean and standard deviation. For categorical predictors, compute the percentages. Are there predictors where the two classes differ substantially?

b. Examine the model performance on the validation set.

i. What is the accuracy rate?

ii. Is one type of misclassification more likely than the other?

iii. Select three customers who were misclassified as acceptors and three who were misclassified as non-acceptors. The goal is to determine why they are misclassified. First, examine their probability of being classified as acceptors: is it close to the threshold of 0.5? If not, compare their predictor values to the summary statistics of the two classes to determine why they were misclassified.

c. As in many marketing campaigns, it is more important to identify customers who will accept the offer rather than customers who will not accept it. Therefore, a good model should be especially accurate at detecting acceptors. Examine the lift chart and decile-wise lift chart for the validation set and interpret them in light of this ranking goal.

d. Compare the results from the discriminant analysis with those from a logistic regression (both with cutoff 0.5 and the same predictors). Examine the confusion matrices, the lift charts, and the decile charts. Which method performs better on your validation set in detecting the acceptors?

e. The bank is planning to continue its campaign by sending its offer to 1000 additional customers. Suppose that the cost of sending the offer is $1 and the profit from an accepted offer is $50. What is the expected profitability of this campaign?

f. The cost of misclassifying a loan acceptor customer as a non-acceptor is much higher than the opposite misclassification cost. To minimize the expected cost of misclassification, should the cutoff value for classification (which is currently at 0.5) be increased or decreased?

Reference no: EM131926086

Questions Cloud

Which repayment method results in higher home equity : What is the difference in total interest payments between the two alternative payment methods? Hint: The total of payment for Option 2 involves.
How much experience must be accumulated by an administrator : How much experience must be accumulated by an administrator with 4 training credits before his or her estimated probability of completing the tasks exceeds 0.5?
Identify a case and write a paper describing the cases focus : What role has each interest group played in American politics? Provide two (2) examples for each group.
Describe the formation of an aqueous ki : Describe the formation of an aqueous KI solution, when solid KI dissolves in water.
Examine the model performance on the validation set : Examine the model performance on the validation set. What is the accuracy rate? Is one type of misclassification more likely than the other?
Calculate time to disintegrate the sample : Calculate time to disintegrate the sample from an initial activity of 8.5 * 10^4 dpm to 7.4 * 10^3 dpm. Here, dpm is disintegrations per minute
Perform a cost benefit analysis on the proposed project : The City of Evanston is considering purchasing a new fleet of police cars. The initial cost of the fleet is $500,000.
Write down the formula for the value of the firm : Write down the formula for the value of the firm (V) at t=0 in terms of g and z. For z=0.195, calculate V for these cases: z=0.05, 0.125, 0.18, 0.19
Articles of confederation and the constitution embodied : Americans in the late eighteenth century believed that government must be kept weak in order to prevent the rise of tyranny and to guarantee

Reviews

Write a Review

Computer Engineering Questions & Answers

  List the top advantages of migrating to ipv6

List the top advantages of migrating to IPv6

  What is the difference between a sequential control

post a 200 to 300-word response to each of the the following question1. what is the difference between a sequential

  Given a relational database with a person table

Given a relational database with a person table that contains an ID, name, and age. What do each of the following return? Select * from person.

  What is the representation of the exponent

What is the representation of the Sign? What is the representation of the Exponent? What is the representation of the Significant?

  Examine the technical merits and demerits of using a

a hypervisor is computer hardware platform virtualization software that allows multiple different operating systems os

  Why words much clearer after running faux increase volume

Run increase Volume on the sound, and then only Maximize on the same sound. Why the words will be much clearer after running faux Increase Volume.

  How can value of a form element be accessed by a php script

How can the value of a form element be accessed by a PHP script? How can a script determine whether a particular cookie exists?

  Imagine that you''re the manager of a small project

suppose that you're the manager of a small project. What baselines would you define for the project and how would you control them, also state what are baselines?

  Describe the components of organizational culture

Name and describe the components of organizational culture. Name and describe the components of an organizational change management plan.

  Write a class called bar chart

Write a class called BarChart that compares the data using a bar graph representation. Allow the parameters to the constructor to specify the bar's width.

  Define the project in terms of the selected framework

Select and describe in detail framework that you used to define and implement system integration project. Define the project in terms of the selected framework.

  Calculate the beamwidth of a microwave dish antenna

Calculate the beamwidth of a microwave dish antenna with a 6-m mouth diameter when used at 5 GHz.

Free Assignment Quote

Assured A++ Grade

Get guaranteed satisfaction & time on delivery in every assignment order you paid with us! We ensure premium quality solution document along with free turntin report!

All rights reserved! Copyrights ©2019-2020 ExpertsMind IT Educational Pvt Ltd